# Digital Edge (`digitaledge`) - **Domain**: `www.digitaledgedc.com` · **Kind**: operator · **Priority**: 2 · **Enabled**: yes - **License / attribution**: Publicly accessible operator pages; factual data only; © Digital Edge (Singapore) Holdings Pte Ltd — attribution “Digital Edge” - **Coverage**: JP, KR, IN, ID, PH, CN, TH ## What it collects Facility records (`kind: facility`) via parser(s) `digitaledge_facility_v1` in `apps/worker/src/connectors/operators2/apac.ts`. Fields published by the pages and captured: name, code, city, countryIso2, plannedPowerMw (campus at full capacity), itCapacityMw when stated, tier, status, description. Every field carries an extraction method in `methods` (e.g. `json-ld:PostalAddress`, `regex:it_mw_v1`, `attr:data-value+MW`, `lookup:city-country`) for provenance. Coordinates are only taken from JSON-LD / embedded JSON / map links published by the operator — never geocoded. Status defaults to `operational` unless the page uses pipeline wording (under construction / coming soon / planned). ## Discovery - Sitemaps: `https://www.digitaledgedc.com/data-center-sitemap.xml`, `https://www.digitaledgedc.com/news-sitemap.xml` - Include: `digitaledgedc\.com/resources/data-center/[a-z]+\d/$`, `digitaledgedc\.com/news/.` - Exclude: `\?`, `#`, `/(cn|jp)/` - Classify: `/resources/data-center/[a-z]+\d/$` → facility_pages (facility_page); `/news/.` → newsroom (press_release) Newsroom URLs (when configured) rely on the GenericConnector news fallback (press_release → `news_event` when data-center relevant). ## Schedule & fetch - Schedule: facility_pages weekly, newsroom daily, sitemap weekly - Fetch: L1→L2, 15 rpm, concurrency 2, robots.txt respected, max 20 premium credits/run. All pages are served at L1 (plain HTTP); `maxLevel` is capped at 2 so false-positive “blocked” heuristics (HubSpot / cookie-consent strings) never spend Scrapfly credits. ## Quirks English facility pages /resources/data-center// publish code, city and description; campus 'MW at full capacity' is recorded as planned power. CN/JP duplicates excluded. robots.txt only disallows some spec-sheet PDFs. ## Verification (live, 2026-09-11) `try-connector --limit N`: discovered 14 URLs (facility_pages=14); sampled pages → 4 entities, 4 valid, 0 rejected, 0 premium credits.