spb/datacenterindex
Public
HTML 53.9%
TypeScript 44.5%
JavaScript 0.6%
SQL 0.5%
1# DataBank (`databank`)23- **Domain**: `www.databank.com` · **Kind**: operator · **Priority**: 1 · **Enabled**: yes4- **License / attribution**: Publicly accessible operator pages; factual data only; © DataBank — attribution “DataBank”5- **Coverage**: US, GB67## What it collects89Facility records (`kind: facility`) via parser(s) `databank_facility_v1` in `apps/worker/src/connectors/operators2/americas.ts`.1011Fields published by the pages and captured: name, code (ATL6), address, city, regionName, postalCode, countryIso2, itCapacityMw (Critical IT Load), buildingSqm (IT sq ft), coolingType, renewableClaim, certifications, description.1213Every field carries an extraction method in `methods` (e.g. `json-ld:PostalAddress`, `regex:it_mw_v1`, `attr:data-value+MW`, `lookup:city-country`) for provenance. Coordinates are only taken from JSON-LD / embedded JSON / map links published by the operator — never geocoded. Status defaults to `operational` unless the page uses pipeline wording (under construction / coming soon / planned).1415## Discovery1617- Sitemaps: `https://www.databank.com/db_data_center-sitemap.xml`, `https://www.databank.com/db-press-release-sitemap.xml`, `https://www.databank.com/db-news-sitemap.xml`18- Include: `/data-centers/[a-z0-9-]+/`, `/press-release`, `/news/`19- Exclude: `\?`, `#`, `/data-centers/data-center-design/`, `/data-centers-map/`20- Classify: `/data-centers/[a-z0-9-]+/[a-z0-9-]+-campus/$` → facility_index (facility_index); `/data-centers/[a-z0-9-]+/(?:[a-z0-9-]+/)?[a-z0-9-]+/$` → facility_pages (facility_page); `/data-centers/[a-z0-9-]+/$` → facility_index (facility_index); `/(press-release|news)` → newsroom (press_release)2122Newsroom URLs (when configured) rely on the GenericConnector news fallback (press_release → `news_event` when data-center relevant).2324## Schedule & fetch2526- Schedule: facility_pages weekly, facility_index monthly, newsroom daily, sitemap weekly27- Fetch: L1→L2, 20 rpm, concurrency 2, robots.txt respected, max 50 premium credits/run. All pages are served at L1 (plain HTTP); `maxLevel` is capped at 2 so false-positive “blocked” heuristics (HubSpot / cookie-consent strings) never spend Scrapfly credits.2829## Quirks3031Facility pages /data-centers/<metro>/[<campus>/]<facility>/ publish code (ATL6), street address, IT sq ft and critical IT load. JSON-LD on these pages is the corporate HQ and is ignored.3233## Verification (live, 2026-09-11)3435`try-connector --limit N`: discovered 572 URLs (facility_index=39, facility_pages=77, newsroom=456); sampled pages → 3 entities, 3 valid, 0 rejected, 0 premium credits.36