spb/datacenterindex
Public
HTML 53.9%
TypeScript 44.5%
JavaScript 0.6%
SQL 0.5%
1# stack — STACK Infrastructure campuses23| | |4|---|---|5| Config | `config/connectors/stack.yaml` |6| Parser | `apps/worker/src/connectors/operators1/stack.ts` → `stack_campus_v1` |7| Kind / priority | operator / 1 |8| License / attribution | Publicly accessible operator pages; factual data only; © STACK Infrastructure |9| Cadence | facility pages weekly · newsroom daily · index monthly · sitemap weekly |10| Default facility type | wholesale, `isHyperscale: true` (operatorKind developer) |1112## What is crawled13- Yoast sitemaps `/campuses-sitemap.xml` (48 campus pages `/locations/<region>/<metro>/<code>/`) and `/news-sitemap.xml` (154); RSS `/feed/`. `/location-sitemap.xml` metro pages → `index`; `/japan/locations/*` (Japanese duplicates of KIX01/TKY01) excluded.14- News `/about/news-press/(press-releases|news)/<slug>/` → GenericConnector news fallback.15- robots.txt has no disallows.1617## Access — maxLevel 2 on purpose18Every stackinfra.com page embeds Cloudflare's `/cdn-cgi/challenge-platform/` beacon script although the page is served in full (HTTP 200 via L1/L2). The shared `looksBlockedOrEmpty()` heuristic (`packages/connectors/src/fetchers.ts`) treats that string as a bot wall, so with `maxLevel: 4` every campus page would be escalated to Scrapfly (30 credits × ~50 pages per run). `fetch.maxLevel` is therefore capped at 2 (direct HTTP only, browser identity fallback), which is all the site needs. Re-raise once the heuristic only fires on short pages.1920## Records21One page → one **campus** record + one record per **building** card:22- Campus: `name` "STACK ATL01 – Alpharetta" (`<h1>ATL01 – Alpharetta</h1>`), `code` ATL01, `campusName` "STACK ATL01", `totalPowerMw` ← campus tile "20MW", `siteAreaHa` ← "18 Acres" / "7 hectares", `plannedPowerMw` ← "designed for 96MW" when it differs, `description` ← the "<CODE> Campus" intro paragraph (the Yoast meta description concatenates page chrome and is skipped when it contains "Schedule a tour"). Key `stack:atl01`.23- Buildings: `h3.title-campus` (ATL01A, ATL01B) with `div.flex.items-end` values "105,000 SQ FT" → `buildingSqm`, "8MW" → `totalPowerMw`; `campusName` = the campus. Key `stack:atl01a`.24- `city` ← h1 after the dash; `regionName` ← gallery caption "ATL01 Campus – Alpharetta, GA" (US state / CA province) → `countryIso2`; otherwise country from the metro segment of the URL via a fixed table of STACK metros (atlanta→US, toronto→CA, milan→IT, zurich/geneva→CH, oslo→NO, frankfurt→DE, copenhagen→DK, stockholm→SE, tokyo/osaka→JP, sydney/melbourne→AU, johor-bahru→MY). Deterministic mapping, not geocoding.25- No street address or coordinates are published → geo null.2627## Verification (2026-09-11)28- Discovery: 202 URLs → `facility_pages=48`, `newsroom=153`, index 1.29- `--url` ATL01: campus 20 MW, 7.28 ha, Alpharetta, GA, US + ATL01A (9,755 m², 8 MW) + ATL01B (12,356 m², 12 MW) — matches the live page.30- `--url` FRA01: campus 96 MW, 6.88 ha (17 acres), Frankfurt, DE + FRA01A 68,262 m² (734,766 sq ft).31- Sweep of 8 spread campus pages → 20 entities (ATL01, MIL06, NVA02 420 MW, MEL01 + 4 buildings, OSL04 + 3, ZUR03, DFW02 500 MW, NVA01): 20/20 valid, all with country. Newsroom sample 4 pages → 6 entities (4 news + 2 project candidates), all valid. Credits: 0.3233## Gaps / TODO34- Campus "critical capacity" figures are design/full-build numbers as published by STACK; buildings without a MW card keep null.35- `maxLevel: 2` workaround (see Access).36