# stack — STACK Infrastructure campuses | | | |---|---| | Config | `config/connectors/stack.yaml` | | Parser | `apps/worker/src/connectors/operators1/stack.ts` → `stack_campus_v1` | | Kind / priority | operator / 1 | | License / attribution | Publicly accessible operator pages; factual data only; © STACK Infrastructure | | Cadence | facility pages weekly · newsroom daily · index monthly · sitemap weekly | | Default facility type | wholesale, `isHyperscale: true` (operatorKind developer) | ## What is crawled - Yoast sitemaps `/campuses-sitemap.xml` (48 campus pages `/locations////`) and `/news-sitemap.xml` (154); RSS `/feed/`. `/location-sitemap.xml` metro pages → `index`; `/japan/locations/*` (Japanese duplicates of KIX01/TKY01) excluded. - News `/about/news-press/(press-releases|news)//` → GenericConnector news fallback. - robots.txt has no disallows. ## Access — maxLevel 2 on purpose Every stackinfra.com page embeds Cloudflare's `/cdn-cgi/challenge-platform/` beacon script although the page is served in full (HTTP 200 via L1/L2). The shared `looksBlockedOrEmpty()` heuristic (`packages/connectors/src/fetchers.ts`) treats that string as a bot wall, so with `maxLevel: 4` every campus page would be escalated to Scrapfly (30 credits × ~50 pages per run). `fetch.maxLevel` is therefore capped at 2 (direct HTTP only, browser identity fallback), which is all the site needs. Re-raise once the heuristic only fires on short pages. ## Records One page → one **campus** record + one record per **building** card: - Campus: `name` "STACK ATL01 – Alpharetta" (`

ATL01 – Alpharetta

`), `code` ATL01, `campusName` "STACK ATL01", `totalPowerMw` ← campus tile "20MW", `siteAreaHa` ← "18 Acres" / "7 hectares", `plannedPowerMw` ← "designed for 96MW" when it differs, `description` ← the " Campus" intro paragraph (the Yoast meta description concatenates page chrome and is skipped when it contains "Schedule a tour"). Key `stack:atl01`. - Buildings: `h3.title-campus` (ATL01A, ATL01B) with `div.flex.items-end` values "105,000 SQ FT" → `buildingSqm`, "8MW" → `totalPowerMw`; `campusName` = the campus. Key `stack:atl01a`. - `city` ← h1 after the dash; `regionName` ← gallery caption "ATL01 Campus – Alpharetta, GA" (US state / CA province) → `countryIso2`; otherwise country from the metro segment of the URL via a fixed table of STACK metros (atlanta→US, toronto→CA, milan→IT, zurich/geneva→CH, oslo→NO, frankfurt→DE, copenhagen→DK, stockholm→SE, tokyo/osaka→JP, sydney/melbourne→AU, johor-bahru→MY). Deterministic mapping, not geocoding. - No street address or coordinates are published → geo null. ## Verification (2026-09-11) - Discovery: 202 URLs → `facility_pages=48`, `newsroom=153`, index 1. - `--url` ATL01: campus 20 MW, 7.28 ha, Alpharetta, GA, US + ATL01A (9,755 m², 8 MW) + ATL01B (12,356 m², 12 MW) — matches the live page. - `--url` FRA01: campus 96 MW, 6.88 ha (17 acres), Frankfurt, DE + FRA01A 68,262 m² (734,766 sq ft). - Sweep of 8 spread campus pages → 20 entities (ATL01, MIL06, NVA02 420 MW, MEL01 + 4 buildings, OSL04 + 3, ZUR03, DFW02 500 MW, NVA01): 20/20 valid, all with country. Newsroom sample 4 pages → 6 entities (4 news + 2 project candidates), all valid. Credits: 0. ## Gaps / TODO - Campus "critical capacity" figures are design/full-build numbers as published by STACK; buildings without a MW card keep null. - `maxLevel: 2` workaround (see Access).