# DataCenterIndex — full repository & data audit (2026-09-11) Internal audit written before the "infrastructure intelligence graph" upgrade. Sections A–K follow the brief. Figures come from the production database on BHS128 (`deploy/bin/psql.sh`) on 2026-09-11 evening, one day after launch. ## A. Current architecture - **Monorepo** pnpm + TypeScript strict/ESM: `packages/core` (types, ids, normalize, geo, confidence, ssrf, api contract), `packages/db` (Drizzle schema, plain SQL migrations, ClickHouse DDL), `packages/connectors` (connector SDK), `apps/worker` (runtime, ingest, scheduler, rankings, CLI), `apps/api` (Fastify 5, `/api/v1` public + `/api/admin` token), `apps/web` (Next 16 App Router, Tailwind v4, MapLibre + OpenFreeMap, hand-rolled SVG charts). - **Rendering**: every public page is ISR (`revalidate` 30–300 s) on top of a typed API client that never throws; map page is a static shell with client fetches to `/api/v1/map` (server-side clustering, ≤ 5 000 points). - **Deployment**: two OVH servers joined by WireGuard. BHS128 = data/public (Caddy edge :8300 → web :8310 / api :8311, Postgres 17, ClickHouse, Redis, MinIO, Prometheus, Grafana). BHS64b = crawl worker (concurrency 6). Public route via the MacLustr Tunnel gateway BHS64. Nightly backup timer (pg_dump, config tar, ClickHouse backup best-effort, MinIO mirror, rsync off-node to BHS64b). No CI, no Alertmanager. - **Finding (ops)**: the production `scheduler` container is crash-looping (851 restarts, exit 0, no log) because the deployed image predates commit `ff67cee` (standalone scheduler entry). Crawling continues because the BHS64b worker embeds the scheduler loop. Fixed by the redeploy at the end of this upgrade. ## B. Current database schema (Postgres) 28 tables. Core: `facilities` (2 912 rows; `it_capacity_mw`, `total_power_mw`, `planned_power_mw`, `mw_is_estimate`, `geo_precision`, `status`, `confidence`, `completeness`, `campus_id`, `merged_into`), `operators` (470), `projects` (324; `planned_mw`, `investment_usd`, `status`), `campuses` (114), `metros` (214), `countries` (250), `cloud_regions` (307), `ixps` (12), `facility_ixps`, `facility_tenants`, `facility_aliases`, `entity_keys` (connector-scoped keys, **global PK on `key`**), `provenance` (38 533 rows, one row per entity×field×source×url, `is_current`, `method`, `extractor_version`), `events` (5 549, fingerprint-deduplicated, `significance` 0–100), `news_items` (4 404), `documents` (7 092) + `document_versions` (8 785, content-hash versions with `diff_summary`), `entity_matches` (1 730: 1 200 auto_created, 81 auto_merged, **449 pending never reviewed**), `rankings` (snapshots with `min_coverage`), `daily_metrics`, `sources`, `connectors`, `connector_runs`, `connector_state`, `project_timeline`, `system_alerts`, `admin_sessions`. ClickHouse holds append-only `observations`, `crawl_log`, `page_changes`, `entity_daily`, `api_requests` (best-effort, silently dropped on failure). Indexes: btree on the obvious FKs/status, trigram GIN on names/cities, generated tsvector on facilities, partial `(lat,lng)`. **No PostGIS** (haversine + bbox pre-filter + geohash). No `claims`/observation table in Postgres; no entity-version table; no `run_id` on `provenance`/`document_versions`. ## C. Current crawler architecture Declarative YAML per source (`config/connectors/*.yaml`) → `GenericConnector` or registered parser/implementation → `discover → fetch (L1 direct → L2 browser identity → L3 Firecrawl → L4 Scrapfly) → archive (MinIO, zstd, per content hash) → extract → normalize → validate → ingest (reconcile, provenance, events)`. Robots.txt honoured, per-host token bucket, conditional GET, content-hash change detection, extraction skipped when hash and `extractor_version` unchanged, adaptive `next_check` with EMA change score, per-document quarantine after 3 × 404/410. BullMQ `crawl` + `maintenance` queues, Redis run locks, per-run/per-connector/per-provider premium budgets, Prometheus metrics, `dci doctor`. ## D. Current connectors 104 YAML configs (87 enabled): 60 operator directories, 20 industry/hyperscaler newsrooms (RSS), 10 cloud-region JSON, 6 government, 3 utilities, 3 datasets (PeeringDB — disabled pending AUP approval —, Wikidata, World Bank), SEC EDGAR full-text, OSM Overpass. Facility rows by connector: OSM 1 346, Equinix 262, Digital Realty 259, Wikidata 190, STACK 78, DataBank 75, NTT 72, EdgeConneX 59, CyrusOne 52, Google 47, QTS 47 … 58 named parsers; **zero configs use the declarative extractor** (dead code path); **zero HTML fixtures**; `packages/core` and `packages/connectors` have no tests. ## E. Current data-quality issues (measured) Facilities: 2 640 operational / 154 unknown / 76 announced / 37 under construction; only **536 (18 %) have any MW**; 797 have no coordinates; 1 377 have `facility_type = unknown`; only 4 flagged AI; 109 duplicate normalized-name pairs and 469 coordinate pairs within ~100 m (OSM building-vs-campus, e.g. atNorth ICE02 vs buildings M01…M16). Confirmed extraction errors: - **Decimal MW parsed as thousands**: DataBank MSP4 shows 779 MW; source says "0.779MW Critical IT Load". Same bug hits DFW6 (675), AUS1 (405) and any "0.xxx MW" figure (`parseMw` strips the dot when three decimals follow). - **Headlines stored as project names**: 182/324 projects are named after an article title. - **Executive appointments as projects**: "STACK Appoints Matt VanderZanden as CEO" = 13 000 MW project; "AirTrunk appoints Laura Coad" = 1 400 MW; VIRTUS, 365 Data Centers, Element Critical, Cassava appointments likewise. - **Portfolio/company figures on one project**: "Microsoft Georgia data center project" 12 GW / $175 B (company capex); "Google Alabama" 17 GW (Southern Company pipeline); Nvidia "Washington" 20 GW (Starcloud funding story). - **Market-research articles as projects**: "Hyperscale Data Center Market to Surpass USD 624.2 Billion" ($80.9 B). - **Wrong operator**: "Meta's Canadian AI Data Center" assigned to Google. - **PPA / sustainability / HQ / workforce stories as projects** (Ormat–Switch PPA 3 400 MW, AirTrunk HQ 1 400 MW, Meta Workforce Academy 5 000 MW). - **Projects have no coordinates at all** (0/324), 158 have a country, 0 link to a facility. - **Broken operator slug** `item` for 中華電信數據通信分公司 (26 facilities). - Ranking coverage gate leaves 3 countries in `countries_known_mw` — the MW coverage story must be visible, not hidden. Structural causes (code): one-row-per-entity with no claim scope; `isIdentifyingExternalId` folds unrelated facilities on `*_code/*_url` keys; `loadCandidates` truncates at 400 rows without ordering; news country via unsafe `countryFromText` ("North America" → US); no HQ guard on project location; `projects.country_iso2` write-once; campus/building double counting in every MW aggregate; equal-authority ping-pong after 30 days creates event churn; `entity_keys.key` global PK; `parserVersion` is `v1` everywhere so parser fixes never invalidate cached extractions; `healthFrom(0,0) = ok`. ## F. Current frontend strengths Unified status/confidence/precision visual language (CSS variables shared by badges, charts and MapLibre paint); per-field provenance popovers + provenance table + source history + sources footer with licences; dense typography with mono tabular figures; URL-as-state everywhere; ⌘K/`/` command palette with API-side query interpretation; server-clustered map with precision rings; full SEO (canonical, per-entity OG images, JSON-LD, sharded sitemaps); a11y depth; an automated UX sweep harness; a substantial `/admin` console (connectors, documents, matches, events, dev tool). ## G. Current frontend weaknesses No explore/query builder, no compare, no evidence drawer (only popover + bottom table), no export, no watchlist, no coverage page, no AI index, no pulse, no power/connectivity layers, no map density/time modes, no virtualization, long single-scroll detail pages without active-section tracking, hand-rolled charts without brushing, keyboard story stops at the palette, project pages lack a state machine/timeline funnel, MapLibre controls < 40 px on touch. ## H. Current API strengths 27 public GET endpoints, envelope `{data, meta, sources}`, zod-validated params, weak ETags + `s-maxage`, rate limits, consistent 404/400 JSON, OpenAPI + Swagger UI, `/map` zoom-tiered clustering, `/search` with interpretation, admin API with constant-time token compare and same-origin proxy with CSRF guard. Weaknesses: no provenance/history/nearby/coverage/export endpoints; list responses carry no `sources`; OpenAPI has no response schemas; placeholder-token check is a no-op. ## I. Schema changes proposed (migration 0003) 1. `claims` — claim-first store: subject (type,id), predicate, value/unit, **scope** (building/facility/campus/metro/ country/portfolio/company/unknown), source/document/url, published/retrieved, confidence, is_estimate, evidence text + offsets, parser name/version, `run_id`, status (current/superseded/rejected/review/unscoped), rejection reason. 2. `quality_flags` — deterministic sanity flags (capacity, investment, scope, location, operator, duplicate, project false-positive) with severity, priority score, status, resolution, linking to entity + claim. 3. `facilities.parent_facility_id`, `facilities.record_scope` (building/facility/campus) for containment-aware aggregation; `facilities.ai_evidence` (confirmed/likely/associated/unknown); `facilities.utility_capacity_mw`, `grid_connection_mw`, `ultimate_campus_mw` (capacity ontology columns; IT/total/planned kept). 4. `projects.project_class` (NEW_BUILD … EXECUTIVE_APPOINTMENT), `projects.evidence_level`, `projects.merged_into`, `projects.lat/lng` geocoded from metro/city with `geo_precision = city`, `projects.ai_evidence`, `projects.investment_scope`. 5. `document_versions.run_id`, `provenance.run_id`, `provenance.scope`; `connectors.quarantine` (bool) + `consecutive_failures`, `blocked_since`. 6. `field_authority` (per-field source-kind tiers) as code table in `packages/core` + methodology page (no table needed). 7. `entity_snapshots` (daily JSON snapshot of global totals, rankings, facility status counts, project stage counts) for "as of" views and regression checks. 8. Indexes: `provenance(is_current)`, `facilities(merged_into)`, `facilities(campus_id)`, `facilities(parent_facility_id)`, `events(significance, detected_at)`, `daily_metrics(metric, dim, day)`, GIN on `news_items.operator_ids`. ## J. Components to preserve Everything in section F/H, the connector SDK and escalation ladder, the content-hash/archive pipeline, entity keys, provenance table, event fingerprinting, the matcher (with fixes), rankings coverage gates, the design tokens, the map architecture, the admin console, deploy scripts, backups, the QA sweep. ## K. Implementation order Phase 1 data quality (parseMw fix, claim/scope layer, project classifier + evidence threshold, capacity/investment sanity engine, external-id allowlist, HQ/country guards, campus containment aggregation, extraction debugger, connector health, fixtures, cleanup of current bad records) → Phase 2 core UX (home 2.0, map 2.0, facility 3.0, operator/market/country/project 2.0, brand, mobile) → Phase 3 intelligence (AI index, pulse, live feed 2.0, compare, explore, coverage, quality dashboards) → Phase 4 graph (connectivity, power, cloud, IXP pages) → Phase 5 historical (time machine, capacity history, announced vs delivered, velocity) → API 2.0 + docs + download → regression checks, security fixes, deploy, final data audit.