spb/datacenterindex
Public
HTML 53.9%
TypeScript 44.5%
JavaScript 0.6%
SQL 0.5%
1# DataCenterIndex — full repository & data audit (2026-09-11)23Internal audit written before the "infrastructure intelligence graph" upgrade. Sections A–K follow the brief. Figures come4from the production database on BHS128 (`deploy/bin/psql.sh`) on 2026-09-11 evening, one day after launch.56## A. Current architecture78- **Monorepo** pnpm + TypeScript strict/ESM: `packages/core` (types, ids, normalize, geo, confidence, ssrf, api contract),9 `packages/db` (Drizzle schema, plain SQL migrations, ClickHouse DDL), `packages/connectors` (connector SDK), `apps/worker`10 (runtime, ingest, scheduler, rankings, CLI), `apps/api` (Fastify 5, `/api/v1` public + `/api/admin` token), `apps/web`11 (Next 16 App Router, Tailwind v4, MapLibre + OpenFreeMap, hand-rolled SVG charts).12- **Rendering**: every public page is ISR (`revalidate` 30–300 s) on top of a typed API client that never throws; map page13 is a static shell with client fetches to `/api/v1/map` (server-side clustering, ≤ 5 000 points).14- **Deployment**: two OVH servers joined by WireGuard. BHS128 = data/public (Caddy edge :8300 → web :8310 / api :8311,15 Postgres 17, ClickHouse, Redis, MinIO, Prometheus, Grafana). BHS64b = crawl worker (concurrency 6). Public route via the16 MacLustr Tunnel gateway BHS64. Nightly backup timer (pg_dump, config tar, ClickHouse backup best-effort, MinIO mirror,17 rsync off-node to BHS64b). No CI, no Alertmanager.18- **Finding (ops)**: the production `scheduler` container is crash-looping (851 restarts, exit 0, no log) because the deployed19 image predates commit `ff67cee` (standalone scheduler entry). Crawling continues because the BHS64b worker embeds the20 scheduler loop. Fixed by the redeploy at the end of this upgrade.2122## B. Current database schema (Postgres)232428 tables. Core: `facilities` (2 912 rows; `it_capacity_mw`, `total_power_mw`, `planned_power_mw`, `mw_is_estimate`,25`geo_precision`, `status`, `confidence`, `completeness`, `campus_id`, `merged_into`), `operators` (470), `projects` (324;26`planned_mw`, `investment_usd`, `status`), `campuses` (114), `metros` (214), `countries` (250), `cloud_regions` (307),27`ixps` (12), `facility_ixps`, `facility_tenants`, `facility_aliases`, `entity_keys` (connector-scoped keys, **global PK on28`key`**), `provenance` (38 533 rows, one row per entity×field×source×url, `is_current`, `method`, `extractor_version`),29`events` (5 549, fingerprint-deduplicated, `significance` 0–100), `news_items` (4 404), `documents` (7 092) +30`document_versions` (8 785, content-hash versions with `diff_summary`), `entity_matches` (1 730: 1 200 auto_created, 8131auto_merged, **449 pending never reviewed**), `rankings` (snapshots with `min_coverage`), `daily_metrics`, `sources`,32`connectors`, `connector_runs`, `connector_state`, `project_timeline`, `system_alerts`, `admin_sessions`.33ClickHouse holds append-only `observations`, `crawl_log`, `page_changes`, `entity_daily`, `api_requests` (best-effort, silently34dropped on failure).3536Indexes: btree on the obvious FKs/status, trigram GIN on names/cities, generated tsvector on facilities, partial `(lat,lng)`.37**No PostGIS** (haversine + bbox pre-filter + geohash). No `claims`/observation table in Postgres; no entity-version table;38no `run_id` on `provenance`/`document_versions`.3940## C. Current crawler architecture4142Declarative YAML per source (`config/connectors/*.yaml`) → `GenericConnector` or registered parser/implementation →43`discover → fetch (L1 direct → L2 browser identity → L3 Firecrawl → L4 Scrapfly) → archive (MinIO, zstd, per content hash)44→ extract → normalize → validate → ingest (reconcile, provenance, events)`. Robots.txt honoured, per-host token bucket,45conditional GET, content-hash change detection, extraction skipped when hash and `extractor_version` unchanged, adaptive46`next_check` with EMA change score, per-document quarantine after 3 × 404/410. BullMQ `crawl` + `maintenance` queues,47Redis run locks, per-run/per-connector/per-provider premium budgets, Prometheus metrics, `dci doctor`.4849## D. Current connectors5051104 YAML configs (87 enabled): 60 operator directories, 20 industry/hyperscaler newsrooms (RSS), 10 cloud-region JSON,526 government, 3 utilities, 3 datasets (PeeringDB — disabled pending AUP approval —, Wikidata, World Bank), SEC EDGAR53full-text, OSM Overpass. Facility rows by connector: OSM 1 346, Equinix 262, Digital Realty 259, Wikidata 190, STACK 78,54DataBank 75, NTT 72, EdgeConneX 59, CyrusOne 52, Google 47, QTS 47 … 58 named parsers; **zero configs use the declarative55extractor** (dead code path); **zero HTML fixtures**; `packages/core` and `packages/connectors` have no tests.5657## E. Current data-quality issues (measured)5859Facilities: 2 640 operational / 154 unknown / 76 announced / 37 under construction; only **536 (18 %) have any MW**;60797 have no coordinates; 1 377 have `facility_type = unknown`; only 4 flagged AI; 109 duplicate normalized-name pairs and61469 coordinate pairs within ~100 m (OSM building-vs-campus, e.g. atNorth ICE02 vs buildings M01…M16).6263Confirmed extraction errors:64- **Decimal MW parsed as thousands**: DataBank MSP4 shows 779 MW; source says "0.779MW Critical IT Load". Same bug hits65 DFW6 (675), AUS1 (405) and any "0.xxx MW" figure (`parseMw` strips the dot when three decimals follow).66- **Headlines stored as project names**: 182/324 projects are named after an article title.67- **Executive appointments as projects**: "STACK Appoints Matt VanderZanden as CEO" = 13 000 MW project;68 "AirTrunk appoints Laura Coad" = 1 400 MW; VIRTUS, 365 Data Centers, Element Critical, Cassava appointments likewise.69- **Portfolio/company figures on one project**: "Microsoft Georgia data center project" 12 GW / $175 B (company capex);70 "Google Alabama" 17 GW (Southern Company pipeline); Nvidia "Washington" 20 GW (Starcloud funding story).71- **Market-research articles as projects**: "Hyperscale Data Center Market to Surpass USD 624.2 Billion" ($80.9 B).72- **Wrong operator**: "Meta's Canadian AI Data Center" assigned to Google.73- **PPA / sustainability / HQ / workforce stories as projects** (Ormat–Switch PPA 3 400 MW, AirTrunk HQ 1 400 MW,74 Meta Workforce Academy 5 000 MW).75- **Projects have no coordinates at all** (0/324), 158 have a country, 0 link to a facility.76- **Broken operator slug** `item` for 中華電信數據通信分公司 (26 facilities).77- Ranking coverage gate leaves 3 countries in `countries_known_mw` — the MW coverage story must be visible, not hidden.7879Structural causes (code): one-row-per-entity with no claim scope; `isIdentifyingExternalId` folds unrelated facilities on80`*_code/*_url` keys; `loadCandidates` truncates at 400 rows without ordering; news country via unsafe `countryFromText`81("North America" → US); no HQ guard on project location; `projects.country_iso2` write-once; campus/building double counting82in every MW aggregate; equal-authority ping-pong after 30 days creates event churn; `entity_keys.key` global PK;83`parserVersion` is `v1` everywhere so parser fixes never invalidate cached extractions; `healthFrom(0,0) = ok`.8485## F. Current frontend strengths8687Unified status/confidence/precision visual language (CSS variables shared by badges, charts and MapLibre paint);88per-field provenance popovers + provenance table + source history + sources footer with licences; dense typography with89mono tabular figures; URL-as-state everywhere; ⌘K/`/` command palette with API-side query interpretation; server-clustered90map with precision rings; full SEO (canonical, per-entity OG images, JSON-LD, sharded sitemaps); a11y depth; an automated91UX sweep harness; a substantial `/admin` console (connectors, documents, matches, events, dev tool).9293## G. Current frontend weaknesses9495No explore/query builder, no compare, no evidence drawer (only popover + bottom table), no export, no watchlist, no96coverage page, no AI index, no pulse, no power/connectivity layers, no map density/time modes, no virtualization,97long single-scroll detail pages without active-section tracking, hand-rolled charts without brushing, keyboard story stops98at the palette, project pages lack a state machine/timeline funnel, MapLibre controls < 40 px on touch.99100## H. Current API strengths10110227 public GET endpoints, envelope `{data, meta, sources}`, zod-validated params, weak ETags + `s-maxage`, rate limits,103consistent 404/400 JSON, OpenAPI + Swagger UI, `/map` zoom-tiered clustering, `/search` with interpretation, admin API with104constant-time token compare and same-origin proxy with CSRF guard. Weaknesses: no provenance/history/nearby/coverage/export105endpoints; list responses carry no `sources`; OpenAPI has no response schemas; placeholder-token check is a no-op.106107## I. Schema changes proposed (migration 0003)1081091. `claims` — claim-first store: subject (type,id), predicate, value/unit, **scope** (building/facility/campus/metro/110 country/portfolio/company/unknown), source/document/url, published/retrieved, confidence, is_estimate, evidence text +111 offsets, parser name/version, `run_id`, status (current/superseded/rejected/review/unscoped), rejection reason.1122. `quality_flags` — deterministic sanity flags (capacity, investment, scope, location, operator, duplicate, project113 false-positive) with severity, priority score, status, resolution, linking to entity + claim.1143. `facilities.parent_facility_id`, `facilities.record_scope` (building/facility/campus) for containment-aware aggregation;115 `facilities.ai_evidence` (confirmed/likely/associated/unknown); `facilities.utility_capacity_mw`, `grid_connection_mw`,116 `ultimate_campus_mw` (capacity ontology columns; IT/total/planned kept).1174. `projects.project_class` (NEW_BUILD … EXECUTIVE_APPOINTMENT), `projects.evidence_level`, `projects.merged_into`,118 `projects.lat/lng` geocoded from metro/city with `geo_precision = city`, `projects.ai_evidence`, `projects.investment_scope`.1195. `document_versions.run_id`, `provenance.run_id`, `provenance.scope`; `connectors.quarantine` (bool) +120 `consecutive_failures`, `blocked_since`.1216. `field_authority` (per-field source-kind tiers) as code table in `packages/core` + methodology page (no table needed).1227. `entity_snapshots` (daily JSON snapshot of global totals, rankings, facility status counts, project stage counts) for123 "as of" views and regression checks.1248. Indexes: `provenance(is_current)`, `facilities(merged_into)`, `facilities(campus_id)`, `facilities(parent_facility_id)`,125 `events(significance, detected_at)`, `daily_metrics(metric, dim, day)`, GIN on `news_items.operator_ids`.126127## J. Components to preserve128129Everything in section F/H, the connector SDK and escalation ladder, the content-hash/archive pipeline, entity keys,130provenance table, event fingerprinting, the matcher (with fixes), rankings coverage gates, the design tokens, the map131architecture, the admin console, deploy scripts, backups, the QA sweep.132133## K. Implementation order134135Phase 1 data quality (parseMw fix, claim/scope layer, project classifier + evidence threshold, capacity/investment sanity136engine, external-id allowlist, HQ/country guards, campus containment aggregation, extraction debugger, connector health,137fixtures, cleanup of current bad records) → Phase 2 core UX (home 2.0, map 2.0, facility 3.0, operator/market/country/project1382.0, brand, mobile) → Phase 3 intelligence (AI index, pulse, live feed 2.0, compare, explore, coverage, quality dashboards)139→ Phase 4 graph (connectivity, power, cloud, IXP pages) → Phase 5 historical (time machine, capacity history, announced vs140delivered, velocity) → API 2.0 + docs + download → regression checks, security fixes, deploy, final data audit.141