DataCenterIndex — full repository & data audit (2026-09-11)
Internal audit written before the "infrastructure intelligence graph" upgrade. Sections A–K follow the brief. Figures come
from the production database on BHS128 (deploy/bin/psql.sh) on 2026-09-11 evening, one day after launch.
A. Current architecture
- Monorepo pnpm + TypeScript strict/ESM:
packages/core(types, ids, normalize, geo, confidence, ssrf, api contract),packages/db(Drizzle schema, plain SQL migrations, ClickHouse DDL),packages/connectors(connector SDK),apps/worker(runtime, ingest, scheduler, rankings, CLI),apps/api(Fastify 5,/api/v1public +/api/admintoken),apps/web(Next 16 App Router, Tailwind v4, MapLibre + OpenFreeMap, hand-rolled SVG charts). - Rendering: every public page is ISR (
revalidate30–300 s) on top of a typed API client that never throws; map page is a static shell with client fetches to/api/v1/map(server-side clustering, ≤ 5 000 points). - Deployment: two OVH servers joined by WireGuard. BHS128 = data/public (Caddy edge :8300 → web :8310 / api :8311, Postgres 17, ClickHouse, Redis, MinIO, Prometheus, Grafana). BHS64b = crawl worker (concurrency 6). Public route via the MacLustr Tunnel gateway BHS64. Nightly backup timer (pg_dump, config tar, ClickHouse backup best-effort, MinIO mirror, rsync off-node to BHS64b). No CI, no Alertmanager.
- Finding (ops): the production
schedulercontainer is crash-looping (851 restarts, exit 0, no log) because the deployed image predates commitff67cee(standalone scheduler entry). Crawling continues because the BHS64b worker embeds the scheduler loop. Fixed by the redeploy at the end of this upgrade.
B. Current database schema (Postgres)
28 tables. Core: facilities (2 912 rows; it_capacity_mw, total_power_mw, planned_power_mw, mw_is_estimate,
geo_precision, status, confidence, completeness, campus_id, merged_into), operators (470), projects (324;
planned_mw, investment_usd, status), campuses (114), metros (214), countries (250), cloud_regions (307),
ixps (12), facility_ixps, facility_tenants, facility_aliases, entity_keys (connector-scoped keys, global PK on
key), provenance (38 533 rows, one row per entity×field×source×url, is_current, method, extractor_version),
events (5 549, fingerprint-deduplicated, significance 0–100), news_items (4 404), documents (7 092) +
document_versions (8 785, content-hash versions with diff_summary), entity_matches (1 730: 1 200 auto_created, 81
auto_merged, 449 pending never reviewed), rankings (snapshots with min_coverage), daily_metrics, sources,
connectors, connector_runs, connector_state, project_timeline, system_alerts, admin_sessions.
ClickHouse holds append-only observations, crawl_log, page_changes, entity_daily, api_requests (best-effort, silently
dropped on failure).
Indexes: btree on the obvious FKs/status, trigram GIN on names/cities, generated tsvector on facilities, partial (lat,lng).
No PostGIS (haversine + bbox pre-filter + geohash). No claims/observation table in Postgres; no entity-version table;
no run_id on provenance/document_versions.
C. Current crawler architecture
Declarative YAML per source (config/connectors/*.yaml) → GenericConnector or registered parser/implementation →
discover → fetch (L1 direct → L2 browser identity → L3 Firecrawl → L4 Scrapfly) → archive (MinIO, zstd, per content hash) → extract → normalize → validate → ingest (reconcile, provenance, events). Robots.txt honoured, per-host token bucket,
conditional GET, content-hash change detection, extraction skipped when hash and extractor_version unchanged, adaptive
next_check with EMA change score, per-document quarantine after 3 × 404/410. BullMQ crawl + maintenance queues,
Redis run locks, per-run/per-connector/per-provider premium budgets, Prometheus metrics, dci doctor.
D. Current connectors
104 YAML configs (87 enabled): 60 operator directories, 20 industry/hyperscaler newsrooms (RSS), 10 cloud-region JSON,
6 government, 3 utilities, 3 datasets (PeeringDB — disabled pending AUP approval —, Wikidata, World Bank), SEC EDGAR
full-text, OSM Overpass. Facility rows by connector: OSM 1 346, Equinix 262, Digital Realty 259, Wikidata 190, STACK 78,
DataBank 75, NTT 72, EdgeConneX 59, CyrusOne 52, Google 47, QTS 47 … 58 named parsers; zero configs use the declarative
extractor (dead code path); zero HTML fixtures; packages/core and packages/connectors have no tests.
E. Current data-quality issues (measured)
Facilities: 2 640 operational / 154 unknown / 76 announced / 37 under construction; only 536 (18 %) have any MW;
797 have no coordinates; 1 377 have facility_type = unknown; only 4 flagged AI; 109 duplicate normalized-name pairs and
469 coordinate pairs within ~100 m (OSM building-vs-campus, e.g. atNorth ICE02 vs buildings M01…M16).
Confirmed extraction errors:
- Decimal MW parsed as thousands: DataBank MSP4 shows 779 MW; source says "0.779MW Critical IT Load". Same bug hits
DFW6 (675), AUS1 (405) and any "0.xxx MW" figure (
parseMwstrips the dot when three decimals follow). - Headlines stored as project names: 182/324 projects are named after an article title.
- Executive appointments as projects: "STACK Appoints Matt VanderZanden as CEO" = 13 000 MW project; "AirTrunk appoints Laura Coad" = 1 400 MW; VIRTUS, 365 Data Centers, Element Critical, Cassava appointments likewise.
- Portfolio/company figures on one project: "Microsoft Georgia data center project" 12 GW / $175 B (company capex); "Google Alabama" 17 GW (Southern Company pipeline); Nvidia "Washington" 20 GW (Starcloud funding story).
- Market-research articles as projects: "Hyperscale Data Center Market to Surpass USD 624.2 Billion" ($80.9 B).
- Wrong operator: "Meta's Canadian AI Data Center" assigned to Google.
- PPA / sustainability / HQ / workforce stories as projects (Ormat–Switch PPA 3 400 MW, AirTrunk HQ 1 400 MW, Meta Workforce Academy 5 000 MW).
- Projects have no coordinates at all (0/324), 158 have a country, 0 link to a facility.
- Broken operator slug
itemfor 中華電信數據通信分公司 (26 facilities). - Ranking coverage gate leaves 3 countries in
countries_known_mw— the MW coverage story must be visible, not hidden.
Structural causes (code): one-row-per-entity with no claim scope; isIdentifyingExternalId folds unrelated facilities on
*_code/*_url keys; loadCandidates truncates at 400 rows without ordering; news country via unsafe countryFromText
("North America" → US); no HQ guard on project location; projects.country_iso2 write-once; campus/building double counting
in every MW aggregate; equal-authority ping-pong after 30 days creates event churn; entity_keys.key global PK;
parserVersion is v1 everywhere so parser fixes never invalidate cached extractions; healthFrom(0,0) = ok.
F. Current frontend strengths
Unified status/confidence/precision visual language (CSS variables shared by badges, charts and MapLibre paint);
per-field provenance popovers + provenance table + source history + sources footer with licences; dense typography with
mono tabular figures; URL-as-state everywhere; ⌘K// command palette with API-side query interpretation; server-clustered
map with precision rings; full SEO (canonical, per-entity OG images, JSON-LD, sharded sitemaps); a11y depth; an automated
UX sweep harness; a substantial /admin console (connectors, documents, matches, events, dev tool).
G. Current frontend weaknesses
No explore/query builder, no compare, no evidence drawer (only popover + bottom table), no export, no watchlist, no coverage page, no AI index, no pulse, no power/connectivity layers, no map density/time modes, no virtualization, long single-scroll detail pages without active-section tracking, hand-rolled charts without brushing, keyboard story stops at the palette, project pages lack a state machine/timeline funnel, MapLibre controls < 40 px on touch.
H. Current API strengths
27 public GET endpoints, envelope {data, meta, sources}, zod-validated params, weak ETags + s-maxage, rate limits,
consistent 404/400 JSON, OpenAPI + Swagger UI, /map zoom-tiered clustering, /search with interpretation, admin API with
constant-time token compare and same-origin proxy with CSRF guard. Weaknesses: no provenance/history/nearby/coverage/export
endpoints; list responses carry no sources; OpenAPI has no response schemas; placeholder-token check is a no-op.
I. Schema changes proposed (migration 0003)
claims— claim-first store: subject (type,id), predicate, value/unit, scope (building/facility/campus/metro/ country/portfolio/company/unknown), source/document/url, published/retrieved, confidence, is_estimate, evidence text + offsets, parser name/version,run_id, status (current/superseded/rejected/review/unscoped), rejection reason.quality_flags— deterministic sanity flags (capacity, investment, scope, location, operator, duplicate, project false-positive) with severity, priority score, status, resolution, linking to entity + claim.facilities.parent_facility_id,facilities.record_scope(building/facility/campus) for containment-aware aggregation;facilities.ai_evidence(confirmed/likely/associated/unknown);facilities.utility_capacity_mw,grid_connection_mw,ultimate_campus_mw(capacity ontology columns; IT/total/planned kept).projects.project_class(NEW_BUILD … EXECUTIVE_APPOINTMENT),projects.evidence_level,projects.merged_into,projects.lat/lnggeocoded from metro/city withgeo_precision = city,projects.ai_evidence,projects.investment_scope.document_versions.run_id,provenance.run_id,provenance.scope;connectors.quarantine(bool) +consecutive_failures,blocked_since.field_authority(per-field source-kind tiers) as code table inpackages/core+ methodology page (no table needed).entity_snapshots(daily JSON snapshot of global totals, rankings, facility status counts, project stage counts) for "as of" views and regression checks.- Indexes:
provenance(is_current),facilities(merged_into),facilities(campus_id),facilities(parent_facility_id),events(significance, detected_at),daily_metrics(metric, dim, day), GIN onnews_items.operator_ids.
J. Components to preserve
Everything in section F/H, the connector SDK and escalation ladder, the content-hash/archive pipeline, entity keys, provenance table, event fingerprinting, the matcher (with fixes), rankings coverage gates, the design tokens, the map architecture, the admin console, deploy scripts, backups, the QA sweep.
K. Implementation order
Phase 1 data quality (parseMw fix, claim/scope layer, project classifier + evidence threshold, capacity/investment sanity engine, external-id allowlist, HQ/country guards, campus containment aggregation, extraction debugger, connector health, fixtures, cleanup of current bad records) → Phase 2 core UX (home 2.0, map 2.0, facility 3.0, operator/market/country/project 2.0, brand, mobile) → Phase 3 intelligence (AI index, pulse, live feed 2.0, compare, explore, coverage, quality dashboards) → Phase 4 graph (connectivity, power, cloud, IXP pages) → Phase 5 historical (time machine, capacity history, announced vs delivered, velocity) → API 2.0 + docs + download → regression checks, security fixes, deploy, final data audit.