SPB Git forge
38commits 1branches 0releases
338.7 MBsize
maindefault branch
2 h agolast push
HTML 53.9% TypeScript 44.5% JavaScript 0.6% SQL 0.5%
11.4 KB

# DataCenterIndex — full repository & data audit (2026-09-11)

Internal audit written before the "infrastructure intelligence graph" upgrade. Sections A–K follow the brief. Figures come from the production database on BHS128 (deploy/bin/psql.sh) on 2026-09-11 evening, one day after launch.

# A. Current architecture

  • Monorepo pnpm + TypeScript strict/ESM: packages/core (types, ids, normalize, geo, confidence, ssrf, api contract), packages/db (Drizzle schema, plain SQL migrations, ClickHouse DDL), packages/connectors (connector SDK), apps/worker (runtime, ingest, scheduler, rankings, CLI), apps/api (Fastify 5, /api/v1 public + /api/admin token), apps/web (Next 16 App Router, Tailwind v4, MapLibre + OpenFreeMap, hand-rolled SVG charts).
  • Rendering: every public page is ISR (revalidate 30–300 s) on top of a typed API client that never throws; map page is a static shell with client fetches to /api/v1/map (server-side clustering, ≤ 5 000 points).
  • Deployment: two OVH servers joined by WireGuard. BHS128 = data/public (Caddy edge :8300 → web :8310 / api :8311, Postgres 17, ClickHouse, Redis, MinIO, Prometheus, Grafana). BHS64b = crawl worker (concurrency 6). Public route via the MacLustr Tunnel gateway BHS64. Nightly backup timer (pg_dump, config tar, ClickHouse backup best-effort, MinIO mirror, rsync off-node to BHS64b). No CI, no Alertmanager.
  • Finding (ops): the production scheduler container is crash-looping (851 restarts, exit 0, no log) because the deployed image predates commit ff67cee (standalone scheduler entry). Crawling continues because the BHS64b worker embeds the scheduler loop. Fixed by the redeploy at the end of this upgrade.

# B. Current database schema (Postgres)

28 tables. Core: facilities (2 912 rows; it_capacity_mw, total_power_mw, planned_power_mw, mw_is_estimate, geo_precision, status, confidence, completeness, campus_id, merged_into), operators (470), projects (324; planned_mw, investment_usd, status), campuses (114), metros (214), countries (250), cloud_regions (307), ixps (12), facility_ixps, facility_tenants, facility_aliases, entity_keys (connector-scoped keys, global PK on key), provenance (38 533 rows, one row per entity×field×source×url, is_current, method, extractor_version), events (5 549, fingerprint-deduplicated, significance 0–100), news_items (4 404), documents (7 092) + document_versions (8 785, content-hash versions with diff_summary), entity_matches (1 730: 1 200 auto_created, 81 auto_merged, 449 pending never reviewed), rankings (snapshots with min_coverage), daily_metrics, sources, connectors, connector_runs, connector_state, project_timeline, system_alerts, admin_sessions. ClickHouse holds append-only observations, crawl_log, page_changes, entity_daily, api_requests (best-effort, silently dropped on failure).

Indexes: btree on the obvious FKs/status, trigram GIN on names/cities, generated tsvector on facilities, partial (lat,lng). No PostGIS (haversine + bbox pre-filter + geohash). No claims/observation table in Postgres; no entity-version table; no run_id on provenance/document_versions.

# C. Current crawler architecture

Declarative YAML per source (config/connectors/*.yaml) → GenericConnector or registered parser/implementation → discover → fetch (L1 direct → L2 browser identity → L3 Firecrawl → L4 Scrapfly) → archive (MinIO, zstd, per content hash) → extract → normalize → validate → ingest (reconcile, provenance, events). Robots.txt honoured, per-host token bucket, conditional GET, content-hash change detection, extraction skipped when hash and extractor_version unchanged, adaptive next_check with EMA change score, per-document quarantine after 3 × 404/410. BullMQ crawl + maintenance queues, Redis run locks, per-run/per-connector/per-provider premium budgets, Prometheus metrics, dci doctor.

# D. Current connectors

104 YAML configs (87 enabled): 60 operator directories, 20 industry/hyperscaler newsrooms (RSS), 10 cloud-region JSON, 6 government, 3 utilities, 3 datasets (PeeringDB — disabled pending AUP approval —, Wikidata, World Bank), SEC EDGAR full-text, OSM Overpass. Facility rows by connector: OSM 1 346, Equinix 262, Digital Realty 259, Wikidata 190, STACK 78, DataBank 75, NTT 72, EdgeConneX 59, CyrusOne 52, Google 47, QTS 47 … 58 named parsers; zero configs use the declarative extractor (dead code path); zero HTML fixtures; packages/core and packages/connectors have no tests.

# E. Current data-quality issues (measured)

Facilities: 2 640 operational / 154 unknown / 76 announced / 37 under construction; only 536 (18 %) have any MW; 797 have no coordinates; 1 377 have facility_type = unknown; only 4 flagged AI; 109 duplicate normalized-name pairs and 469 coordinate pairs within ~100 m (OSM building-vs-campus, e.g. atNorth ICE02 vs buildings M01…M16).

Confirmed extraction errors:

  • Decimal MW parsed as thousands: DataBank MSP4 shows 779 MW; source says "0.779MW Critical IT Load". Same bug hits DFW6 (675), AUS1 (405) and any "0.xxx MW" figure (parseMw strips the dot when three decimals follow).
  • Headlines stored as project names: 182/324 projects are named after an article title.
  • Executive appointments as projects: "STACK Appoints Matt VanderZanden as CEO" = 13 000 MW project; "AirTrunk appoints Laura Coad" = 1 400 MW; VIRTUS, 365 Data Centers, Element Critical, Cassava appointments likewise.
  • Portfolio/company figures on one project: "Microsoft Georgia data center project" 12 GW / $175 B (company capex); "Google Alabama" 17 GW (Southern Company pipeline); Nvidia "Washington" 20 GW (Starcloud funding story).
  • Market-research articles as projects: "Hyperscale Data Center Market to Surpass USD 624.2 Billion" ($80.9 B).
  • Wrong operator: "Meta's Canadian AI Data Center" assigned to Google.
  • PPA / sustainability / HQ / workforce stories as projects (Ormat–Switch PPA 3 400 MW, AirTrunk HQ 1 400 MW, Meta Workforce Academy 5 000 MW).
  • Projects have no coordinates at all (0/324), 158 have a country, 0 link to a facility.
  • Broken operator slug item for 中華電信數據通信分公司 (26 facilities).
  • Ranking coverage gate leaves 3 countries in countries_known_mw — the MW coverage story must be visible, not hidden.

Structural causes (code): one-row-per-entity with no claim scope; isIdentifyingExternalId folds unrelated facilities on *_code/*_url keys; loadCandidates truncates at 400 rows without ordering; news country via unsafe countryFromText ("North America" → US); no HQ guard on project location; projects.country_iso2 write-once; campus/building double counting in every MW aggregate; equal-authority ping-pong after 30 days creates event churn; entity_keys.key global PK; parserVersion is v1 everywhere so parser fixes never invalidate cached extractions; healthFrom(0,0) = ok.

# F. Current frontend strengths

Unified status/confidence/precision visual language (CSS variables shared by badges, charts and MapLibre paint); per-field provenance popovers + provenance table + source history + sources footer with licences; dense typography with mono tabular figures; URL-as-state everywhere; ⌘K// command palette with API-side query interpretation; server-clustered map with precision rings; full SEO (canonical, per-entity OG images, JSON-LD, sharded sitemaps); a11y depth; an automated UX sweep harness; a substantial /admin console (connectors, documents, matches, events, dev tool).

# G. Current frontend weaknesses

No explore/query builder, no compare, no evidence drawer (only popover + bottom table), no export, no watchlist, no coverage page, no AI index, no pulse, no power/connectivity layers, no map density/time modes, no virtualization, long single-scroll detail pages without active-section tracking, hand-rolled charts without brushing, keyboard story stops at the palette, project pages lack a state machine/timeline funnel, MapLibre controls < 40 px on touch.

# H. Current API strengths

27 public GET endpoints, envelope {data, meta, sources}, zod-validated params, weak ETags + s-maxage, rate limits, consistent 404/400 JSON, OpenAPI + Swagger UI, /map zoom-tiered clustering, /search with interpretation, admin API with constant-time token compare and same-origin proxy with CSRF guard. Weaknesses: no provenance/history/nearby/coverage/export endpoints; list responses carry no sources; OpenAPI has no response schemas; placeholder-token check is a no-op.

# I. Schema changes proposed (migration 0003)

  1. claims — claim-first store: subject (type,id), predicate, value/unit, scope (building/facility/campus/metro/ country/portfolio/company/unknown), source/document/url, published/retrieved, confidence, is_estimate, evidence text + offsets, parser name/version, run_id, status (current/superseded/rejected/review/unscoped), rejection reason.
  2. quality_flags — deterministic sanity flags (capacity, investment, scope, location, operator, duplicate, project false-positive) with severity, priority score, status, resolution, linking to entity + claim.
  3. facilities.parent_facility_id, facilities.record_scope (building/facility/campus) for containment-aware aggregation; facilities.ai_evidence (confirmed/likely/associated/unknown); facilities.utility_capacity_mw, grid_connection_mw, ultimate_campus_mw (capacity ontology columns; IT/total/planned kept).
  4. projects.project_class (NEW_BUILD … EXECUTIVE_APPOINTMENT), projects.evidence_level, projects.merged_into, projects.lat/lng geocoded from metro/city with geo_precision = city, projects.ai_evidence, projects.investment_scope.
  5. document_versions.run_id, provenance.run_id, provenance.scope; connectors.quarantine (bool) + consecutive_failures, blocked_since.
  6. field_authority (per-field source-kind tiers) as code table in packages/core + methodology page (no table needed).
  7. entity_snapshots (daily JSON snapshot of global totals, rankings, facility status counts, project stage counts) for "as of" views and regression checks.
  8. Indexes: provenance(is_current), facilities(merged_into), facilities(campus_id), facilities(parent_facility_id), events(significance, detected_at), daily_metrics(metric, dim, day), GIN on news_items.operator_ids.

# J. Components to preserve

Everything in section F/H, the connector SDK and escalation ladder, the content-hash/archive pipeline, entity keys, provenance table, event fingerprinting, the matcher (with fixes), rankings coverage gates, the design tokens, the map architecture, the admin console, deploy scripts, backups, the QA sweep.

# K. Implementation order

Phase 1 data quality (parseMw fix, claim/scope layer, project classifier + evidence threshold, capacity/investment sanity engine, external-id allowlist, HQ/country guards, campus containment aggregation, extraction debugger, connector health, fixtures, cleanup of current bad records) → Phase 2 core UX (home 2.0, map 2.0, facility 3.0, operator/market/country/project 2.0, brand, mobile) → Phase 3 intelligence (AI index, pulse, live feed 2.0, compare, explore, coverage, quality dashboards) → Phase 4 graph (connectivity, power, cloud, IXP pages) → Phase 5 historical (time machine, capacity history, announced vs delivered, velocity) → API 2.0 + docs + download → regression checks, security fixes, deploy, final data audit.