import type { AuthorityField, ClaimScope, SourceKind } from "@dci/core"; import { AI_EVIDENCE_LEVELS, AUTHORITY_FIELDS, CAPACITY_PREDICATE_LABEL, CLAIM_SCOPES, FIELD_AUTHORITY, INVESTMENT_PREDICATE_LABEL, PROJECT_CLASSES, SOURCE_KINDS, isSiteScope } from "@dci/core"; import Link from "next/link"; import type { ReactNode } from "react"; import { AiMark, ScopeBadge, StageBadge, TierMark } from "@/components/ui/badge"; import { Section } from "@/components/ui/card"; import { TBody, THead, Table, Td, Th, Tr } from "@/components/ui/table"; import { AI_EVIDENCE_LABEL, AUTHORITY_TIER_LABEL, PROJECT_STAGE_ORDER, SOURCE_KIND_LABEL } from "@/lib/labels"; import { routes } from "@/lib/routes"; /* ------------------------------------------------------------------------------------------ Methodology §9–17: the claim-first architecture (docs/CLAIMS.md), rendered from the same code tables the pipeline uses (@dci/core claims.ts) so the page can never drift from the rules. ------------------------------------------------------------------------------------------ */ const linkCls = "text-ink underline decoration-hair-2 hover:text-accent"; function P({ children }: { children: ReactNode }) { return
{children}
; } function Code({ children }: { children: ReactNode }) { return{children};
}
const CLAIM_COLUMNS: Array<[string, string]> = [
["subject_type, subject_id", "facility / campus / project / operator / market / country the figure is about"],
["predicate", "what is asserted — it_capacity_mw, critical_power_mw, utility_capacity_mw, grid_connection_mw, current_power_mw, planned_power_mw, ultimate_campus_mw, phase_mw, project_investment_usd, campus_investment_usd, company_investment_usd, country_program_usd, multi_year_capex_usd, deal_value_usd, status…"],
["value, value_text, unit", "the figure (MW, USD) or text"],
["scope, scope_reason", "building · facility · campus · metro · country · portfolio · company · unknown, and the rule that decided"],
["evidence_text, evidence_start, evidence_end", "the supporting sentence and its offsets in the parsed text — no sentence, no numerical claim (structured spec tables record the extraction method instead)"],
["source_id, url, document_id, published_at, retrieved_at", "where and when"],
["authority_tier", "A–E for this field (see field-level authority)"],
["parser_name, parser_version, run_id", "who extracted it — enables reprocessing and rollback"],
["status", "current · superseded (same source, same page, new value) · rejected (review / rollback) · review (critical sanity flag) · unscoped (describes a portfolio, a company, a country, a metro, or its scope is unknown)"],
];
const SCOPE_MEANING: RecordSince 2026-09-12 every numerical figure the index extracts is stored as a claim before it can become a displayed value. A claim is one assertion by one document about one subject, with its scope, its semantics, the sentence that supports it, the parser that produced it and the field-level authority of its source. Reconciliation then decides which claim (if any) populates a column. Nothing is deleted when sources disagree — every claim stays visible in the evidence drawer.
| Column | Meaning |
|---|---|
{c}
|
{m} |
The winner behind each displayed value is marked (is_winner); run_id on claims, provenance rows, events and document versions ties every change to a connector run, which can be rolled back as a whole. Daily snapshots keep global totals, per-operator and per-country totals, rankings and stage counts for “as of” views and automated regression checks (facilities −5 %, known MW +20 %, one operator +10 GW in a day, one connector > 500 records a day, > 200 location changes a day).
Only site scopes (building, facility, campus) may populate a facility’s or project’s capacity and investment columns. Portfolio, company, country, metro and unknown scopes are stored and shown as evidence, never summed. Scope is classified deterministically from the sentence (classifyScope): company → portfolio → country → metro → campus → building → facility → the record’s own scope when the sentence has no signal (structured parsers) → unknown. A campus figure is never written to a building record; a building’s figure never stands for its campus.
| Scope | Populates a site’s figures? | Meaning |
|---|---|---|
|
|
{isSiteScope(s) ? yes — site scope : no — evidence only} | {SCOPE_MEANING[s]} {isSiteScope(s) ? "(site scope)" : "(evidence only)"} |
IT load, critical power, utility capacity, grid connection, current power, planned capacity, ultimate build-out and phase capacity are different predicates and different columns. The displayed figure carries capacityScope and capacitySemantics so a page can say what “300 MW” means. Utility and grid figures are supply-side — never IT load and never comparable with IT capacity. When a sentence gives no cue the predicate defaults from the field / record status and the claim is flagged mw_semantics_default.
| Capacity predicate | Label |
|---|---|
{k}
|
{v} |
| Investment predicate | Label |
|---|---|
{k}
|
{v} |
Only project and campus investment may stand for a site; deal values, capex, company investment and country programmes are stored as evidence.
capacitySanity and investmentSanity run on every claim. Blocking flags keep the figure out of the columns (it stays a claim). Non-blocking critical flags assign the figure but open a quality flag prioritised by impact; the figure stays visible and marked until a reviewer resolves or dismisses the flag. Every entity page lists its open flags in the data-quality block.
{c}
{c} · {sev}
Review priority weighs the figure’s magnitude, how many pages display it and the number of sources that disagree; the admin console works the queue from the top.
Before a news article can create a project, classifyProjectEvent labels it. Only physical classes may create or modify a project; associated classes attach an event to an identifiable project; the rest never become projects. This is what keeps executive appointments, market-research articles and capex guidance out of the pipeline.
{g.note}
projectTransition(from, to) accepts forward moves along the funnel; delayed and cancelled are side branches; leaving delayed resumes where the project was. A backward move needs a source that outranks the stored one for the status field, otherwise it is flagged status_backward. Stage dates (permit_filed_on, approved_on, construction_started_on, opened_on) are set from the announcement that moved the stage — earliest evidence wins.
A single keyword never confirms an AI site. Each facility and project carries an evidence level; AI counts and MW on the site cover confirmed + likely only, and figures are published, site-scoped values (containment-aware), never extrapolated.
See the AI infrastructure index.
Every aggregate uses the containment-aware facility view: a campus whose buildings publish figures contributes no MW itself; a building without a figure under a campus with one is “covered”; a campus with buildings is not counted as an extra facility. Each MW aggregate exposes its coverage (facilities with any figure ÷ facilities) and rankings gate on it (≥ 5 facilities, ≥ 35 %). Market concentration (HHI, top-3 / top-5 shares) is computed on facility counts and, separately, on known MW with its coverage — the MW view is labelled partial when coverage is low. Momentum is shown as separate counters (projects announced, MW entering construction, new entrants, openings, cloud additions, grid events), never collapsed into a score.
Browse the rankings and the coverage report.