Docs: pointers from reconciliation, project extraction, rankings and crawl operations to the claim-first layer
4 changed files +8 −0
modified
docs/CRAWL-OPERATIONS.md
+2 −0
@@ -1,5 +1,7 @@ | ||
| 1 | 1 | # Crawl operations — DataCenterIndex.io |
| 2 | 2 | |
| 3 | +> **2026-09-12 additions.** Connector health now has `blocked`, `schema_change`, `no_new_content` states and a `quarantine` flag (`dci quarantine <id> on|off`, auto-set after 5 consecutive failed/blocked runs or a non-full run creating ≥ 500 records — everything runs, nothing is published, `connector_runs.quarantined = true`). The effective extractor version is the YAML `parserVersion` plus the versions of the parsers actually used: bump a parser `version` and `dci reprocess <id> --stale` re-extracts only the documents extracted with an older version. `dci trace <doc>` shows every stage for one document; `dci quality` / `dci snapshot` / `dci gaps` run the data-quality jobs (also scheduled: quality every 6 h, snapshot + regression checks daily 00:45 UTC). 429 responses never escalate and honour `retry-after`; robots.txt failures fail closed to L1/L2; `noarchive` pages are never archived. | |
| 4 | + | |
| 3 | 5 | Runbook for the crawl runtime (`apps/worker`) in production: the crawl worker on **BHS64b** |
| 4 | 6 | (`compose.crawl.yml`, service `worker`, `/healthz` + `/metrics` on `10.68.0.2:8320`), the dedicated |
| 5 | 7 | **scheduler** and the maintenance-only worker on **BHS128** (`compose.data.yml`, services `scheduler` :8321 and |
modified
docs/PROJECT-EXTRACTION.md
+2 −0
@@ -1,5 +1,7 @@ | ||
| 1 | 1 | # Announcement → project extraction (`news_article_v1`, `news_planning_pdf_v1`, `news_edgar_fts_v1`) |
| 2 | 2 | |
| 3 | +> **2026-09-12 — announcement classifier first.** Before any project is created, `classifyProjectEvent` (packages/core/src/claims.ts) labels the announcement (NEW_BUILD, EXPANSION, CONSTRUCTION_START, PERMIT, LAND_ACQUISITION, GRID_CONNECTION = physical; POWER_AGREEMENT, FINANCING, ACQUISITION, PARTNERSHIP, CUSTOMER_AGREEMENT = associated events only; EXECUTIVE_APPOINTMENT, SUSTAINABILITY, PRODUCT_NEWS, GENERAL_COMPANY_NEWS, UNKNOWN = never a project). Creation needs (explicit name or operator) + location + development verb; company domiciles ("Denver-based", "headquartered in") are stripped before locating; the headline MW/$ figures are stored as claims with scope and the supporting sentence — portfolio / company / country figures never populate `planned_mw` / `investment_usd`. Lifecycle changes follow `projectTransition`. Details in [CLAIMS.md](CLAIMS.md); debug any article with `pnpm dci trace <url>`. | |
| 4 | + | |
| 3 | 5 | News, government, utility and filing sources feed the **live change feed** (`news_event`) and the **project pipeline** |
| 4 | 6 | (`project`). This document describes how a press release or a planning report becomes structured records, what is |
| 5 | 7 | extracted, with which method (provenance), and — above all — what is *not* inferred. |
modified
docs/RANKINGS.md
+2 −0
@@ -1,5 +1,7 @@ | ||
| 1 | 1 | # Rankings — methodology |
| 2 | 2 | |
| 3 | +> **2026-09-12 — containment-aware aggregation.** Every aggregate reads `FACILITY_VIEW` (apps/worker/src/rankings.ts): a campus row whose buildings publish MW contributes nothing itself, a building covered by its campus figure counts as covered, a campus with buildings is not an extra facility. Project figures are summed only from site-scoped claims (`projectPlannedMw`). Hidden / merged projects are excluded everywhere. New keys: metros_ai, metros_cloud_regions, metros_projects, operators_ai, operators_projects, operators_project_mw, countries_project_mw, countries_projects, countries_ixps. | |
| 4 | + | |
| 3 | 5 | Rankings are recomputed by `computeRankings()` (`apps/worker/src/rankings.ts`) and served with the methodology |
| 4 | 6 | text below (`rankings.methodology`, also returned in the API `meta`). They are only as good as the coverage of the |
| 5 | 7 | index; every row therefore exposes `coverage`, the share of the entity's facilities that carry any MW figure. |
modified
docs/RECONCILIATION.md
+2 −0
@@ -1,5 +1,7 @@ | ||
| 1 | 1 | # Entity reconciliation |
| 2 | 2 | |
| 3 | +> **2026-09-12 — claim-first layer.** Numeric fields (capacity, investment) now flow through the claim store described in [CLAIMS.md](CLAIMS.md): scope + semantics + evidence sentence + field-level authority, sanity engine, then the merge policy below. External-id folding is allowlist-only (`IDENTIFYING_EXTERNAL_ID_KEYS`), candidates are relevance-ordered (same operator → same name → nearest), and `rule:campus-vs-building` now links the building to its campus (`parent_facility_id`, `record_scope`) instead of opening a duplicate review. `provenance.is_winner` marks the observation behind each displayed value; `run_id` ties every write to a connector run (rollback). | |
| 4 | + | |
| 3 | 5 | How a normalized record coming out of a connector becomes (or updates) a canonical facility, operator, project, |
| 4 | 6 | cloud region or IXP — and how conflicting values are merged. Implementation: `apps/worker/src/ingest/` |
| 5 | 7 | (`match.ts` holds the pure scoring/merge logic, `facilities.ts` wires it to Postgres). |
| 6 | 8 | |