SPB Git forge
38commits 1branches 0releases
338.7 MBsize
maindefault branch
2 h agolast push
HTML 53.9% TypeScript 44.5% JavaScript 0.6% SQL 0.5%
11.1 KB

# Entity reconciliation

2026-09-12 — claim-first layer. Numeric fields (capacity, investment) now flow through the claim store described in CLAIMS.md: scope + semantics + evidence sentence + field-level authority, sanity engine, then the merge policy below. External-id folding is allowlist-only (IDENTIFYING_EXTERNAL_ID_KEYS), candidates are relevance-ordered (same operator → same name → nearest), and rule:campus-vs-building now links the building to its campus (parent_facility_id, record_scope) instead of opening a duplicate review. provenance.is_winner marks the observation behind each displayed value; run_id ties every write to a connector run (rollback).

How a normalized record coming out of a connector becomes (or updates) a canonical facility, operator, project, cloud region or IXP — and how conflicting values are merged. Implementation: apps/worker/src/ingest/ (match.ts holds the pure scoring/merge logic, facilities.ts wires it to Postgres).

# Guarantees

  • Nothing is invented and nothing is lost. Every incoming value is written to provenance even when it does not win the merge; ambiguous matches are kept as new records and queued for review, never discarded.
  • Idempotent. Re-ingesting the same document yields 0 created rows, 0 events and the same provenance rows (refreshed last_observed). Connector keys (peeringdb:fac:12) are mapped in entity_keys.
  • Atomic per batch, tolerant per entity. A batch runs in one transaction; each entity runs in a savepoint, so a bad record is rolled back and counted as rejected without failing the batch. Dry runs roll back everything.

# Concurrency

Connectors run in parallel, so two batches may describe the same operator at the same moment. Operator resolution takes a transaction-scoped advisory lock on the normalized (canonical) name before looking up or inserting, which serializes creation across batches; mergeOperators() folds duplicates that slipped through before the lock existed. Facilities are not locked (a dataset batch can carry thousands of them and Postgres' lock table is finite): two connectors creating the same unseen facility within the same seconds can still produce a pair of records that only a later near-duplicate sweep or an admin merge (mergeFacilities) will fold.

# Facilities

# Resolution order

  1. Connector key — entity_keys.key = record.key → same facility (following merged_into). Always wins.
  2. Shared external ids — any externalIds entry (peeringdb_fac, osm, wikidata, …) already stored on a live facility → same facility.
  3. Candidate scoring — candidates are live facilities that (a) lie in a 40 km bounding box around the incoming coordinates, or (b) share the normalized name (or an alias) in the same country, or (c) belong to the same operator in the same country (so facility codes like DC12 can be compared). Up to 400 candidates are scored.

# Score

Weighted sum, weights redistributed over the signals available on both sides:

Signal Weight Value
Name 0.40 1 for equal normalized names (noise words removed), equal code-preserving names or an alias hit; else max(token Jaccard, 0.85 × containment)
Facility codes (DC12, FR5, LD8) 0.15 1 when both names carry codes and one matches, 0 when they conflict; skipped when either side has no code
Operator 0.20 1 same operator, 0.5 unknown on either side, 0 different
Distance 0.15 1 − 0.4·d/r inside the match radius r = matchRadiusKm(precisionA, precisionB) (exact 250 m … metro 40 km), decaying to 0 at 4r; 0.6 for "same city" without coordinates
Address 0.10 normalized address equality → 1; same house number → ≥ 0.75; different house numbers → ≤ 0.4; else token Jaccard

Rules applied on top of the weighted sum:

  • Operator + code rule: same operator, matching facility code and geographically compatible (inside the radius or no coordinates) → the name signal is raised to 0.95. This is what merges DC12 into Equinix DC12.
  • Code conflict: same operator but conflicting codes (DC12 vs DC13) caps the name signal at 0.45.
  • Different countries cap the score at 0.30.
  • Too far: precise coordinates more than 4× the match radius apart cap the score at 0.50.
  • Different known operators never merge — the score is capped below the pending threshold — unless the normalized name and the normalized address are both exact (an operator change on the same building).

# Decision thresholds

Score Decision What happens
≥ 0.92 auto-merge The incoming record updates the candidate; entity_matches gets an auto_merged row (score + reasons) for audit.
0.60 – 0.92 pending A new facility is created with confidence = unverified, and an entity_matches row with status = pending, the candidate JSON (incl. createdFacilityId), the matched facility id, the score and the reasons. Nothing is lost; the admin decides.
< 0.60 create New facility (auto_created row recorded when a candidate scored ≥ 0.30, for tuning).

Why "create + pending" rather than "hold": the record is visible immediately (flagged unverified), its provenance is preserved, and approving the match later is a pure fold (mergeFacilities). Holding would hide real facilities for days when the review queue is slow.

# Admin review flow

GET /api/admin/matches?status=pending lists the queue. For each row:

  • Approve → mergeFacilities(createdFacilityId, matchedFacilityId): keys, aliases, provenance (deduplicated on entity/field/source/url), tenants, IXPs, projects and events are moved to the surviving facility; the created one gets merged_into; its external ids are folded; derived fields are recomputed; the match row becomes approved.
  • Reject → the match row becomes rejected; the created facility simply stays a separate record and its confidence is recomputed on the next observation (it is no longer forced to unverified).

A facility keeps confidence = unverified for as long as a pending match references it as createdFacilityId.

# Field merge policy

Each incoming field is compared with the observation that currently backs the stored value (from provenance, is_current = true, joined with sources.kind):

  1. Empty stored value → take the incoming one.
  2. Same source re-observing → take it (a source may correct itself).
  3. Measured beats estimate (isEstimate), whatever the source.
  4. Otherwise higher authority wins: operator / government / filing 4 · utility / cloud provider 3.5 · registry 3 · dataset 2.5 · community 2 · secondary 1.5 · news 1 (estimates −1.5). Ties → most recent observation.

Special cases:

  • MW (itCapacityMw, totalPowerMw, plannedPowerMw): a non-primary source (dataset, community, secondary, news) never overwrites a figure from an operator / government / filing / utility / cloud-provider / registry source. mw_is_estimate is true only when every current MW observation is an estimate.
  • Coordinates: never replaced by a less precise point (PRECISION_RANK: exact > parcel > street > approximate > city > metro > unknown). Same precision → newest wins (same source always refreshes its own point).
  • Flags is_ai, is_hyperscale: true sticks. A facility operated by a hyperscaler is hyperscale.
  • Certifications, aliases, external ids: accumulated (set union).
  • Name: same policy as other fields; the previous name becomes an alias.
  • Status / type: unknown never overwrites a known value.

# Derived fields (recomputed after every write)

  • completeness 0–100: geo 15 (exact/parcel/street; 8 for city/metro/approximate), operator 10, address 10, status 10, MW 20 (14 when estimate; planned-only 12/8), type 5, opened 10, website 5, description 5, tenants/IXPs 10.
  • confidence = computeConfidence() with the best source kind as base and the number of distinct primary source kinds as corroborations; verified needs ≥ 2 primary kinds, estimated when every observation is an estimate.
  • source_count = distinct sources with current provenance; last_verified = latest observation by a primary source.
  • metro_id = nearest seeded metro whose radius covers the point (same country), else the metro whose alias matches the city; country_iso2 falls back to the metro's country; geohash (precision 7).

# Operators

Order: connector key → external ids → curated canonical table (canonical-operators.ts: ~175 brands with aliases, e.g. Interxion/Telx/DuPont Fabros → Digital Realty, RagingWire/e-shelter/Gyron/NetMagic → NTT Global Data Centers) → exact normalized name → alias match (case-insensitive) → trigram similarity ≥ 0.92 and the same website domain → create. Kind inference: curated kind, else caller hint (carrier / cloud), else name heuristics. Hyperscalers (AWS, Microsoft, Google, Meta, Oracle, Alibaba Cloud, Tencent Cloud, Apple, IBM, Huawei Cloud) get kind = hyperscaler; is_cloud_provider / is_carrier are set from the table or the role hint and only ever turn on. Incoming names that differ from the canonical name are appended to aliases. News text can only link to curated or already-known operators; it never creates one.

# Projects

Key → external ids → same sourceUrl → (same country, same or unknown operator, normalized-name trigram similarity ≥ 0.85). Timeline rows are deduplicated on (project, date, type, sha256(description)). Status changes emit project_status_changed; planned MW, expected opening, investment and operator changes follow the same event rules as facilities. last_update moves whenever a column or a timeline row changes.

# Cloud regions, IXPs, campuses, tenants

  • Cloud regions are unique on (provider, code); the provider is resolved as an operator with the cloud hint.
  • IXPs: key → external ids → normalized name (or long name) in the same country. facility_ixps links come from the IXP's facilityKeys (resolved through entity_keys) or from a facility page's ixps list.
  • Campuses: key → external ids → normalized name with compatible operator/country.
  • Tenants: carriers / cloudProviders on a facility resolve to operators (carrier / cloud hint, AS12345 parsed into asn) and fill facility_tenants; carriers_count / networks_count are recomputed.

# Events

Tracked facility fields: status, IT/total/planned MW, opened/announced/construction-start dates, operator, owner, facility type, name. Significance: status 90 · operator 85 · owner 70 · MW change ≥ 20 % 80 else 50 · dates 60 · others 20 (first-time values are capped at 50; enrichments of name/type/operator are not events). Fingerprint = sha256(entityType|entityId|eventType|JSON(newValue)|day) — the same change observed by several pages on the same day is one event. facility_discovered (65 when MW ≥ 50 or in pipeline, else 40) and project_announced (75 when ≥ 100 MW planned) are emitted for new records. News items keyed by URL emit one event per article (fingerprint on the URL) when they carry an event type with significance ≥ 40.