spb/datacenterindex
Public
HTML 53.9%
TypeScript 44.5%
JavaScript 0.6%
SQL 0.5%
1# Connector authoring guide23A connector is one source (an operator website, a publisher feed, a government newsroom, a registry API, a dataset) described by **one YAML file** in `config/connectors/<id>.yaml`, optionally backed by code registered from `apps/worker/src/connectors/<group>/index.ts`. The runtime (`apps/worker`) turns it into: discover → fetch → archive → extract → normalize → validate → reconcile → publish. This guide covers the YAML, the two code hooks (parser / implementation), the fixture-based golden tests, live dry-runs, quarantine, parser versioning, health states and the legal rules. Everything below is grounded in the code paths named in brackets.45Per-source notes (what the site publishes, quirks, verification log) live in `docs/connectors/<id>.md` — one file per connector, written when the connector ships.67## 1. The YAML (`connectorConfigSchema`, `packages/connectors/src/config.ts`)89`parseConnectorConfig(text, file)` validates with zod and applies defaults; `loadConnectorConfigs(dir)` loads the whole directory and rejects duplicate ids. `pnpm dci validate` runs the same schema (plus the `superRefine` cross-field rules: `fetch.level ≤ fetch.maxLevel`, seed / classify `minLevel ≤ fetch.maxLevel`, every regex compiles).1011| Field | Type / default | Meaning |12|---|---|---|13| `id` | `^[a-z0-9][a-z0-9-]*$` | Connector id = file name = fixture dir = `connectors.id`. Keys are prefixed with it (`databank:dfw6`). |14| `name`, `domain` | string | Source name and host (`domain` drives the per-host rate limiter, `configureHost`). |15| `kind` | `SOURCE_KINDS` (`operator`, `news`, `government`, `filing`, `utility`, `cloud_provider`, `registry`, …) | Drives default provenance confidence (`high` for operator/government/filing/utility/cloud_provider/registry, else `moderate`) and the third-party rules (`kind: news`, § 8). |16| `mode` | `html \| sitemap \| rss \| pdf \| json \| hybrid \| api \| dataset` (default `hybrid`) | Informational (`Connector.type`). |17| `enabled` | bool, default `true` | Parked connectors keep their YAML with `enabled: false`. |18| `priority` | 1 (best) … 5, default 3 | Source priority used by reconciliation when sources disagree. |19| `license`, `attribution`, `homepage`, `notes`, `coverage{countries,operators,global}` | optional | Recorded on `sources`; `coverage` feeds admin + scheduling. |20| `fetch.level` / `fetch.maxLevel` | 1–4, defaults 1 / 2 | Starting escalation level and ceiling: L1 direct, L2 direct + browser identity (`userAgent: browser`), L3 Firecrawl, L4 Scrapfly rendered. The runtime escalates only on 403/429/JS-shell detection. |21| `fetch.renderJs`, `country`, `waitForSelector`, `accept`, `headers`, `timeoutMs`, `maxBytes` | optional | Passed to the fetchers. |22| `fetch.userAgent` | `bot \| browser`, default `bot` | `browser` sends `BROWSER_UA` at L2. |23| `fetch.rpm`, `fetch.concurrency` | 1–600 (default 30), 1–8 (default 2) | Per-host rate limit (`packages/connectors/src/ratelimit.ts`). |24| `fetch.respectRobots` | default `true` | Never set to `false` for a public website (§ 8). |25| `fetch.maxCreditsPerRun` | default 200 | Premium (Scrapfly/Firecrawl) credit budget per run; `maxLevelForBudget` drops to L2 once spent. |26| `fetch.maxCreditsPerDay` | optional | Per-connector UTC-day budget (Redis counter, all workers) on top of the global `DCI_SCRAPFLY_DAILY_BUDGET` / `DCI_FIRECRAWL_DAILY_BUDGET`. |27| `discovery.sitemap` | `true` (robots.txt sitemaps) or list of URLs | Sitemap / sitemap-index walk. Discovery fetches (`sitemap`, `rss` groups) never use premium fetchers (`DISCOVERY_GROUPS`). |28| `discovery.rss` | list of feed URLs | RSS/Atom items → documents (title/published/summary hints in `doc.meta`). |29| `discovery.seeds` | strings or `{url, group, pageType, minLevel}` | Fixed pages (index pages, single documents). |30| `discovery.follow` | regex list | Internal links followed **from index pages only**. |31| `discovery.include` / `exclude` | regex lists | URL allow / deny filters applied to everything discovered. |32| `discovery.classify` | `[{pattern, group, pageType?, priority?, minLevel?}]` | First matching pattern assigns the schedule group and a page-type hint. |33| `discovery.maxUrlsPerRun` | default 2000 | Cap per run. |34| `discovery.pagination` | `{template, start=1, max=20}` | `?page={n}` / `/page/{n}` appended to index seeds until no new links. |35| `schedule` | `{group: interval}`, default `{seed: weekly}` | Interval per group: `hourly`, `daily`, `weekly`, `biweekly`, `monthly`, `quarterly`, `never` or `30m` / `3h` / `2d` / `1w` (`intervalMs`). The adaptive scheduler halves it for pages that change and stretches it up to ×4 for static ones (`apps/worker/src/scheduling.ts`). |36| `extractors.<name>` | see below | Which pages produce which records. |37| `defaults` | `{operatorName, operatorKind, countryIso2, facilityType, status, isHyperscale, providerName}` | Merged into every normalized entity when the page did not say otherwise (`GenericConnector.normalize`). |38| `implementation` | registered name | Replaces the generic pipeline (§ 2.2). |39| `params` | free object | Passed to the implementation. |40| `parserVersion` | default `v1` | Human-readable version prefix (§ 6). |4142### Extractors (`extractorSchema`)4344```yaml45extractors:46 facility:47 match: ["/data-centers/[a-z0-9-]+/"] # regex on finalUrl or url (case-insensitive); absent = every page48 pageTypes: [facility_page] # additionally restrict to classified page types49 kind: facility # facility | operator | campus | cloud_region | ixp | project | news_event | country50 parser: databank_facility_v1 # code-backed (§ 2.1) — OR declarative `fields:` below51 params: { keepText: true } # passed to Parser.parse(doc, ctx, params)52 address: # declarative example53 kind: facility54 each: "div.location-card" # one record per matching container (optional)55 key: "acme:{code}" # stable key template; {url} = finalUrl; default acme:<urlFingerprint>56 minCertainty: 0.5 # certainty = filled fields / max(3, 0.6 × declared fields)57 fields:58 name: "h1" # a bare string is a CSS selector (text)59 code: { selector: "h1", regex: "\\b([A-Z]{2,4}\\d{1,3})\\b", group: 1 }60 itCapacityMw: { selector: ".stat-power", transform: mw }61 lat: { jsonld: "Place", jsonPath: "$.geo.latitude", transform: float }62 city: { meta: "og:locality" }63```6465`selectorRule` accepts `selector`, `attr`, `regex` / `regexFlags` / `group`, `all`, `jsonPath` (embedded JSON / JSON-LD), `jsonld` (@type to search), `meta` (`<meta property|name>`), `transform` (`mw sqm ha money date status type country int float trim lower url bool`) and `default` — evaluated by `evalRule` in `packages/connectors/src/extract.ts`. Field names are the `NormalizedFacility` / `NormalizedProject` / … attributes (`name`, `address`, `city`, `regionName`, `postalCode`, `countryIso2` or `country`, `lat`/`lng` + `geoPrecision`, `itCapacityMw`, `totalPowerMw`, `plannedPowerMw`, `buildingSqm`, `siteAreaHa`, `pue`, `rackCount`, `tier`, `status`/`statusText`, `facilityType`, `certifications`, `cloudProviders`, `renewableClaim`, `coolingType`, `openedOn`, `description`, `externalIds`, …).6667When a page matches **no** extractor but is classified as an announcement (press release, project, expansion, planning document, …) and is data-center relevant, `GenericConnector.extract` emits a fallback `news_event` (and, for non-`news` kinds only, a `project` when `fallbackProjectGate` sees a MW figure and a pipeline verb in the headline).6869## 2. Code hooks7071### 2.1 Parser (`registerParser`, `Parser` in `packages/connectors/src/types.ts`)7273```ts74export interface Parser {75 name: string; // referenced from YAML: extractors.<x>.parser76 version: string; // part of the effective extractor version (§ 6) — bump when the output changes77 pageTypes?: PageType[];78 parse(doc: RawDocument, ctx: ConnectorContext, params?: Record<string, unknown>): ExtractedRecord[] | Promise<ExtractedRecord[]>;79}80```8182A parser is a pure function of the archived document: **no network, no geocoding, no guessed figures** — it surfaces what the page publishes, with a per-field `methods` map for provenance (`{ itCapacityMw: "regex:it_mw_v1", lat: "json-ld:GeoCoordinates" }`). Parsers live in `apps/worker/src/connectors/<group>/<file>.ts` and are registered by the group's `index.ts`:8384- `operators1/` (Equinix, Digital Realty, NTT, CyrusOne, QTS, Vantage, STACK, CoreSite, Switch): each file exports a `Parser` object; `operators1/index.ts` `register()` calls `registerParser` for each. Helpers in `operators1/shared.ts` (`parseNaAddress`, `parseEuAddress`, `parseCityLine`, `mwAfter`, `mwWithContext`, `areaSqm`, `certificationsFrom`, `record(kind, key, url, data, methods, certainty)` which drops empty fields and sets `pageType: facility_page`).85- `operators2/` (DataBank, EdgeConneX, TierPoint, NEXTDC, AirTrunk, atNorth, CloudHQ, …, grouped in `americas.ts` / `emea.ts` / `apac.ts`): parsers are declared with **`facilityParser(name, (page, doc) => ExtractedRecord[], version = OPERATORS2_PARSER_VERSION)`** from `operators2/shared.ts`, which registers on import and hands you a `Page` (`$`, `html`, visible `text` with nav/header/footer removed, `title`, `h1`, `desc`, `url`). Build records with the **`Rec`** builder: `new Rec().set(field, value, method)` ignores null/empty/NaN and records the method; `.done(key, url, certainty = 0.85, kind = "facility")`. Other helpers: `itMw`, `capacityMw`, `mw` (handles `1.125 MW` decimals, `MWs`, and "at least" figures `10+MW` / `20MW+` — pass the raw string, `parseMw` reads the stated figure), `sqm`, `ha`, `pue`, `tier`, `certifications`, `usAddress`, `ldWithAddress` / `applyLdAddress` / `applyLdGeo` / `breadcrumbLeaf` (JSON-LD), `countryOf` / `countryFromCity` (deterministic metro → ISO table, not geocoding), `statusOf`, `key(op, ...parts)`. `operators2/index.ts` `register()` only checks that every name in `AMERICAS_PARSERS` / `EMEA_PARSERS` / `APAC_PARSERS` is registered. A stat regex that reads a figure before its label must accept the `+` suffix: `/(\d[\d.,]*\+?\s*MW\+?)\s*\n?\s*IT Capacity/`.86- `news/` (`news_article_v1`, `news_planning_pdf_v1`, `news_edgar_fts_v1` + the `news_edgar_fts` implementation): article → `news_event` (+ `project` when `qualifiesAsProject`) via `articleContent(doc)` → `extractAnnouncement(title, text)` (`extract-project.ts`, unit-tested in `extract-project.test.ts`). `articleContent` uses `mainText` (`packages/connectors/src/extract.ts`): the largest `<article>` / `<main>` block after removing nav / header / footer / aside / related / recommended / trending / popular / sidebar / comments blocks, cut at the first trailing teaser heading ("More in …", "Related articles", "Tags"). Operators are looked up in the title plus the first `OPERATOR_SCOPE_CHARS` (4 000) characters only; money figures in a company-background sentence (`MONEY_PORTFOLIO_RE`: "investment volume", "revenue", "assets under management"…) are never the project's investment. Params: `keepText` (public-sector sources only), `minProjectMw` (5), `minInvestmentUsd` (50 000 000), `minAcres` (100), `lenient`.87- `cloud/` (AWS, Azure, GCP, Oracle, IBM, Alibaba, Tencent, Meta regions) and `datasets/` (PeeringDB, Wikidata, World Bank, OSM Overpass): mostly `implementation`s.8889`apps/worker/src/connectors/index.ts` `registerAllConnectors()` imports every `<group>/index.ts` found on disk and calls its `register()`; a group that throws is reported and the others still load. Make `register()` idempotent (guard with `listParsers()` or a module flag) — the worker, the CLI, `try-connector` and the fixture tests may all call it. Add a new group by creating `apps/worker/src/connectors/<group>/index.ts` with `export function register(): void`.9091### 2.2 Implementation (`registerImplementation(name, (cfg) => Connector)`)9293For APIs and datasets where discovery/fetch/extract do not fit the generic HTML pipeline, register a factory returning a full `Connector` (`discover`, `fetch`, `extract`, `normalize`, `validate`; see `datasets/peeringdb.ts`). Reference it with `implementation: <name>` in the YAML; `buildConnector` (`apps/worker/src/configs.ts`) prefers it over `GenericConnector`. Keep `validate` delegating to `validateEntities` (`packages/connectors/src/generic.ts`) unless you have a reason not to.9495### 2.3 Validation (`validateEntities`)9697Hard errors reject the record (implausible MW `≤ 0` or `> 10 000`; a single facility `> 1 000 MW` without a campus designation; PUE outside 1–3; building `> 500 000 m²` without campus; integer coordinates flagged `exact`; title-like facility names such as "Data Centers in Virginia | Acme"; opening year `< 1950`; dates `> now + 12 y`; bad ISO-2). Warnings are recorded (no country and no geo, IT > total power, …). A connector must produce **zero errors** on its fixtures and on a live dry-run before it ships.9899## 3. Fixtures and golden tests100101Real archived bodies replay through the exact production path — `parseConnectorConfig` → registered parsers → `connector.extract` → `connector.normalize` → `validateEntities` — with **no network**. Test: `apps/worker/src/connectors/fixtures.test.ts`. Data: `apps/worker/fixtures/<connectorId>/`.102103### 3.1 Layout104105```106apps/worker/fixtures/107 databank/108 dfw6-love-field.html # raw archived body (html | json | xml | pdf | txt | md), < 400 KB109 dfw6-love-field.expected.json # sibling expectations110 datacenterdynamics/111 sb-gruppe-bucharest.html112 sb-gruppe-bucharest.expected.json113```114115The directory name **is** the connector id (the test loads `config/connectors/<dir>.yaml`). Adding a fixture = adding the two files; no code changes.116117### 3.2 Pulling a body from the production archive (data node BHS128)118119```bash120# 1. pick a document (Postgres in the dci-postgres-1 container)121ssh BHS128 'docker exec -i dci-postgres-1 psql -U dci -d dci -Atc "122 select id, url, content_type, extract_count, last_fetched from documents123 where connector_id = '\''databank'\'' and storage_key is not null and extract_count > 0 and status_code = 200124 and url ~ '\''/data-centers/[a-z0-9-]+/'\''125 order by last_fetched desc limit 10"'126127# 2. raw archived body through the admin API (x-dci-admin-token from deploy/.env.data)128ssh BHS128 'TOK=$(sed -n "s/^DCI_ADMIN_TOKEN=//p" /srv/dci/app/deploy/.env.data);129 curl -s -H "x-dci-admin-token: $TOK" http://10.67.0.60:8300/api/admin/documents/<doc_id>/raw' \130 > apps/worker/fixtures/<connector>/<name>.html131# fallback: cd /srv/dci/app && docker compose -f deploy/compose.data.yml --env-file deploy/.env.data run --rm cli inspect <doc_id> --body --json132```133134Rules for the body: keep it byte-for-byte (no reformatting, no CRLF normalisation); keep it under 400 KB — if a page is bigger, strip `<script>` blocks *only when the parser does not read them* (Equinix and Digital Realty parsers read JS constants / `__NEXT_DATA__`; NTT is selector-only, so `ntt/chicago-ch3.html` is script-stripped); never remove the region the parser needs (address block, MW stat, JSON-LD). Redact personal data of individuals (journalist bylines, staff e-mails); functional mailboxes (`sales@`, `press@`) and executive quotes in press releases stay. Check for e-mails with `grep -oiE "[a-z0-9._%+-]+@[a-z0-9.-]+\.[a-z]{2,}" <file>`.135136### 3.3 `expected.json`137138```jsonc139{140 "url": "https://www.databank.com/data-centers/dallas/love-field/", // used as url + finalUrl of the RawDocument141 "documentId": "doc_352f7e263b5afaed", // documents.id on the data node142 "fetchedAt": "2026-09-11T08:02:55.127Z", // documents.last_fetched143 "contentType": "text/html; charset=UTF-8", // documents.content_type (a /pdf/ type switches on PDF parsing)144 "note": "why this page: '0.675MW Critical IT Load' — the decimal must survive",145 "counts": { "facility": 1 }, // asserted per entityType; omit a type to leave it unasserted146 "facilities": [ // also: projects, newsEvents, cloudRegions, operators147 {148 "key": "databank:dfw6", // entities are matched by key149 "name": "Dallas Love Field Data Center (DFW6)",150 "operatorName": "DataBank", "city": "Dallas", "regionName": "TX", "countryIso2": "US",151 "itCapacityMw": 0.675, "totalPowerMw": null, // null = absent or null; numbers are compared exactly152 "mentions.countriesIso2": ["US"], // dotted paths reach nested fields153 "_todo": { // CORRECT values the parser does not produce yet → it.fails154 "totalPowerMw": 36,155 "_reason": "the figure precedes its label ('36 MW total power'); the stat regex only reads label-first"156 }157 }158 ],159 "thirdParty": true, // kind: news — summaries ≤ THIRD_PARTY_SUMMARY_MAX, no article text in records160 "announcement": { // news / government: extractAnnouncement(title, text) facts161 "class": "NEW_BUILD", "mayCreateProject": true, "headlineMw": 12, "status": "announced",162 "operator": "Applied Digital", "country": "RO", "city": "Bucharest", "money": { "amount": 7500000000, "currency": "USD" },163 "_todo": { "investmentUsd": null, "_reason": "…" }164 }165}166```167168Announcement facts available: `title`, `published`, `class`, `mayCreateProject`, `physical`, `headlineMw`, `mwAll`, `status`, `operator`, `operators`, `country`, `city`, `region`, `money`, `investmentUsd`, `acreage`, `phaseCount`, `expectedOpening`, `projectName`, `explicitName`, `relevance`, `leadRelevance`, `titleSignal`, `nonProjectTitle`, `pageType`, `eventType`.169170Every fixture always asserts **zero validation errors**. Expected values must be what the parser produces today **and** what the source page says: verify each figure against the body before writing it. When the parser is wrong, put the *correct* value under `_todo` with a `_reason` — the test runs as `it.fails` and starts failing the day the parser is fixed, which is the signal to promote the field. Never assert a value you believe is wrong.171172### 3.4 Running173174```bash175pnpm --filter @dci/worker test # everything (fixtures included)176pnpm --filter @dci/worker exec vitest run src/connectors/fixtures.test.ts -t databank # one connector177```178179Golden coverage on 2026-09-12: 21 fixtures over 18 connectors (equinix, digitalrealty, stack, databank ×2, ntt, edgeconnex ×2, cyrusone, qts ×2, vantage, coresite, airtrunk, nextdc ×2, atnorth, cloudhq, datacenterfrontier, datacenterdynamics, loudoun-county), no open `_todo` (the nine gaps recorded on 2026-09-12 — EdgeConneX city from the h1, NEXTDC `20MW+`, QTS figure-first capacity and template-copied map pin, DCD teaser operator and portfolio money — were fixed the same day and promoted). No planning PDF fixture yet: the archive holds no document with a PDF content type.180181## 4. Live dry-run182183```bash184pnpm tsx scripts/try-connector.ts config/connectors/<id>.yaml [--limit 5] [--url <single url>] [--verbose] [--json out.json]185pnpm dci run <id> --dry-run --limit 5 [--url u] # same through the worker CLI (needs the DB env)186pnpm dci trace <doc-id-or-url> [--live] # one document end to end187```188189`try-connector` registers every `apps/worker/src/connectors/<group>/index.ts`, builds the connector from the YAML (`implementation` or `GenericConnector`), runs `discover` (or the single `--url`), fetches with the real robots / rate-limit / escalation policy through `createTestContext` (in-memory state, nothing persisted, `dryRun: true`) and prints records, entities and the `ValidationReport`. Nothing ships without a live dry-run yielding valid entities and **0 rejected**; record the result in `docs/connectors/<id>.md` ("Verification (live, <date>)").190191## 5. Quarantine192193Two different things share the word:194195- **Connector quarantine** (`connectors.quarantine`, `pnpm dci quarantine <id> [on|off]`): the run does everything — discover, fetch, archive, extract, validate — but the ingest transaction is **rolled back**, so nothing is published; the run is stored with `quarantined = true` for review. `runConnector` (`apps/worker/src/pipeline.ts`) reads the flag before each run. It is switched on **automatically** when a connector fails or is blocked **5 runs in a row** (`consecutive_failures ≥ 5`) or when a **non-`full` run creates ≥ 500 records** (`suspiciousSpike`): both are logged as `… — quarantined pending review`. New connectors are worth running quarantined for the first full pass.196- **Document quarantine** (`documents.quarantined`): `nextCheckAfter` (`apps/worker/src/scheduling.ts`) backs a 404/410 off ×4 and after **3 consecutive** gone responses (`MAX_CONSECUTIVE_GONE`) sets `nextCheck = null` and `quarantined = true` — the URL is never scheduled again. Other errors back off ×2^(n−1) (cap ×8); premium retries after two failures are limited to every 4th attempt (`premiumAllowedAfterErrors`). Weekly check in `docs/CRAWL-OPERATIONS.md`: `select count(*) from documents where quarantined` should grow slowly.197198## 6. Parser versioning and re-extraction199200The **effective extractor version** stored on every document (`documents.extractor_version`) is computed by `effectiveExtractorVersion(cfg, connector)` in `apps/worker/src/configs.ts`:201202```203<cfg.parserVersion>+<sha256("news_article_v1@news_v3|databank_facility_v1@v2|impl:peeringdb@v1")[0:10]>204```205206i.e. the YAML `parserVersion` plus a hash of every referenced `Parser.name@Parser.version` (and `impl:<name>@<connector.parserVersion>`). Consequences:207208- **Bumping `version` on the `Parser` object is what invalidates cached extractions.** `shouldSkipExtraction` (`scheduling.ts`) skips a document only when its content fingerprint is unchanged **and** the stored extractor version equals the current one (and the run is not `--force`). Change a regex without bumping the version and unchanged pages keep their old records until their content changes.209- Re-run the archive without network: `pnpm dci reprocess <id> [--limit n] [--group g] --stale` re-extracts from stored bodies (MinIO) only the documents whose `extractor_version` differs from the current one; without `--stale` everything is re-extracted. `pnpm dci run <id> --task reprocess` is the same thing.210- Declarative connectors have no parser: bump `parserVersion` in the YAML.211212## 7. Health states (`connectorHealthFrom`, `apps/worker/src/scheduling.ts`)213214`connectors.health` after each run (admin health center; `docs/CRAWL-OPERATIONS.md` § "Reading a bad run"):215216| State | Rule |217|---|---|218| `ok` | run status not failed, fetch failure rate < 20 % (`healthFrom`) |219| `degraded` | failure rate 20–50 % |220| `failing` | run `status = failed`, or failure rate ≥ 50 % |221| `blocked` | ≥ 50 % of fetched pages came back blocked (403/429/CAPTCHA/JS shell) — fix the fetch level / user agent, never bypass |222| `schema_change` | ≥ 5 pages fetched fine and passed to extractors but **0 entities** on a connector that has documents — the site's HTML changed: re-check selectors against a fresh fixture |223| `no_new_content` | discovery yielded 0 URLs and 0 fetches although a previous run discovered some — sitemap/RSS moved or emptied |224| `never_run` | no run yet (also the state of `enabled: false` connectors) |225226## 8. Legal and content rules (non-negotiable)227228- `fetch.respectRobots: true` — `createTestContext` and the runtime check `robots.txt` before every fetch and honour `crawl-delay`; a disallowed URL yields `error.code = robots_disallow`, not a fetch.229- No authentication, paywall or CAPTCHA bypass; escalation (L3/L4) is for JS shells and bot walls on public pages only, within the credit budgets.230- Conditional GET (ETag / Last-Modified) and content hashing: unchanged pages are never re-extracted, premium credits are never spent re-fetching unchanged content.231- **Third-party publishers (`kind: news`)** keep only title, date, link, extracted facts and a summary of at most **`THIRD_PARTY_SUMMARY_MAX = 400` characters** (`trimSummary`, `packages/connectors/src/generic.ts`); the article text is never stored (`news_article_v1` `keepText` defaults to false; the fixture test enforces both when `thirdParty: true`). Only public-sector sources (`kind: government`, `utility`, `filing`) may set `keepText: true`.232- `license` / `attribution` are mandatory in the YAML and are shown with every record (`sources.license`, `sources.attribution`).233- Never invent data: no geocoding, no guessed capacities, no default coordinates. City-level coordinates are `geoPrecision: city`, operator-published markers `exact` / `parcel`. An estimate is flagged `isEstimate`.234235## 9. Checklist for a new connector2362371. `config/connectors/<id>.yaml` — id, name, domain, kind, priority, license/attribution, fetch, discovery (sitemap first, then include/exclude/classify), schedule, extractors, defaults. `pnpm dci validate`.2382. Declarative `fields` when the page is regular; otherwise a parser in `apps/worker/src/connectors/<group>/<file>.ts` registered in `<group>/index.ts` `register()` (`facilityParser` + `Rec` for facilities), or an `implementation` for APIs / datasets.2393. Live dry-run: `pnpm tsx scripts/try-connector.ts config/connectors/<id>.yaml --limit 5` → valid entities, 0 rejected, 0 unexpected premium credits.2404. Fixture: archive a real body under `apps/worker/fixtures/<id>/`, write `expected.json` from what the parser produces **after** checking each value against the page, `pnpm --filter @dci/worker test`.2415. `docs/connectors/<id>.md` — what it collects, discovery, schedule & fetch, quirks, verification log.2426. First production run quarantined (`pnpm dci quarantine <id> on`, run, review `connector_runs`, then `off`).2437. Any later change to a parser's output → bump `Parser.version` → `pnpm dci reprocess <id> --stale`.244