# Connector authoring guide A connector is one source (an operator website, a publisher feed, a government newsroom, a registry API, a dataset) described by **one YAML file** in `config/connectors/.yaml`, optionally backed by code registered from `apps/worker/src/connectors//index.ts`. The runtime (`apps/worker`) turns it into: discover → fetch → archive → extract → normalize → validate → reconcile → publish. This guide covers the YAML, the two code hooks (parser / implementation), the fixture-based golden tests, live dry-runs, quarantine, parser versioning, health states and the legal rules. Everything below is grounded in the code paths named in brackets. Per-source notes (what the site publishes, quirks, verification log) live in `docs/connectors/.md` — one file per connector, written when the connector ships. ## 1. The YAML (`connectorConfigSchema`, `packages/connectors/src/config.ts`) `parseConnectorConfig(text, file)` validates with zod and applies defaults; `loadConnectorConfigs(dir)` loads the whole directory and rejects duplicate ids. `pnpm dci validate` runs the same schema (plus the `superRefine` cross-field rules: `fetch.level ≤ fetch.maxLevel`, seed / classify `minLevel ≤ fetch.maxLevel`, every regex compiles). | Field | Type / default | Meaning | |---|---|---| | `id` | `^[a-z0-9][a-z0-9-]*$` | Connector id = file name = fixture dir = `connectors.id`. Keys are prefixed with it (`databank:dfw6`). | | `name`, `domain` | string | Source name and host (`domain` drives the per-host rate limiter, `configureHost`). | | `kind` | `SOURCE_KINDS` (`operator`, `news`, `government`, `filing`, `utility`, `cloud_provider`, `registry`, …) | Drives default provenance confidence (`high` for operator/government/filing/utility/cloud_provider/registry, else `moderate`) and the third-party rules (`kind: news`, § 8). | | `mode` | `html \| sitemap \| rss \| pdf \| json \| hybrid \| api \| dataset` (default `hybrid`) | Informational (`Connector.type`). | | `enabled` | bool, default `true` | Parked connectors keep their YAML with `enabled: false`. | | `priority` | 1 (best) … 5, default 3 | Source priority used by reconciliation when sources disagree. | | `license`, `attribution`, `homepage`, `notes`, `coverage{countries,operators,global}` | optional | Recorded on `sources`; `coverage` feeds admin + scheduling. | | `fetch.level` / `fetch.maxLevel` | 1–4, defaults 1 / 2 | Starting escalation level and ceiling: L1 direct, L2 direct + browser identity (`userAgent: browser`), L3 Firecrawl, L4 Scrapfly rendered. The runtime escalates only on 403/429/JS-shell detection. | | `fetch.renderJs`, `country`, `waitForSelector`, `accept`, `headers`, `timeoutMs`, `maxBytes` | optional | Passed to the fetchers. | | `fetch.userAgent` | `bot \| browser`, default `bot` | `browser` sends `BROWSER_UA` at L2. | | `fetch.rpm`, `fetch.concurrency` | 1–600 (default 30), 1–8 (default 2) | Per-host rate limit (`packages/connectors/src/ratelimit.ts`). | | `fetch.respectRobots` | default `true` | Never set to `false` for a public website (§ 8). | | `fetch.maxCreditsPerRun` | default 200 | Premium (Scrapfly/Firecrawl) credit budget per run; `maxLevelForBudget` drops to L2 once spent. | | `fetch.maxCreditsPerDay` | optional | Per-connector UTC-day budget (Redis counter, all workers) on top of the global `DCI_SCRAPFLY_DAILY_BUDGET` / `DCI_FIRECRAWL_DAILY_BUDGET`. | | `discovery.sitemap` | `true` (robots.txt sitemaps) or list of URLs | Sitemap / sitemap-index walk. Discovery fetches (`sitemap`, `rss` groups) never use premium fetchers (`DISCOVERY_GROUPS`). | | `discovery.rss` | list of feed URLs | RSS/Atom items → documents (title/published/summary hints in `doc.meta`). | | `discovery.seeds` | strings or `{url, group, pageType, minLevel}` | Fixed pages (index pages, single documents). | | `discovery.follow` | regex list | Internal links followed **from index pages only**. | | `discovery.include` / `exclude` | regex lists | URL allow / deny filters applied to everything discovered. | | `discovery.classify` | `[{pattern, group, pageType?, priority?, minLevel?}]` | First matching pattern assigns the schedule group and a page-type hint. | | `discovery.maxUrlsPerRun` | default 2000 | Cap per run. | | `discovery.pagination` | `{template, start=1, max=20}` | `?page={n}` / `/page/{n}` appended to index seeds until no new links. | | `schedule` | `{group: interval}`, default `{seed: weekly}` | Interval per group: `hourly`, `daily`, `weekly`, `biweekly`, `monthly`, `quarterly`, `never` or `30m` / `3h` / `2d` / `1w` (`intervalMs`). The adaptive scheduler halves it for pages that change and stretches it up to ×4 for static ones (`apps/worker/src/scheduling.ts`). | | `extractors.` | see below | Which pages produce which records. | | `defaults` | `{operatorName, operatorKind, countryIso2, facilityType, status, isHyperscale, providerName}` | Merged into every normalized entity when the page did not say otherwise (`GenericConnector.normalize`). | | `implementation` | registered name | Replaces the generic pipeline (§ 2.2). | | `params` | free object | Passed to the implementation. | | `parserVersion` | default `v1` | Human-readable version prefix (§ 6). | ### Extractors (`extractorSchema`) ```yaml extractors: facility: match: ["/data-centers/[a-z0-9-]+/"] # regex on finalUrl or url (case-insensitive); absent = every page pageTypes: [facility_page] # additionally restrict to classified page types kind: facility # facility | operator | campus | cloud_region | ixp | project | news_event | country parser: databank_facility_v1 # code-backed (§ 2.1) — OR declarative `fields:` below params: { keepText: true } # passed to Parser.parse(doc, ctx, params) address: # declarative example kind: facility each: "div.location-card" # one record per matching container (optional) key: "acme:{code}" # stable key template; {url} = finalUrl; default acme: minCertainty: 0.5 # certainty = filled fields / max(3, 0.6 × declared fields) fields: name: "h1" # a bare string is a CSS selector (text) code: { selector: "h1", regex: "\\b([A-Z]{2,4}\\d{1,3})\\b", group: 1 } itCapacityMw: { selector: ".stat-power", transform: mw } lat: { jsonld: "Place", jsonPath: "$.geo.latitude", transform: float } city: { meta: "og:locality" } ``` `selectorRule` accepts `selector`, `attr`, `regex` / `regexFlags` / `group`, `all`, `jsonPath` (embedded JSON / JSON-LD), `jsonld` (@type to search), `meta` (``), `transform` (`mw sqm ha money date status type country int float trim lower url bool`) and `default` — evaluated by `evalRule` in `packages/connectors/src/extract.ts`. Field names are the `NormalizedFacility` / `NormalizedProject` / … attributes (`name`, `address`, `city`, `regionName`, `postalCode`, `countryIso2` or `country`, `lat`/`lng` + `geoPrecision`, `itCapacityMw`, `totalPowerMw`, `plannedPowerMw`, `buildingSqm`, `siteAreaHa`, `pue`, `rackCount`, `tier`, `status`/`statusText`, `facilityType`, `certifications`, `cloudProviders`, `renewableClaim`, `coolingType`, `openedOn`, `description`, `externalIds`, …). When a page matches **no** extractor but is classified as an announcement (press release, project, expansion, planning document, …) and is data-center relevant, `GenericConnector.extract` emits a fallback `news_event` (and, for non-`news` kinds only, a `project` when `fallbackProjectGate` sees a MW figure and a pipeline verb in the headline). ## 2. Code hooks ### 2.1 Parser (`registerParser`, `Parser` in `packages/connectors/src/types.ts`) ```ts export interface Parser { name: string; // referenced from YAML: extractors..parser version: string; // part of the effective extractor version (§ 6) — bump when the output changes pageTypes?: PageType[]; parse(doc: RawDocument, ctx: ConnectorContext, params?: Record): ExtractedRecord[] | Promise; } ``` A parser is a pure function of the archived document: **no network, no geocoding, no guessed figures** — it surfaces what the page publishes, with a per-field `methods` map for provenance (`{ itCapacityMw: "regex:it_mw_v1", lat: "json-ld:GeoCoordinates" }`). Parsers live in `apps/worker/src/connectors//.ts` and are registered by the group's `index.ts`: - `operators1/` (Equinix, Digital Realty, NTT, CyrusOne, QTS, Vantage, STACK, CoreSite, Switch): each file exports a `Parser` object; `operators1/index.ts` `register()` calls `registerParser` for each. Helpers in `operators1/shared.ts` (`parseNaAddress`, `parseEuAddress`, `parseCityLine`, `mwAfter`, `mwWithContext`, `areaSqm`, `certificationsFrom`, `record(kind, key, url, data, methods, certainty)` which drops empty fields and sets `pageType: facility_page`). - `operators2/` (DataBank, EdgeConneX, TierPoint, NEXTDC, AirTrunk, atNorth, CloudHQ, …, grouped in `americas.ts` / `emea.ts` / `apac.ts`): parsers are declared with **`facilityParser(name, (page, doc) => ExtractedRecord[], version = OPERATORS2_PARSER_VERSION)`** from `operators2/shared.ts`, which registers on import and hands you a `Page` (`$`, `html`, visible `text` with nav/header/footer removed, `title`, `h1`, `desc`, `url`). Build records with the **`Rec`** builder: `new Rec().set(field, value, method)` ignores null/empty/NaN and records the method; `.done(key, url, certainty = 0.85, kind = "facility")`. Other helpers: `itMw`, `capacityMw`, `mw` (handles `1.125 MW` decimals, `MWs`, and "at least" figures `10+MW` / `20MW+` — pass the raw string, `parseMw` reads the stated figure), `sqm`, `ha`, `pue`, `tier`, `certifications`, `usAddress`, `ldWithAddress` / `applyLdAddress` / `applyLdGeo` / `breadcrumbLeaf` (JSON-LD), `countryOf` / `countryFromCity` (deterministic metro → ISO table, not geocoding), `statusOf`, `key(op, ...parts)`. `operators2/index.ts` `register()` only checks that every name in `AMERICAS_PARSERS` / `EMEA_PARSERS` / `APAC_PARSERS` is registered. A stat regex that reads a figure before its label must accept the `+` suffix: `/(\d[\d.,]*\+?\s*MW\+?)\s*\n?\s*IT Capacity/`. - `news/` (`news_article_v1`, `news_planning_pdf_v1`, `news_edgar_fts_v1` + the `news_edgar_fts` implementation): article → `news_event` (+ `project` when `qualifiesAsProject`) via `articleContent(doc)` → `extractAnnouncement(title, text)` (`extract-project.ts`, unit-tested in `extract-project.test.ts`). `articleContent` uses `mainText` (`packages/connectors/src/extract.ts`): the largest `
` / `
` block after removing nav / header / footer / aside / related / recommended / trending / popular / sidebar / comments blocks, cut at the first trailing teaser heading ("More in …", "Related articles", "Tags"). Operators are looked up in the title plus the first `OPERATOR_SCOPE_CHARS` (4 000) characters only; money figures in a company-background sentence (`MONEY_PORTFOLIO_RE`: "investment volume", "revenue", "assets under management"…) are never the project's investment. Params: `keepText` (public-sector sources only), `minProjectMw` (5), `minInvestmentUsd` (50 000 000), `minAcres` (100), `lenient`. - `cloud/` (AWS, Azure, GCP, Oracle, IBM, Alibaba, Tencent, Meta regions) and `datasets/` (PeeringDB, Wikidata, World Bank, OSM Overpass): mostly `implementation`s. `apps/worker/src/connectors/index.ts` `registerAllConnectors()` imports every `/index.ts` found on disk and calls its `register()`; a group that throws is reported and the others still load. Make `register()` idempotent (guard with `listParsers()` or a module flag) — the worker, the CLI, `try-connector` and the fixture tests may all call it. Add a new group by creating `apps/worker/src/connectors//index.ts` with `export function register(): void`. ### 2.2 Implementation (`registerImplementation(name, (cfg) => Connector)`) For APIs and datasets where discovery/fetch/extract do not fit the generic HTML pipeline, register a factory returning a full `Connector` (`discover`, `fetch`, `extract`, `normalize`, `validate`; see `datasets/peeringdb.ts`). Reference it with `implementation: ` in the YAML; `buildConnector` (`apps/worker/src/configs.ts`) prefers it over `GenericConnector`. Keep `validate` delegating to `validateEntities` (`packages/connectors/src/generic.ts`) unless you have a reason not to. ### 2.3 Validation (`validateEntities`) Hard errors reject the record (implausible MW `≤ 0` or `> 10 000`; a single facility `> 1 000 MW` without a campus designation; PUE outside 1–3; building `> 500 000 m²` without campus; integer coordinates flagged `exact`; title-like facility names such as "Data Centers in Virginia | Acme"; opening year `< 1950`; dates `> now + 12 y`; bad ISO-2). Warnings are recorded (no country and no geo, IT > total power, …). A connector must produce **zero errors** on its fixtures and on a live dry-run before it ships. ## 3. Fixtures and golden tests Real archived bodies replay through the exact production path — `parseConnectorConfig` → registered parsers → `connector.extract` → `connector.normalize` → `validateEntities` — with **no network**. Test: `apps/worker/src/connectors/fixtures.test.ts`. Data: `apps/worker/fixtures//`. ### 3.1 Layout ``` apps/worker/fixtures/ databank/ dfw6-love-field.html # raw archived body (html | json | xml | pdf | txt | md), < 400 KB dfw6-love-field.expected.json # sibling expectations datacenterdynamics/ sb-gruppe-bucharest.html sb-gruppe-bucharest.expected.json ``` The directory name **is** the connector id (the test loads `config/connectors/.yaml`). Adding a fixture = adding the two files; no code changes. ### 3.2 Pulling a body from the production archive (data node BHS128) ```bash # 1. pick a document (Postgres in the dci-postgres-1 container) ssh BHS128 'docker exec -i dci-postgres-1 psql -U dci -d dci -Atc " select id, url, content_type, extract_count, last_fetched from documents where connector_id = '\''databank'\'' and storage_key is not null and extract_count > 0 and status_code = 200 and url ~ '\''/data-centers/[a-z0-9-]+/'\'' order by last_fetched desc limit 10"' # 2. raw archived body through the admin API (x-dci-admin-token from deploy/.env.data) ssh BHS128 'TOK=$(sed -n "s/^DCI_ADMIN_TOKEN=//p" /srv/dci/app/deploy/.env.data); curl -s -H "x-dci-admin-token: $TOK" http://10.67.0.60:8300/api/admin/documents//raw' \ > apps/worker/fixtures//.html # fallback: cd /srv/dci/app && docker compose -f deploy/compose.data.yml --env-file deploy/.env.data run --rm cli inspect --body --json ``` Rules for the body: keep it byte-for-byte (no reformatting, no CRLF normalisation); keep it under 400 KB — if a page is bigger, strip `