Connector authoring guide
A connector is one source (an operator website, a publisher feed, a government newsroom, a registry API, a dataset) described by one YAML file in config/connectors/<id>.yaml, optionally backed by code registered from apps/worker/src/connectors/<group>/index.ts. The runtime (apps/worker) turns it into: discover → fetch → archive → extract → normalize → validate → reconcile → publish. This guide covers the YAML, the two code hooks (parser / implementation), the fixture-based golden tests, live dry-runs, quarantine, parser versioning, health states and the legal rules. Everything below is grounded in the code paths named in brackets.
Per-source notes (what the site publishes, quirks, verification log) live in docs/connectors/<id>.md — one file per connector, written when the connector ships.
1. The YAML (connectorConfigSchema, packages/connectors/src/config.ts)
parseConnectorConfig(text, file) validates with zod and applies defaults; loadConnectorConfigs(dir) loads the whole directory and rejects duplicate ids. pnpm dci validate runs the same schema (plus the superRefine cross-field rules: fetch.level ≤ fetch.maxLevel, seed / classify minLevel ≤ fetch.maxLevel, every regex compiles).
| Field | Type / default | Meaning |
|---|---|---|
id |
^[a-z0-9][a-z0-9-]*$ |
Connector id = file name = fixture dir = connectors.id. Keys are prefixed with it (databank:dfw6). |
name, domain |
string | Source name and host (domain drives the per-host rate limiter, configureHost). |
kind |
SOURCE_KINDS (operator, news, government, filing, utility, cloud_provider, registry, …) |
Drives default provenance confidence (high for operator/government/filing/utility/cloud_provider/registry, else moderate) and the third-party rules (kind: news, § 8). |
mode |
html | sitemap | rss | pdf | json | hybrid | api | dataset (default hybrid) |
Informational (Connector.type). |
enabled |
bool, default true |
Parked connectors keep their YAML with enabled: false. |
priority |
1 (best) … 5, default 3 | Source priority used by reconciliation when sources disagree. |
license, attribution, homepage, notes, coverage{countries,operators,global} |
optional | Recorded on sources; coverage feeds admin + scheduling. |
fetch.level / fetch.maxLevel |
1–4, defaults 1 / 2 | Starting escalation level and ceiling: L1 direct, L2 direct + browser identity (userAgent: browser), L3 Firecrawl, L4 Scrapfly rendered. The runtime escalates only on 403/429/JS-shell detection. |
fetch.renderJs, country, waitForSelector, accept, headers, timeoutMs, maxBytes |
optional | Passed to the fetchers. |
fetch.userAgent |
bot | browser, default bot |
browser sends BROWSER_UA at L2. |
fetch.rpm, fetch.concurrency |
1–600 (default 30), 1–8 (default 2) | Per-host rate limit (packages/connectors/src/ratelimit.ts). |
fetch.respectRobots |
default true |
Never set to false for a public website (§ 8). |
fetch.maxCreditsPerRun |
default 200 | Premium (Scrapfly/Firecrawl) credit budget per run; maxLevelForBudget drops to L2 once spent. |
fetch.maxCreditsPerDay |
optional | Per-connector UTC-day budget (Redis counter, all workers) on top of the global DCI_SCRAPFLY_DAILY_BUDGET / DCI_FIRECRAWL_DAILY_BUDGET. |
discovery.sitemap |
true (robots.txt sitemaps) or list of URLs |
Sitemap / sitemap-index walk. Discovery fetches (sitemap, rss groups) never use premium fetchers (DISCOVERY_GROUPS). |
discovery.rss |
list of feed URLs | RSS/Atom items → documents (title/published/summary hints in doc.meta). |
discovery.seeds |
strings or {url, group, pageType, minLevel} |
Fixed pages (index pages, single documents). |
discovery.follow |
regex list | Internal links followed from index pages only. |
discovery.include / exclude |
regex lists | URL allow / deny filters applied to everything discovered. |
discovery.classify |
[{pattern, group, pageType?, priority?, minLevel?}] |
First matching pattern assigns the schedule group and a page-type hint. |
discovery.maxUrlsPerRun |
default 2000 | Cap per run. |
discovery.pagination |
{template, start=1, max=20} |
?page={n} / /page/{n} appended to index seeds until no new links. |
schedule |
{group: interval}, default {seed: weekly} |
Interval per group: hourly, daily, weekly, biweekly, monthly, quarterly, never or 30m / 3h / 2d / 1w (intervalMs). The adaptive scheduler halves it for pages that change and stretches it up to ×4 for static ones (apps/worker/src/scheduling.ts). |
extractors.<name> |
see below | Which pages produce which records. |
defaults |
{operatorName, operatorKind, countryIso2, facilityType, status, isHyperscale, providerName} |
Merged into every normalized entity when the page did not say otherwise (GenericConnector.normalize). |
implementation |
registered name | Replaces the generic pipeline (§ 2.2). |
params |
free object | Passed to the implementation. |
parserVersion |
default v1 |
Human-readable version prefix (§ 6). |
Extractors (extractorSchema)
extractors:
facility:
match: ["/data-centers/[a-z0-9-]+/"] # regex on finalUrl or url (case-insensitive); absent = every page
pageTypes: [facility_page] # additionally restrict to classified page types
kind: facility # facility | operator | campus | cloud_region | ixp | project | news_event | country
parser: databank_facility_v1 # code-backed (§ 2.1) — OR declarative `fields:` below
params: { keepText: true } # passed to Parser.parse(doc, ctx, params)
address: # declarative example
kind: facility
each: "div.location-card" # one record per matching container (optional)
key: "acme:{code}" # stable key template; {url} = finalUrl; default acme:<urlFingerprint>
minCertainty: 0.5 # certainty = filled fields / max(3, 0.6 × declared fields)
fields:
name: "h1" # a bare string is a CSS selector (text)
code: { selector: "h1", regex: "\\b([A-Z]{2,4}\\d{1,3})\\b", group: 1 }
itCapacityMw: { selector: ".stat-power", transform: mw }
lat: { jsonld: "Place", jsonPath: "$.geo.latitude", transform: float }
city: { meta: "og:locality" }selectorRule accepts selector, attr, regex / regexFlags / group, all, jsonPath (embedded JSON / JSON-LD), jsonld (@type to search), meta (<meta property|name>), transform (mw sqm ha money date status type country int float trim lower url bool) and default — evaluated by evalRule in packages/connectors/src/extract.ts. Field names are the NormalizedFacility / NormalizedProject / … attributes (name, address, city, regionName, postalCode, countryIso2 or country, lat/lng + geoPrecision, itCapacityMw, totalPowerMw, plannedPowerMw, buildingSqm, siteAreaHa, pue, rackCount, tier, status/statusText, facilityType, certifications, cloudProviders, renewableClaim, coolingType, openedOn, description, externalIds, …).
When a page matches no extractor but is classified as an announcement (press release, project, expansion, planning document, …) and is data-center relevant, GenericConnector.extract emits a fallback news_event (and, for non-news kinds only, a project when fallbackProjectGate sees a MW figure and a pipeline verb in the headline).
2. Code hooks
2.1 Parser (registerParser, Parser in packages/connectors/src/types.ts)
export interface Parser {
name: string; // referenced from YAML: extractors.<x>.parser
version: string; // part of the effective extractor version (§ 6) — bump when the output changes
pageTypes?: PageType[];
parse(doc: RawDocument, ctx: ConnectorContext, params?: Record<string, unknown>): ExtractedRecord[] | Promise<ExtractedRecord[]>;
}A parser is a pure function of the archived document: no network, no geocoding, no guessed figures — it surfaces what the page publishes, with a per-field methods map for provenance ({ itCapacityMw: "regex:it_mw_v1", lat: "json-ld:GeoCoordinates" }). Parsers live in apps/worker/src/connectors/<group>/<file>.ts and are registered by the group's index.ts:
operators1/(Equinix, Digital Realty, NTT, CyrusOne, QTS, Vantage, STACK, CoreSite, Switch): each file exports aParserobject;operators1/index.tsregister()callsregisterParserfor each. Helpers inoperators1/shared.ts(parseNaAddress,parseEuAddress,parseCityLine,mwAfter,mwWithContext,areaSqm,certificationsFrom,record(kind, key, url, data, methods, certainty)which drops empty fields and setspageType: facility_page).operators2/(DataBank, EdgeConneX, TierPoint, NEXTDC, AirTrunk, atNorth, CloudHQ, …, grouped inamericas.ts/emea.ts/apac.ts): parsers are declared withfacilityParser(name, (page, doc) => ExtractedRecord[], version = OPERATORS2_PARSER_VERSION)fromoperators2/shared.ts, which registers on import and hands you aPage($,html, visibletextwith nav/header/footer removed,title,h1,desc,url). Build records with theRecbuilder:new Rec().set(field, value, method)ignores null/empty/NaN and records the method;.done(key, url, certainty = 0.85, kind = "facility"). Other helpers:itMw,capacityMw,mw(handles1.125 MWdecimals,MWs, and "at least" figures10+MW/20MW+— pass the raw string,parseMwreads the stated figure),sqm,ha,pue,tier,certifications,usAddress,ldWithAddress/applyLdAddress/applyLdGeo/breadcrumbLeaf(JSON-LD),countryOf/countryFromCity(deterministic metro → ISO table, not geocoding),statusOf,key(op, ...parts).operators2/index.tsregister()only checks that every name inAMERICAS_PARSERS/EMEA_PARSERS/APAC_PARSERSis registered. A stat regex that reads a figure before its label must accept the+suffix:/(\d[\d.,]*\+?\s*MW\+?)\s*\n?\s*IT Capacity/.news/(news_article_v1,news_planning_pdf_v1,news_edgar_fts_v1+ thenews_edgar_ftsimplementation): article →news_event(+projectwhenqualifiesAsProject) viaarticleContent(doc)→extractAnnouncement(title, text)(extract-project.ts, unit-tested inextract-project.test.ts).articleContentusesmainText(packages/connectors/src/extract.ts): the largest<article>/<main>block after removing nav / header / footer / aside / related / recommended / trending / popular / sidebar / comments blocks, cut at the first trailing teaser heading ("More in …", "Related articles", "Tags"). Operators are looked up in the title plus the firstOPERATOR_SCOPE_CHARS(4 000) characters only; money figures in a company-background sentence (MONEY_PORTFOLIO_RE: "investment volume", "revenue", "assets under management"…) are never the project's investment. Params:keepText(public-sector sources only),minProjectMw(5),minInvestmentUsd(50 000 000),minAcres(100),lenient.cloud/(AWS, Azure, GCP, Oracle, IBM, Alibaba, Tencent, Meta regions) anddatasets/(PeeringDB, Wikidata, World Bank, OSM Overpass): mostlyimplementations.
apps/worker/src/connectors/index.ts registerAllConnectors() imports every <group>/index.ts found on disk and calls its register(); a group that throws is reported and the others still load. Make register() idempotent (guard with listParsers() or a module flag) — the worker, the CLI, try-connector and the fixture tests may all call it. Add a new group by creating apps/worker/src/connectors/<group>/index.ts with export function register(): void.
2.2 Implementation (registerImplementation(name, (cfg) => Connector))
For APIs and datasets where discovery/fetch/extract do not fit the generic HTML pipeline, register a factory returning a full Connector (discover, fetch, extract, normalize, validate; see datasets/peeringdb.ts). Reference it with implementation: <name> in the YAML; buildConnector (apps/worker/src/configs.ts) prefers it over GenericConnector. Keep validate delegating to validateEntities (packages/connectors/src/generic.ts) unless you have a reason not to.
2.3 Validation (validateEntities)
Hard errors reject the record (implausible MW ≤ 0 or > 10 000; a single facility > 1 000 MW without a campus designation; PUE outside 1–3; building > 500 000 m² without campus; integer coordinates flagged exact; title-like facility names such as "Data Centers in Virginia | Acme"; opening year < 1950; dates > now + 12 y; bad ISO-2). Warnings are recorded (no country and no geo, IT > total power, …). A connector must produce zero errors on its fixtures and on a live dry-run before it ships.
3. Fixtures and golden tests
Real archived bodies replay through the exact production path — parseConnectorConfig → registered parsers → connector.extract → connector.normalize → validateEntities — with no network. Test: apps/worker/src/connectors/fixtures.test.ts. Data: apps/worker/fixtures/<connectorId>/.
3.1 Layout
apps/worker/fixtures/
databank/
dfw6-love-field.html # raw archived body (html | json | xml | pdf | txt | md), < 400 KB
dfw6-love-field.expected.json # sibling expectations
datacenterdynamics/
sb-gruppe-bucharest.html
sb-gruppe-bucharest.expected.jsonThe directory name is the connector id (the test loads config/connectors/<dir>.yaml). Adding a fixture = adding the two files; no code changes.
3.2 Pulling a body from the production archive (data node BHS128)
# 1. pick a document (Postgres in the dci-postgres-1 container)
ssh BHS128 'docker exec -i dci-postgres-1 psql -U dci -d dci -Atc "
select id, url, content_type, extract_count, last_fetched from documents
where connector_id = '\''databank'\'' and storage_key is not null and extract_count > 0 and status_code = 200
and url ~ '\''/data-centers/[a-z0-9-]+/'\''
order by last_fetched desc limit 10"'
# 2. raw archived body through the admin API (x-dci-admin-token from deploy/.env.data)
ssh BHS128 'TOK=$(sed -n "s/^DCI_ADMIN_TOKEN=//p" /srv/dci/app/deploy/.env.data);
curl -s -H "x-dci-admin-token: $TOK" http://10.67.0.60:8300/api/admin/documents/<doc_id>/raw' \
> apps/worker/fixtures/<connector>/<name>.html
# fallback: cd /srv/dci/app && docker compose -f deploy/compose.data.yml --env-file deploy/.env.data run --rm cli inspect <doc_id> --body --jsonRules for the body: keep it byte-for-byte (no reformatting, no CRLF normalisation); keep it under 400 KB — if a page is bigger, strip <script> blocks only when the parser does not read them (Equinix and Digital Realty parsers read JS constants / __NEXT_DATA__; NTT is selector-only, so ntt/chicago-ch3.html is script-stripped); never remove the region the parser needs (address block, MW stat, JSON-LD). Redact personal data of individuals (journalist bylines, staff e-mails); functional mailboxes (sales@, press@) and executive quotes in press releases stay. Check for e-mails with grep -oiE "[a-z0-9._%+-]+@[a-z0-9.-]+\.[a-z]{2,}" <file>.
3.3 expected.json
{
"url": "https://www.databank.com/data-centers/dallas/love-field/", // used as url + finalUrl of the RawDocument
"documentId": "doc_352f7e263b5afaed", // documents.id on the data node
"fetchedAt": "2026-09-11T08:02:55.127Z", // documents.last_fetched
"contentType": "text/html; charset=UTF-8", // documents.content_type (a /pdf/ type switches on PDF parsing)
"note": "why this page: '0.675MW Critical IT Load' — the decimal must survive",
"counts": { "facility": 1 }, // asserted per entityType; omit a type to leave it unasserted
"facilities": [ // also: projects, newsEvents, cloudRegions, operators
{
"key": "databank:dfw6", // entities are matched by key
"name": "Dallas Love Field Data Center (DFW6)",
"operatorName": "DataBank", "city": "Dallas", "regionName": "TX", "countryIso2": "US",
"itCapacityMw": 0.675, "totalPowerMw": null, // null = absent or null; numbers are compared exactly
"mentions.countriesIso2": ["US"], // dotted paths reach nested fields
"_todo": { // CORRECT values the parser does not produce yet → it.fails
"totalPowerMw": 36,
"_reason": "the figure precedes its label ('36 MW total power'); the stat regex only reads label-first"
}
}
],
"thirdParty": true, // kind: news — summaries ≤ THIRD_PARTY_SUMMARY_MAX, no article text in records
"announcement": { // news / government: extractAnnouncement(title, text) facts
"class": "NEW_BUILD", "mayCreateProject": true, "headlineMw": 12, "status": "announced",
"operator": "Applied Digital", "country": "RO", "city": "Bucharest", "money": { "amount": 7500000000, "currency": "USD" },
"_todo": { "investmentUsd": null, "_reason": "…" }
}
}Announcement facts available: title, published, class, mayCreateProject, physical, headlineMw, mwAll, status, operator, operators, country, city, region, money, investmentUsd, acreage, phaseCount, expectedOpening, projectName, explicitName, relevance, leadRelevance, titleSignal, nonProjectTitle, pageType, eventType.
Every fixture always asserts zero validation errors. Expected values must be what the parser produces today and what the source page says: verify each figure against the body before writing it. When the parser is wrong, put the correct value under _todo with a _reason — the test runs as it.fails and starts failing the day the parser is fixed, which is the signal to promote the field. Never assert a value you believe is wrong.
3.4 Running
pnpm --filter @dci/worker test # everything (fixtures included)
pnpm --filter @dci/worker exec vitest run src/connectors/fixtures.test.ts -t databank # one connectorGolden coverage on 2026-09-12: 21 fixtures over 18 connectors (equinix, digitalrealty, stack, databank ×2, ntt, edgeconnex ×2, cyrusone, qts ×2, vantage, coresite, airtrunk, nextdc ×2, atnorth, cloudhq, datacenterfrontier, datacenterdynamics, loudoun-county), no open _todo (the nine gaps recorded on 2026-09-12 — EdgeConneX city from the h1, NEXTDC 20MW+, QTS figure-first capacity and template-copied map pin, DCD teaser operator and portfolio money — were fixed the same day and promoted). No planning PDF fixture yet: the archive holds no document with a PDF content type.
4. Live dry-run
pnpm tsx scripts/try-connector.ts config/connectors/<id>.yaml [--limit 5] [--url <single url>] [--verbose] [--json out.json]
pnpm dci run <id> --dry-run --limit 5 [--url u] # same through the worker CLI (needs the DB env)
pnpm dci trace <doc-id-or-url> [--live] # one document end to endtry-connector registers every apps/worker/src/connectors/<group>/index.ts, builds the connector from the YAML (implementation or GenericConnector), runs discover (or the single --url), fetches with the real robots / rate-limit / escalation policy through createTestContext (in-memory state, nothing persisted, dryRun: true) and prints records, entities and the ValidationReport. Nothing ships without a live dry-run yielding valid entities and 0 rejected; record the result in docs/connectors/<id>.md ("Verification (live, )").
5. Quarantine
Two different things share the word:
- Connector quarantine (
connectors.quarantine,pnpm dci quarantine <id> [on|off]): the run does everything — discover, fetch, archive, extract, validate — but the ingest transaction is rolled back, so nothing is published; the run is stored withquarantined = truefor review.runConnector(apps/worker/src/pipeline.ts) reads the flag before each run. It is switched on automatically when a connector fails or is blocked 5 runs in a row (consecutive_failures ≥ 5) or when a non-fullrun creates ≥ 500 records (suspiciousSpike): both are logged as… — quarantined pending review. New connectors are worth running quarantined for the first full pass. - Document quarantine (
documents.quarantined):nextCheckAfter(apps/worker/src/scheduling.ts) backs a 404/410 off ×4 and after 3 consecutive gone responses (MAX_CONSECUTIVE_GONE) setsnextCheck = nullandquarantined = true— the URL is never scheduled again. Other errors back off ×2^(n−1) (cap ×8); premium retries after two failures are limited to every 4th attempt (premiumAllowedAfterErrors). Weekly check indocs/CRAWL-OPERATIONS.md:select count(*) from documents where quarantinedshould grow slowly.
6. Parser versioning and re-extraction
The effective extractor version stored on every document (documents.extractor_version) is computed by effectiveExtractorVersion(cfg, connector) in apps/worker/src/configs.ts:
<cfg.parserVersion>+<sha256("news_article_v1@news_v3|databank_facility_v1@v2|impl:peeringdb@v1")[0:10]>i.e. the YAML parserVersion plus a hash of every referenced Parser.name@Parser.version (and impl:<name>@<connector.parserVersion>). Consequences:
- Bumping
versionon theParserobject is what invalidates cached extractions.shouldSkipExtraction(scheduling.ts) skips a document only when its content fingerprint is unchanged and the stored extractor version equals the current one (and the run is not--force). Change a regex without bumping the version and unchanged pages keep their old records until their content changes. - Re-run the archive without network:
pnpm dci reprocess <id> [--limit n] [--group g] --stalere-extracts from stored bodies (MinIO) only the documents whoseextractor_versiondiffers from the current one; without--staleeverything is re-extracted.pnpm dci run <id> --task reprocessis the same thing. - Declarative connectors have no parser: bump
parserVersionin the YAML.
7. Health states (connectorHealthFrom, apps/worker/src/scheduling.ts)
connectors.health after each run (admin health center; docs/CRAWL-OPERATIONS.md § "Reading a bad run"):
| State | Rule |
|---|---|
ok |
run status not failed, fetch failure rate < 20 % (healthFrom) |
degraded |
failure rate 20–50 % |
failing |
run status = failed, or failure rate ≥ 50 % |
blocked |
≥ 50 % of fetched pages came back blocked (403/429/CAPTCHA/JS shell) — fix the fetch level / user agent, never bypass |
schema_change |
≥ 5 pages fetched fine and passed to extractors but 0 entities on a connector that has documents — the site's HTML changed: re-check selectors against a fresh fixture |
no_new_content |
discovery yielded 0 URLs and 0 fetches although a previous run discovered some — sitemap/RSS moved or emptied |
never_run |
no run yet (also the state of enabled: false connectors) |
8. Legal and content rules (non-negotiable)
fetch.respectRobots: true—createTestContextand the runtime checkrobots.txtbefore every fetch and honourcrawl-delay; a disallowed URL yieldserror.code = robots_disallow, not a fetch.- No authentication, paywall or CAPTCHA bypass; escalation (L3/L4) is for JS shells and bot walls on public pages only, within the credit budgets.
- Conditional GET (ETag / Last-Modified) and content hashing: unchanged pages are never re-extracted, premium credits are never spent re-fetching unchanged content.
- Third-party publishers (
kind: news) keep only title, date, link, extracted facts and a summary of at mostTHIRD_PARTY_SUMMARY_MAX = 400characters (trimSummary,packages/connectors/src/generic.ts); the article text is never stored (news_article_v1keepTextdefaults to false; the fixture test enforces both whenthirdParty: true). Only public-sector sources (kind: government,utility,filing) may setkeepText: true. license/attributionare mandatory in the YAML and are shown with every record (sources.license,sources.attribution).- Never invent data: no geocoding, no guessed capacities, no default coordinates. City-level coordinates are
geoPrecision: city, operator-published markersexact/parcel. An estimate is flaggedisEstimate.
9. Checklist for a new connector
config/connectors/<id>.yaml— id, name, domain, kind, priority, license/attribution, fetch, discovery (sitemap first, then include/exclude/classify), schedule, extractors, defaults.pnpm dci validate.- Declarative
fieldswhen the page is regular; otherwise a parser inapps/worker/src/connectors/<group>/<file>.tsregistered in<group>/index.tsregister()(facilityParser+Recfor facilities), or animplementationfor APIs / datasets. - Live dry-run:
pnpm tsx scripts/try-connector.ts config/connectors/<id>.yaml --limit 5→ valid entities, 0 rejected, 0 unexpected premium credits. - Fixture: archive a real body under
apps/worker/fixtures/<id>/, writeexpected.jsonfrom what the parser produces after checking each value against the page,pnpm --filter @dci/worker test. docs/connectors/<id>.md— what it collects, discovery, schedule & fetch, quirks, verification log.- First production run quarantined (
pnpm dci quarantine <id> on, run, reviewconnector_runs, thenoff). - Any later change to a parser's output → bump
Parser.version→pnpm dci reprocess <id> --stale.