SPB Git forge
38commits 1branches 0releases
338.7 MBsize
maindefault branch
2 h agolast push
HTML 53.9% TypeScript 44.5% JavaScript 0.6% SQL 0.5%
25.7 KB

# Connector authoring guide

A connector is one source (an operator website, a publisher feed, a government newsroom, a registry API, a dataset) described by one YAML file in config/connectors/<id>.yaml, optionally backed by code registered from apps/worker/src/connectors/<group>/index.ts. The runtime (apps/worker) turns it into: discover → fetch → archive → extract → normalize → validate → reconcile → publish. This guide covers the YAML, the two code hooks (parser / implementation), the fixture-based golden tests, live dry-runs, quarantine, parser versioning, health states and the legal rules. Everything below is grounded in the code paths named in brackets.

Per-source notes (what the site publishes, quirks, verification log) live in docs/connectors/<id>.md — one file per connector, written when the connector ships.

# 1. The YAML (connectorConfigSchema, packages/connectors/src/config.ts)

parseConnectorConfig(text, file) validates with zod and applies defaults; loadConnectorConfigs(dir) loads the whole directory and rejects duplicate ids. pnpm dci validate runs the same schema (plus the superRefine cross-field rules: fetch.level ≤ fetch.maxLevel, seed / classify minLevel ≤ fetch.maxLevel, every regex compiles).

Field Type / default Meaning
id ^[a-z0-9][a-z0-9-]*$ Connector id = file name = fixture dir = connectors.id. Keys are prefixed with it (databank:dfw6).
name, domain string Source name and host (domain drives the per-host rate limiter, configureHost).
kind SOURCE_KINDS (operator, news, government, filing, utility, cloud_provider, registry, …) Drives default provenance confidence (high for operator/government/filing/utility/cloud_provider/registry, else moderate) and the third-party rules (kind: news, § 8).
mode html | sitemap | rss | pdf | json | hybrid | api | dataset (default hybrid) Informational (Connector.type).
enabled bool, default true Parked connectors keep their YAML with enabled: false.
priority 1 (best) … 5, default 3 Source priority used by reconciliation when sources disagree.
license, attribution, homepage, notes, coverage{countries,operators,global} optional Recorded on sources; coverage feeds admin + scheduling.
fetch.level / fetch.maxLevel 1–4, defaults 1 / 2 Starting escalation level and ceiling: L1 direct, L2 direct + browser identity (userAgent: browser), L3 Firecrawl, L4 Scrapfly rendered. The runtime escalates only on 403/429/JS-shell detection.
fetch.renderJs, country, waitForSelector, accept, headers, timeoutMs, maxBytes optional Passed to the fetchers.
fetch.userAgent bot | browser, default bot browser sends BROWSER_UA at L2.
fetch.rpm, fetch.concurrency 1–600 (default 30), 1–8 (default 2) Per-host rate limit (packages/connectors/src/ratelimit.ts).
fetch.respectRobots default true Never set to false for a public website (§ 8).
fetch.maxCreditsPerRun default 200 Premium (Scrapfly/Firecrawl) credit budget per run; maxLevelForBudget drops to L2 once spent.
fetch.maxCreditsPerDay optional Per-connector UTC-day budget (Redis counter, all workers) on top of the global DCI_SCRAPFLY_DAILY_BUDGET / DCI_FIRECRAWL_DAILY_BUDGET.
discovery.sitemap true (robots.txt sitemaps) or list of URLs Sitemap / sitemap-index walk. Discovery fetches (sitemap, rss groups) never use premium fetchers (DISCOVERY_GROUPS).
discovery.rss list of feed URLs RSS/Atom items → documents (title/published/summary hints in doc.meta).
discovery.seeds strings or {url, group, pageType, minLevel} Fixed pages (index pages, single documents).
discovery.follow regex list Internal links followed from index pages only.
discovery.include / exclude regex lists URL allow / deny filters applied to everything discovered.
discovery.classify [{pattern, group, pageType?, priority?, minLevel?}] First matching pattern assigns the schedule group and a page-type hint.
discovery.maxUrlsPerRun default 2000 Cap per run.
discovery.pagination {template, start=1, max=20} ?page={n} / /page/{n} appended to index seeds until no new links.
schedule {group: interval}, default {seed: weekly} Interval per group: hourly, daily, weekly, biweekly, monthly, quarterly, never or 30m / 3h / 2d / 1w (intervalMs). The adaptive scheduler halves it for pages that change and stretches it up to ×4 for static ones (apps/worker/src/scheduling.ts).
extractors.<name> see below Which pages produce which records.
defaults {operatorName, operatorKind, countryIso2, facilityType, status, isHyperscale, providerName} Merged into every normalized entity when the page did not say otherwise (GenericConnector.normalize).
implementation registered name Replaces the generic pipeline (§ 2.2).
params free object Passed to the implementation.
parserVersion default v1 Human-readable version prefix (§ 6).

# Extractors (extractorSchema)

yaml
extractors:
  facility:
    match: ["/data-centers/[a-z0-9-]+/"]        # regex on finalUrl or url (case-insensitive); absent = every page
    pageTypes: [facility_page]                    # additionally restrict to classified page types
    kind: facility                                # facility | operator | campus | cloud_region | ixp | project | news_event | country
    parser: databank_facility_v1                  # code-backed (§ 2.1) — OR declarative `fields:` below
    params: { keepText: true }                    # passed to Parser.parse(doc, ctx, params)
  address:                                        # declarative example
    kind: facility
    each: "div.location-card"                     # one record per matching container (optional)
    key: "acme:{code}"                            # stable key template; {url} = finalUrl; default acme:<urlFingerprint>
    minCertainty: 0.5                             # certainty = filled fields / max(3, 0.6 × declared fields)
    fields:
      name: "h1"                                  # a bare string is a CSS selector (text)
      code: { selector: "h1", regex: "\\b([A-Z]{2,4}\\d{1,3})\\b", group: 1 }
      itCapacityMw: { selector: ".stat-power", transform: mw }
      lat: { jsonld: "Place", jsonPath: "$.geo.latitude", transform: float }
      city: { meta: "og:locality" }

selectorRule accepts selector, attr, regex / regexFlags / group, all, jsonPath (embedded JSON / JSON-LD), jsonld (@type to search), meta (<meta property|name>), transform (mw sqm ha money date status type country int float trim lower url bool) and default — evaluated by evalRule in packages/connectors/src/extract.ts. Field names are the NormalizedFacility / NormalizedProject / … attributes (name, address, city, regionName, postalCode, countryIso2 or country, lat/lng + geoPrecision, itCapacityMw, totalPowerMw, plannedPowerMw, buildingSqm, siteAreaHa, pue, rackCount, tier, status/statusText, facilityType, certifications, cloudProviders, renewableClaim, coolingType, openedOn, description, externalIds, …).

When a page matches no extractor but is classified as an announcement (press release, project, expansion, planning document, …) and is data-center relevant, GenericConnector.extract emits a fallback news_event (and, for non-news kinds only, a project when fallbackProjectGate sees a MW figure and a pipeline verb in the headline).

# 2. Code hooks

# 2.1 Parser (registerParser, Parser in packages/connectors/src/types.ts)

ts
export interface Parser {
  name: string;      // referenced from YAML: extractors.<x>.parser
  version: string;   // part of the effective extractor version (§ 6) — bump when the output changes
  pageTypes?: PageType[];
  parse(doc: RawDocument, ctx: ConnectorContext, params?: Record<string, unknown>): ExtractedRecord[] | Promise<ExtractedRecord[]>;
}

A parser is a pure function of the archived document: no network, no geocoding, no guessed figures — it surfaces what the page publishes, with a per-field methods map for provenance ({ itCapacityMw: "regex:it_mw_v1", lat: "json-ld:GeoCoordinates" }). Parsers live in apps/worker/src/connectors/<group>/<file>.ts and are registered by the group's index.ts:

  • operators1/ (Equinix, Digital Realty, NTT, CyrusOne, QTS, Vantage, STACK, CoreSite, Switch): each file exports a Parser object; operators1/index.ts register() calls registerParser for each. Helpers in operators1/shared.ts (parseNaAddress, parseEuAddress, parseCityLine, mwAfter, mwWithContext, areaSqm, certificationsFrom, record(kind, key, url, data, methods, certainty) which drops empty fields and sets pageType: facility_page).
  • operators2/ (DataBank, EdgeConneX, TierPoint, NEXTDC, AirTrunk, atNorth, CloudHQ, …, grouped in americas.ts / emea.ts / apac.ts): parsers are declared with facilityParser(name, (page, doc) => ExtractedRecord[], version = OPERATORS2_PARSER_VERSION) from operators2/shared.ts, which registers on import and hands you a Page ($, html, visible text with nav/header/footer removed, title, h1, desc, url). Build records with the Rec builder: new Rec().set(field, value, method) ignores null/empty/NaN and records the method; .done(key, url, certainty = 0.85, kind = "facility"). Other helpers: itMw, capacityMw, mw (handles 1.125 MW decimals, MWs, and "at least" figures 10+MW / 20MW+ — pass the raw string, parseMw reads the stated figure), sqm, ha, pue, tier, certifications, usAddress, ldWithAddress / applyLdAddress / applyLdGeo / breadcrumbLeaf (JSON-LD), countryOf / countryFromCity (deterministic metro → ISO table, not geocoding), statusOf, key(op, ...parts). operators2/index.ts register() only checks that every name in AMERICAS_PARSERS / EMEA_PARSERS / APAC_PARSERS is registered. A stat regex that reads a figure before its label must accept the + suffix: /(\d[\d.,]*\+?\s*MW\+?)\s*\n?\s*IT Capacity/.
  • news/ (news_article_v1, news_planning_pdf_v1, news_edgar_fts_v1 + the news_edgar_fts implementation): article → news_event (+ project when qualifiesAsProject) via articleContent(doc) → extractAnnouncement(title, text) (extract-project.ts, unit-tested in extract-project.test.ts). articleContent uses mainText (packages/connectors/src/extract.ts): the largest <article> / <main> block after removing nav / header / footer / aside / related / recommended / trending / popular / sidebar / comments blocks, cut at the first trailing teaser heading ("More in …", "Related articles", "Tags"). Operators are looked up in the title plus the first OPERATOR_SCOPE_CHARS (4 000) characters only; money figures in a company-background sentence (MONEY_PORTFOLIO_RE: "investment volume", "revenue", "assets under management"…) are never the project's investment. Params: keepText (public-sector sources only), minProjectMw (5), minInvestmentUsd (50 000 000), minAcres (100), lenient.
  • cloud/ (AWS, Azure, GCP, Oracle, IBM, Alibaba, Tencent, Meta regions) and datasets/ (PeeringDB, Wikidata, World Bank, OSM Overpass): mostly implementations.

apps/worker/src/connectors/index.ts registerAllConnectors() imports every <group>/index.ts found on disk and calls its register(); a group that throws is reported and the others still load. Make register() idempotent (guard with listParsers() or a module flag) — the worker, the CLI, try-connector and the fixture tests may all call it. Add a new group by creating apps/worker/src/connectors/<group>/index.ts with export function register(): void.

# 2.2 Implementation (registerImplementation(name, (cfg) => Connector))

For APIs and datasets where discovery/fetch/extract do not fit the generic HTML pipeline, register a factory returning a full Connector (discover, fetch, extract, normalize, validate; see datasets/peeringdb.ts). Reference it with implementation: <name> in the YAML; buildConnector (apps/worker/src/configs.ts) prefers it over GenericConnector. Keep validate delegating to validateEntities (packages/connectors/src/generic.ts) unless you have a reason not to.

# 2.3 Validation (validateEntities)

Hard errors reject the record (implausible MW ≤ 0 or > 10 000; a single facility > 1 000 MW without a campus designation; PUE outside 1–3; building > 500 000 m² without campus; integer coordinates flagged exact; title-like facility names such as "Data Centers in Virginia | Acme"; opening year < 1950; dates > now + 12 y; bad ISO-2). Warnings are recorded (no country and no geo, IT > total power, …). A connector must produce zero errors on its fixtures and on a live dry-run before it ships.

# 3. Fixtures and golden tests

Real archived bodies replay through the exact production path — parseConnectorConfig → registered parsers → connector.extract → connector.normalize → validateEntities — with no network. Test: apps/worker/src/connectors/fixtures.test.ts. Data: apps/worker/fixtures/<connectorId>/.

# 3.1 Layout

text
apps/worker/fixtures/
  databank/
    dfw6-love-field.html            # raw archived body (html | json | xml | pdf | txt | md), < 400 KB
    dfw6-love-field.expected.json   # sibling expectations
  datacenterdynamics/
    sb-gruppe-bucharest.html
    sb-gruppe-bucharest.expected.json

The directory name is the connector id (the test loads config/connectors/<dir>.yaml). Adding a fixture = adding the two files; no code changes.

# 3.2 Pulling a body from the production archive (data node BHS128)

bash
# 1. pick a document (Postgres in the dci-postgres-1 container)
ssh BHS128 'docker exec -i dci-postgres-1 psql -U dci -d dci -Atc "
  select id, url, content_type, extract_count, last_fetched from documents
  where connector_id = '\''databank'\'' and storage_key is not null and extract_count > 0 and status_code = 200
    and url ~ '\''/data-centers/[a-z0-9-]+/'\''
  order by last_fetched desc limit 10"'

# 2. raw archived body through the admin API (x-dci-admin-token from deploy/.env.data)
ssh BHS128 'TOK=$(sed -n "s/^DCI_ADMIN_TOKEN=//p" /srv/dci/app/deploy/.env.data);
  curl -s -H "x-dci-admin-token: $TOK" http://10.67.0.60:8300/api/admin/documents/<doc_id>/raw' \
  > apps/worker/fixtures/<connector>/<name>.html
# fallback: cd /srv/dci/app && docker compose -f deploy/compose.data.yml --env-file deploy/.env.data run --rm cli inspect <doc_id> --body --json

Rules for the body: keep it byte-for-byte (no reformatting, no CRLF normalisation); keep it under 400 KB — if a page is bigger, strip <script> blocks only when the parser does not read them (Equinix and Digital Realty parsers read JS constants / __NEXT_DATA__; NTT is selector-only, so ntt/chicago-ch3.html is script-stripped); never remove the region the parser needs (address block, MW stat, JSON-LD). Redact personal data of individuals (journalist bylines, staff e-mails); functional mailboxes (sales@, press@) and executive quotes in press releases stay. Check for e-mails with grep -oiE "[a-z0-9._%+-]+@[a-z0-9.-]+\.[a-z]{2,}" <file>.

# 3.3 expected.json

jsonc
{
  "url": "https://www.databank.com/data-centers/dallas/love-field/",   // used as url + finalUrl of the RawDocument
  "documentId": "doc_352f7e263b5afaed",                               // documents.id on the data node
  "fetchedAt": "2026-09-11T08:02:55.127Z",                            // documents.last_fetched
  "contentType": "text/html; charset=UTF-8",                          // documents.content_type (a /pdf/ type switches on PDF parsing)
  "note": "why this page: '0.675MW Critical IT Load' — the decimal must survive",
  "counts": { "facility": 1 },                     // asserted per entityType; omit a type to leave it unasserted
  "facilities": [                                  // also: projects, newsEvents, cloudRegions, operators
    {
      "key": "databank:dfw6",                      // entities are matched by key
      "name": "Dallas Love Field Data Center (DFW6)",
      "operatorName": "DataBank", "city": "Dallas", "regionName": "TX", "countryIso2": "US",
      "itCapacityMw": 0.675, "totalPowerMw": null,   // null = absent or null; numbers are compared exactly
      "mentions.countriesIso2": ["US"],             // dotted paths reach nested fields
      "_todo": {                                   // CORRECT values the parser does not produce yet → it.fails
        "totalPowerMw": 36,
        "_reason": "the figure precedes its label ('36 MW total power'); the stat regex only reads label-first"
      }
    }
  ],
  "thirdParty": true,                              // kind: news — summaries ≤ THIRD_PARTY_SUMMARY_MAX, no article text in records
  "announcement": {                                // news / government: extractAnnouncement(title, text) facts
    "class": "NEW_BUILD", "mayCreateProject": true, "headlineMw": 12, "status": "announced",
    "operator": "Applied Digital", "country": "RO", "city": "Bucharest", "money": { "amount": 7500000000, "currency": "USD" },
    "_todo": { "investmentUsd": null, "_reason": "…" }
  }
}

Announcement facts available: title, published, class, mayCreateProject, physical, headlineMw, mwAll, status, operator, operators, country, city, region, money, investmentUsd, acreage, phaseCount, expectedOpening, projectName, explicitName, relevance, leadRelevance, titleSignal, nonProjectTitle, pageType, eventType.

Every fixture always asserts zero validation errors. Expected values must be what the parser produces today and what the source page says: verify each figure against the body before writing it. When the parser is wrong, put the correct value under _todo with a _reason — the test runs as it.fails and starts failing the day the parser is fixed, which is the signal to promote the field. Never assert a value you believe is wrong.

# 3.4 Running

bash
pnpm --filter @dci/worker test                                        # everything (fixtures included)
pnpm --filter @dci/worker exec vitest run src/connectors/fixtures.test.ts -t databank   # one connector

Golden coverage on 2026-09-12: 21 fixtures over 18 connectors (equinix, digitalrealty, stack, databank ×2, ntt, edgeconnex ×2, cyrusone, qts ×2, vantage, coresite, airtrunk, nextdc ×2, atnorth, cloudhq, datacenterfrontier, datacenterdynamics, loudoun-county), no open _todo (the nine gaps recorded on 2026-09-12 — EdgeConneX city from the h1, NEXTDC 20MW+, QTS figure-first capacity and template-copied map pin, DCD teaser operator and portfolio money — were fixed the same day and promoted). No planning PDF fixture yet: the archive holds no document with a PDF content type.

# 4. Live dry-run

bash
pnpm tsx scripts/try-connector.ts config/connectors/<id>.yaml [--limit 5] [--url <single url>] [--verbose] [--json out.json]
pnpm dci run <id> --dry-run --limit 5 [--url u]      # same through the worker CLI (needs the DB env)
pnpm dci trace <doc-id-or-url> [--live]              # one document end to end

try-connector registers every apps/worker/src/connectors/<group>/index.ts, builds the connector from the YAML (implementation or GenericConnector), runs discover (or the single --url), fetches with the real robots / rate-limit / escalation policy through createTestContext (in-memory state, nothing persisted, dryRun: true) and prints records, entities and the ValidationReport. Nothing ships without a live dry-run yielding valid entities and 0 rejected; record the result in docs/connectors/<id>.md ("Verification (live, )").

# 5. Quarantine

Two different things share the word:

  • Connector quarantine (connectors.quarantine, pnpm dci quarantine <id> [on|off]): the run does everything — discover, fetch, archive, extract, validate — but the ingest transaction is rolled back, so nothing is published; the run is stored with quarantined = true for review. runConnector (apps/worker/src/pipeline.ts) reads the flag before each run. It is switched on automatically when a connector fails or is blocked 5 runs in a row (consecutive_failures ≥ 5) or when a non-full run creates ≥ 500 records (suspiciousSpike): both are logged as … — quarantined pending review. New connectors are worth running quarantined for the first full pass.
  • Document quarantine (documents.quarantined): nextCheckAfter (apps/worker/src/scheduling.ts) backs a 404/410 off ×4 and after 3 consecutive gone responses (MAX_CONSECUTIVE_GONE) sets nextCheck = null and quarantined = true — the URL is never scheduled again. Other errors back off ×2^(n−1) (cap ×8); premium retries after two failures are limited to every 4th attempt (premiumAllowedAfterErrors). Weekly check in docs/CRAWL-OPERATIONS.md: select count(*) from documents where quarantined should grow slowly.

# 6. Parser versioning and re-extraction

The effective extractor version stored on every document (documents.extractor_version) is computed by effectiveExtractorVersion(cfg, connector) in apps/worker/src/configs.ts:

text
<cfg.parserVersion>+<sha256("news_article_v1@news_v3|databank_facility_v1@v2|impl:peeringdb@v1")[0:10]>

i.e. the YAML parserVersion plus a hash of every referenced Parser.name@Parser.version (and impl:<name>@<connector.parserVersion>). Consequences:

  • Bumping version on the Parser object is what invalidates cached extractions. shouldSkipExtraction (scheduling.ts) skips a document only when its content fingerprint is unchanged and the stored extractor version equals the current one (and the run is not --force). Change a regex without bumping the version and unchanged pages keep their old records until their content changes.
  • Re-run the archive without network: pnpm dci reprocess <id> [--limit n] [--group g] --stale re-extracts from stored bodies (MinIO) only the documents whose extractor_version differs from the current one; without --stale everything is re-extracted. pnpm dci run <id> --task reprocess is the same thing.
  • Declarative connectors have no parser: bump parserVersion in the YAML.

# 7. Health states (connectorHealthFrom, apps/worker/src/scheduling.ts)

connectors.health after each run (admin health center; docs/CRAWL-OPERATIONS.md § "Reading a bad run"):

State Rule
ok run status not failed, fetch failure rate < 20 % (healthFrom)
degraded failure rate 20–50 %
failing run status = failed, or failure rate ≥ 50 %
blocked ≥ 50 % of fetched pages came back blocked (403/429/CAPTCHA/JS shell) — fix the fetch level / user agent, never bypass
schema_change ≥ 5 pages fetched fine and passed to extractors but 0 entities on a connector that has documents — the site's HTML changed: re-check selectors against a fresh fixture
no_new_content discovery yielded 0 URLs and 0 fetches although a previous run discovered some — sitemap/RSS moved or emptied
never_run no run yet (also the state of enabled: false connectors)
  • fetch.respectRobots: true — createTestContext and the runtime check robots.txt before every fetch and honour crawl-delay; a disallowed URL yields error.code = robots_disallow, not a fetch.
  • No authentication, paywall or CAPTCHA bypass; escalation (L3/L4) is for JS shells and bot walls on public pages only, within the credit budgets.
  • Conditional GET (ETag / Last-Modified) and content hashing: unchanged pages are never re-extracted, premium credits are never spent re-fetching unchanged content.
  • Third-party publishers (kind: news) keep only title, date, link, extracted facts and a summary of at most THIRD_PARTY_SUMMARY_MAX = 400 characters (trimSummary, packages/connectors/src/generic.ts); the article text is never stored (news_article_v1 keepText defaults to false; the fixture test enforces both when thirdParty: true). Only public-sector sources (kind: government, utility, filing) may set keepText: true.
  • license / attribution are mandatory in the YAML and are shown with every record (sources.license, sources.attribution).
  • Never invent data: no geocoding, no guessed capacities, no default coordinates. City-level coordinates are geoPrecision: city, operator-published markers exact / parcel. An estimate is flagged isEstimate.

# 9. Checklist for a new connector

  1. config/connectors/<id>.yaml — id, name, domain, kind, priority, license/attribution, fetch, discovery (sitemap first, then include/exclude/classify), schedule, extractors, defaults. pnpm dci validate.
  2. Declarative fields when the page is regular; otherwise a parser in apps/worker/src/connectors/<group>/<file>.ts registered in <group>/index.ts register() (facilityParser + Rec for facilities), or an implementation for APIs / datasets.
  3. Live dry-run: pnpm tsx scripts/try-connector.ts config/connectors/<id>.yaml --limit 5 → valid entities, 0 rejected, 0 unexpected premium credits.
  4. Fixture: archive a real body under apps/worker/fixtures/<id>/, write expected.json from what the parser produces after checking each value against the page, pnpm --filter @dci/worker test.
  5. docs/connectors/<id>.md — what it collects, discovery, schedule & fetch, quirks, verification log.
  6. First production run quarantined (pnpm dci quarantine <id> on, run, review connector_runs, then off).
  7. Any later change to a parser's output → bump Parser.version → pnpm dci reprocess <id> --stale.