Connectors — families, SDK, fixtures, verified vendor endpoints
A connector is a strategy for one kind of public source; a sensor binds a connector to one URL for one company
(spec §8–9, §105–106). All connectors live in src/companyatlas/connectors/, register themselves with the SDK
(sdk/connector.py) and are auto-loaded (connectors.load_all()); catlas connectors lists them and upserts the connectors table.
Families (as of connector-v1 ids)
| id | family | picks when | output |
|---|---|---|---|
generic-html-v1 |
any HTML surface | fallback for every surface (surfaces=("*",)) |
text + semantic blocks + surface-typed lists: people (leadership/about), plans (pricing), locations (locations/contact), news (newsroom/blog/changelog/research/IR), jobs (careers, list/table/card + JSON-LD), products (products/services/solutions); discovered links on every page |
sitemap-v1 |
sitemap*.xml(.gz), sitemap_index.xml, /sitemaps/ |
URL pattern | URL set (bounded by discovery_max_sitemap_urls), classified DiscoveredUrls, one block per URL (diff = new/gone URLs) |
feed-v1 |
RSS / Atom / JSON Feed | URL pattern (/feed, /rss, .xml, feed.json) or surface feed |
news items (title, url, published_at, summary, language) |
jsonld-jobs-v1 |
JobPosting JSON-LD / microdata | non-ATS jobs_board URLs (or explicitly by discovery) |
jobs |
greenhouse-v1 lever-v1 ashby-v1 smartrecruiters-v1 workable-v1 workday-v1 recruitee-v1 personio-v1 teamtailor-v1 |
structured job boards | API URL pattern (set by discovery via connectors/_util.ats_sensor_spec) |
jobs, one stable block per job |
statuspage-v1 |
Atlassian Statuspage | /api/v2/summary.json |
components / incidents as blocks + incident news |
Connector selection: connector.for_surface(surface, url) scores every registered connector (url_pattern match +100, category
+20, listed surfaces +10, "*" +5, plus priority) and returns the best; the generic HTML connector is the floor.
Verified vendor endpoints (2026-09-12, one real board each, then saved as trimmed fixtures — tests never hit the network)
| vendor | endpoint | verified board | fixture |
|---|---|---|---|
| Greenhouse | GET https://boards-api.greenhouse.io/v1/boards/{token}/jobs?content=false (+/departments, /offices optional) |
stripe |
greenhouse/stripe_jobs.json |
| Lever | GET https://api.lever.co/v0/postings/{token}?mode=json (api.eu.lever.co for EU boards) |
palantir |
lever/palantir_postings.json |
| Ashby | GET https://api.ashbyhq.com/posting-api/job-board/{token} |
ashby |
ashby/ashby_board.json |
| SmartRecruiters | GET https://api.smartrecruiters.com/v1/companies/{token}/postings?limit=100&offset=N (offset pagination, ≤ 10 pages) |
smartrecruiters |
smartrecruiters/smartrecruiters_postings.json |
| Workable | GET https://apply.workable.com/api/v1/widget/accounts/{token} |
epignosis |
workable/epignosis_widget.json |
| Workday | POST https://{tenant}.{wdN}.myworkdayjobs.com/wday/cxs/{tenant}/{site}/jobs body {"appliedFacets":{},"limit":20,"offset":N,"searchText":""} (bounded 400 jobs) — uses Fetcher.post_json (SSRF-checked, governor) |
nvidia.wd5 / NVIDIAExternalCareerSite |
workday/nvidia_jobs_page1.json |
| Recruitee | GET https://{token}.recruitee.com/api/offers/ |
vandebron |
recruitee/vandebron_offers.json |
| Personio | GET https://{token}.jobs.personio.de/xml (<workzag-jobs>) |
personio |
personio/personio_jobs.xml |
| Teamtailor | GET https://{career-site-host}/jobs.json (JSON Feed 1.1 with _jobposting) — {token}.teamtailor.com boards that 404 keep the HTML careers sensor |
career.teamtailor.com |
teamtailor/teamtailor_jobs_feed.json |
| Statuspage | GET https://status.<host>/api/v2/summary.json |
githubstatus.com |
statuspage/github_summary.json |
All of these are the endpoints the vendors' own public pages call (spec §13 mode C). Nothing needs a key; nothing bypasses a
challenge. Every request still goes through fetch.Fetcher (SSRF guard, robots.txt, per-domain governor, size caps).
Job normalisation (every board → ExtractedJob)
connectors/_util.finish_job fills derived fields without overriding vendor data: parse_location (city / region / ISO-2
country only when stated — "San Francisco, CA" keeps region CA and no country because CA is also Canada), remote /
hybrid, seniority_guess, employment_type_norm. The pipeline sets is_ai (taxonomy.AI_KEYWORDS) and is_engineering.
Job identity = job_fingerprint(title, location_text, external_id | url); structured boards also give each job a stable block
key so the block diff mirrors the job delta.
Precision rules (connectors/_precision.py, precision-v1, 2026-09-13)
Production samples showed the generic HTML connector capturing navigation / call-to-action anchors as jobs (Meer info, Early careers bij …), cookie-consent categories as locations (Performance & Analytics | Allows use of behavioural data…), swapped person cards
(Chair Emeritus | Warner Bros. Discovery), marketing headings as plans (Win your market with … | Talk to sales) and truncated
price_text. The precision module holds deterministic, multilingual (EN/FR/DE/NL/ES/IT/PT/JA) validators and normalisers that
generic_html.py applies while extracting and services/pipeline._drop_corrupt_entities re-applies as the last line of defence for
every connector (job rules only for generic-html-v1 — structured boards are trusted). Every rule returns a Verdict(ok, reason), so
scripts/audit_extractions.py can explain rejections on stored rows. Limits are module constants (no magic numbers).
| entity | rejected | required | normalised |
|---|---|---|---|
| jobs (HTML) | CTA / nav stoplist (read more, apply now, meer info, en savoir plus, mehr erfahren, alle stellen, ver más, scopri di più, 詳細を見る …), category prefixes (early careers, careers at …, why …), single-word titles, titles ending with an ellipsis, prose-like titles (≥ 4 words with function words) without role vocabulary, > 140 chars |
role vocabulary (engineer, manager, analyst, nurse, driver, ingénieur, entwickler, verpleegkundige, ingeniero, エンジニア … or a format marker (m/w/d), Vollzeit), or a job-like URL (/job/, /jobs/<slug>, /careers/x/y, /vacature, /stelle, /offre, /emploi, /empleo, /lavoro, numeric id, ATS hosts), or an explicit location / department cell |
(m/w/d), (h/f), (all genders), - Apply now and bracketed ids stripped from the display title; the id becomes external_id (so it stays in the fingerprint). Bare known cities (München) count as a location cell |
| people | names that are role titles or generic labels (Contact, Team, Board of Directors, Leadership, お問い合わせ), digits, < 2 or > 5 tokens |
name = 2–5 capitalised tokens (particles / initials allowed) or a short CJK name without role words | name ↔ title swapped when the card put the title first; link-label titles (Contact, Read more, Bio, LinkedIn, プロフィール) and bio sentences → title=None; role_category / is_executive recomputed after the swap (JA titles: 代表取締役社長 → ceo, 社外取締役 → board …) |
| locations | cookie / consent / nav vocabulary (performance, analytics, marketing, necessary, preferences, cookies, contact us, careers, privacy, newsletter, datenschutz, voorkeuren …), names > 80 chars or > 8 words, sentences |
at least one of: ISO country, known city (≈ 400 major cities + exonyms, n-gram match), postal code (UK/NL/CA/numeric), street address (EN/FR/DE/NL/ES/IT patterns), explicit kind (headquarters, factory, warehouse, lab, store … or a kind word in the name) |
city cleared unless it is a plausible city (Level 12 is not); `Oman |
| pricing plans | names > 40 chars or > 5 words, terminal punctuation, ?, ` |
, eyebrows (Most popular, New), generic headings (Plans, Pricing, Tarifs, Preise), marketing verbs (win, get, discover, unlock, découvrez, entdecken…), plans without a parsed price /contact_sales` / free tier |
— |
| products | nav words (Overview, Products, Solutions, All products, Learn more, Produits, 製品一覧), > 80 chars / > 10 words, sentences and slogans |
— | ® / ™ kept |
| news | pagination (Next, Older posts », Page 2, 3), category labels (Press releases, Blog posts, Actualités, ニュース), Read more |
≥ 3 words or ≥ 15 chars unless published_at or a /YYYY/MM/ URL confirms an article |
the title is the first non-date line of the card |
Normalizer changes made for these rules (sdk/normalize.py): a .job-list-style wrapper whose only child is a linked / card-like
<ul> or <table> is a container (its rows become job_listing blocks, inheriting the wrapper's kind) while a feature <ul> inside a
pricing card is not; card text has one line per direct child so inline <a>title</a><span>city</span> cells do not merge. Stripping
gender markers changes the display title (and therefore the fingerprint) of affected HTML jobs once — expect a one-off
removed + added pair for (m/w/d)-style listings on the first crawl after deployment.
Fixtures reproducing the production noise: generic_html/careers_cta_noise.html (NL careers page: CTA anchors, single word, ellipsis
- 5 real listings incl.
(m/v/x),[12933],- Apply now),careers_de.html(German(m/w/d)listings +Alle Stellen/Mehr erfahren),leadership_swapped.html(title-first card,Contactlink,Chair Emeritus | Warner Bros. Discovery,Board of Directors),leadership_ja.html(Japanese officers page),locations_cookie_noise.html(consent categories inside a locations page,Contact Us | KR,Oman | Hormuz Grand Hotel, Dutch address),pricing_marketing_noise.html(marketing hero +Talk to sales,Worry-free roaming.+ prose price, eyebrow, real Free / Starter / Team / Business / Enterprise),products_news_noise.html. Tests:tests/test_extraction_precision.py(noise gone, every real item still extracted). Audit / purge of stored rows:scripts/audit_extractions.py— samples HTML-sourced rows per table, prints a precision estimate (Wilson 95 % interval) and the rejections by reason;--all-sourceswidens the read-only audit to enrichment rows (Wikidata people);--purgedeletes failing HTML-sourced rows of people / products / pricing_plans / locations / jobs only (never ATS jobs, enrichment rows or news). Run it before and after deploying to measure the effect on production.
Adding a connector family
- Create
src/companyatlas/connectors/<family>.py; subclassConnector(orAtsConnectorfor JSON/XML boards — implementparse_jobs(data, sensor)only) and decorate with@register. DeclareConnectorMeta(connector_id="<family>-v1", category=Surface…, fetch_mode, url_pattern, priority, accept, default_interval_s). Overridefetch()only for pagination / POST; never open sockets — use theFetcheryou receive (get,request,post_json). - Verify the endpoint shape against one real public source (a handful of polite requests), then save the response trimmed to
~20 items under
fixtures/connectors/<family>/(strip descriptions / personal data). - Add tests in
tests/test_connectors_*.pythat load the fixture withfetch.file_resultand assert counts, ids, urls, locations. Tests must not touch the network (-m liveis opt-in only). - If discovery should create sensors for it, extend
connectors/_util.ats_sensor_spec(vendor → API URL + connector id + config) and the detection regexes inurls.ATS_PATTERNS/services/discovery.EMBED_ATS_RE. - Behaviour change that alters extraction materially → new id (
<family>-v2). Observations and snapshots store the id asconnector_version, so history stays reproducible;connector.get("<family>-v1")keeps resolving to the newest of the family.
Fixture policy
- Hand-written HTML for each generic surface lives in
fixtures/connectors/generic_html/(pricing.html+pricing_v2.html,careers.html+careers_v2.html,leadership.html,locations.html,newsroom.html,homepage.html,legal_terms.html,careers_jsonld.html);sitemap/andfeed/hold XML samples. Pairs (*_v2) exist so the diff engine and the pipeline can be tested for real change semantics (price change, jobs added/removed). - Vendor fixtures are captured responses, trimmed; keep the top-level shape intact (pagination fields included).
- Regression rule (spec §148): when changing the normalizer, the diff engine or an extractor, run the suite — the fixtures are the historical behaviour contract.