ODATA (odata)
- Domain:
odatacolocation.com· Kind: operator · Priority: 2 · Enabled: yes - License / attribution: Publicly accessible operator pages; factual data only; © ODATA — attribution “ODATA”
- Coverage: BR, MX, CL, CO
What it collects
Facility records (kind: facility) via parser(s) odata_facility_v1 in apps/worker/src/connectors/operators2/americas.ts.
Fields published by the pages and captured: name, code, city, countryIso2, lat/lng (exact, Google Maps link), itCapacityMw (IT power), buildingSqm (floor area), tier, certifications.
Every field carries an extraction method in methods (e.g. json-ld:PostalAddress, regex:it_mw_v1, attr:data-value+MW, lookup:city-country) for provenance. Coordinates are only taken from JSON-LD / embedded JSON / map links published by the operator — never geocoded. Status defaults to operational unless the page uses pipeline wording (under construction / coming soon / planned).
Discovery
- Sitemaps:
https://odatacolocation.com/data-center-sitemap.xml,https://odatacolocation.com/imprensa-sitemap.xml - Include:
/en/blog/data-center/dc-[a-z]{2}\d{2}(-\d)?/$,/en/press,/en/news - Exclude:
\?,# - Classify:
/en/blog/data-center/dc-[a-z]{2}\d{2}(-\d)?/$→ facility_pages (facility_page);/en/(press|news)→ newsroom (press_release)
Newsroom URLs (when configured) rely on the GenericConnector news fallback (press_release → news_event when data-center relevant).
Schedule & fetch
- Schedule: facility_pages weekly, newsroom daily, sitemap weekly
- Fetch: L1→L2, 15 rpm, concurrency 2, robots.txt respected, max 20 premium credits/run. All pages are served at L1 (plain HTTP);
maxLevelis capped at 2 so false-positive “blocked” heuristics (HubSpot / cookie-consent strings) never spend Scrapfly credits.
Quirks
English facility pages /en/blog/data-center/dc-/ publish IT power (MW), floor area, Tier III and a Google Maps link with coordinates. Portuguese/Spanish duplicates are excluded. Country is only set when the page names it (BR/MX/CL/CO).
Verification (live, 2026-09-11)
try-connector --limit N: discovered 12 URLs (facility_pages=12); sampled pages → 6 entities, 6 valid, 0 rejected, 0 premium credits.