SPB Git forge
38commits 1branches 0releases
338.7 MBsize
maindefault branch
3 h agolast push
HTML 53.9% TypeScript 44.5% JavaScript 0.6% SQL 0.5%
2.3 KB

# CloudHQ (cloudhq)

  • Domain: cloudhq.com · Kind: operator · Priority: 2 · Enabled: yes
  • License / attribution: Publicly accessible operator pages; factual data only; © CloudHQ — attribution “CloudHQ”
  • Coverage: US, GB, NL, IT, DE, FR, JP, TH, BR, MX, PY

# What it collects

Facility records (kind: facility) via parser(s) cloudhq_campus_v1 in apps/worker/src/connectors/operators2/americas.ts.

Fields published by the pages and captured: name, code, lat/lng (parcel, published coordinate line), itCapacityMw (total critical IT load), buildingSqm, status (RFS timing), countryIso2/city when named.

Every field carries an extraction method in methods (e.g. json-ld:PostalAddress, regex:it_mw_v1, attr:data-value+MW, lookup:city-country) for provenance. Coordinates are only taken from JSON-LD / embedded JSON / map links published by the operator — never geocoded. Status defaults to operational unless the page uses pipeline wording (under construction / coming soon / planned).

# Discovery

  • Sitemaps: https://cloudhq.com/sitemaps.xml
  • Seeds: /campuses/, /campuses/?sf_paged=2, /campuses/?sf_paged=3
  • Link following: /campus/[a-z0-9-]+/$
  • Include: /campus/[a-z0-9-]+/$, /news/, /press/
  • Exclude: \?, #
  • Classify: /campus/[a-z0-9-]+/$ → facility_pages (facility_page); /(news|press)/. → newsroom (press_release)

Newsroom URLs (when configured) rely on the GenericConnector news fallback (press_release → news_event when data-center relevant).

# Schedule & fetch

  • Schedule: facility_pages weekly, newsroom daily, sitemap weekly
  • Fetch: L1→L2, 15 rpm, concurrency 2, robots.txt respected, max 20 premium credits/run. All pages are served at L1 (plain HTTP); maxLevel is capped at 2 so false-positive “blocked” heuristics (HubSpot / cookie-consent strings) never spend Scrapfly credits.

# Quirks

The campus post type is absent from sitemaps.xml, so campus pages are discovered by following links from /campuses/ (3 pages). Campus pages /campus// publish a coordinate line (lat, lng), total critical IT load (MW), square feet and RFS timing (→ under_construction when RFS is in the future).

# Verification (live, 2026-09-11)

try-connector --limit N: discovered 23 URLs (seed=3, facility_pages=20); sampled pages → 1 entities, 1 valid, 0 rejected, 0 premium credits.