# CloudHQ (`cloudhq`) - **Domain**: `cloudhq.com` · **Kind**: operator · **Priority**: 2 · **Enabled**: yes - **License / attribution**: Publicly accessible operator pages; factual data only; © CloudHQ — attribution “CloudHQ” - **Coverage**: US, GB, NL, IT, DE, FR, JP, TH, BR, MX, PY ## What it collects Facility records (`kind: facility`) via parser(s) `cloudhq_campus_v1` in `apps/worker/src/connectors/operators2/americas.ts`. Fields published by the pages and captured: name, code, lat/lng (parcel, published coordinate line), itCapacityMw (total critical IT load), buildingSqm, status (RFS timing), countryIso2/city when named. Every field carries an extraction method in `methods` (e.g. `json-ld:PostalAddress`, `regex:it_mw_v1`, `attr:data-value+MW`, `lookup:city-country`) for provenance. Coordinates are only taken from JSON-LD / embedded JSON / map links published by the operator — never geocoded. Status defaults to `operational` unless the page uses pipeline wording (under construction / coming soon / planned). ## Discovery - Sitemaps: `https://cloudhq.com/sitemaps.xml` - Seeds: `/campuses/`, `/campuses/?sf_paged=2`, `/campuses/?sf_paged=3` - Link following: `/campus/[a-z0-9-]+/$` - Include: `/campus/[a-z0-9-]+/$`, `/news/`, `/press/` - Exclude: `\?`, `#` - Classify: `/campus/[a-z0-9-]+/$` → facility_pages (facility_page); `/(news|press)/.` → newsroom (press_release) Newsroom URLs (when configured) rely on the GenericConnector news fallback (press_release → `news_event` when data-center relevant). ## Schedule & fetch - Schedule: facility_pages weekly, newsroom daily, sitemap weekly - Fetch: L1→L2, 15 rpm, concurrency 2, robots.txt respected, max 20 premium credits/run. All pages are served at L1 (plain HTTP); `maxLevel` is capped at 2 so false-positive “blocked” heuristics (HubSpot / cookie-consent strings) never spend Scrapfly credits. ## Quirks The `campus` post type is absent from sitemaps.xml, so campus pages are discovered by following links from /campuses/ (3 pages). Campus pages /campus// publish a coordinate line (lat, lng), total critical IT load (MW), square feet and RFS timing (→ under_construction when RFS is in the future). ## Verification (live, 2026-09-11) `try-connector --limit N`: discovered 23 URLs (seed=3, facility_pages=20); sampled pages → 1 entities, 1 valid, 0 rejected, 0 premium credits.