# STACK Infrastructure — campus pages (campus + building records) + news/press feed id: stack name: STACK Infrastructure domain: www.stackinfra.com kind: operator mode: hybrid priority: 1 license: "Publicly accessible operator pages; factual data only; © STACK Infrastructure" attribution: "STACK Infrastructure" homepage: https://www.stackinfra.com/locations/ notes: > Open WordPress site. Yoast sitemaps: campuses-sitemap.xml (/locations////), location-sitemap.xml (metro indexes), news-sitemap.xml; RSS at /feed/. Campus pages publish campus MW + acreage and a card per building (code, SQ FT, MW). No street address or coordinates. The parser emits one campus record plus one record per building (campusName = STACK ). Japanese pages under /japan/ are excluded (duplicates of KIX01/TKY01). robots.txt has no disallows. maxLevel is capped at 2 (direct HTTP only): every stackinfra.com page embeds Cloudflare's "/cdn-cgi/challenge-platform/" beacon, which the shared looksBlockedOrEmpty() heuristic misreads as a bot wall and would escalate each of the ~50 campus pages to Scrapfly (30 credits each) although direct fetches return the full page (HTTP 200, verified). coverage: { global: true, operators: ["STACK Infrastructure"] } fetch: level: 1 maxLevel: 2 rpm: 20 concurrency: 2 respectRobots: true maxCreditsPerRun: 40 discovery: sitemap: ["/campuses-sitemap.xml", "/news-sitemap.xml"] rss: ["/feed/"] include: ["/locations/", "/about/news-press/"] exclude: ["\\?", "/japan/", "/take-a-tour", "/about/news-press/?$"] classify: - { pattern: "/locations/(americas|emea|asia-pacific)/[a-z-]+/[a-z0-9-]+/?$", group: facility_pages, pageType: facility_page, priority: 60 } - { pattern: "/locations/(americas|emea|asia-pacific)/[a-z-]+/?$", group: index, pageType: facility_index, priority: 20 } - { pattern: "/locations/?$", group: index, pageType: facility_index, priority: 10 } - { pattern: "/about/news-press/(press-releases|news)/[a-z0-9._%-]+/?$", group: newsroom, pageType: press_release, priority: 70 } maxUrlsPerRun: 1000 schedule: facility_pages: weekly newsroom: daily index: monthly sitemap: weekly rss: daily extractors: campus: # selected by URL only (classifyPage() may mislabel deep facility URLs; the URL is authoritative) match: ["/locations/(americas|emea|asia-pacific)/[a-z-]+/[a-z0-9-]+/?$"] kind: facility parser: stack_campus_v1 defaults: operatorName: STACK Infrastructure operatorKind: developer facilityType: wholesale isHyperscale: true parserVersion: v1