SPB Git forge
38commits 1branches 0releases
338.7 MBsize
maindefault branch
3 h agolast push
HTML 53.9% TypeScript 44.5% JavaScript 0.6% SQL 0.5%
2.4 KB · 23 lines markdown
Rendered Raw Blame History
1# ironmountain — Iron Mountain Data Centers (DISABLED)23| | |4|---|---|5| Config | `config/connectors/ironmountain.yaml` (`enabled: false`) |6| Parser | none shipped |7| Kind / priority | operator / 1 |8| License / attribution | Publicly accessible operator pages; factual data only; © Iron Mountain Incorporated |910## Why it is disabled (2026-09-11)11- Every direct request to `www.ironmountain.com` — `robots.txt`, `/sitemap.xml`, `/data-centers`, `/data-centers/global-data-center-locations`, with the bot UA **and** with a browser UA — returns HTTP 429 with a "Vercel Security Checkpoint" JavaScript challenge. This is a hard bot wall, not a rate limit we can wait out.12- Scrapfly with ASP does pass (HTTP 200), but each request is billed **30 credits** (residential proxy 25 + browser 5, even with `render_js=false`). Discovery alone needs robots.txt + the sitemap index + 9 sitemap shards (`sitemap1.xml … sitemap9.xml`) ≈ 330 credits before the first facility page, i.e. more than the whole 150-credit run budget, and every weekly recrawl would cost the same again. That violates the "premium only when blocked, never for bulk crawling" rule and the daily budget.13- Two probe requests were spent to establish this (60 credits): robots.txt and the sitemap index. robots.txt (fetched through Scrapfly) allows `/data-centers/`, disallows Sitecore/system folders, `/zh-cn/`, `?localize=` and some landing pages, and blocks the `Scrapy` and `BoogleBot2` agents.1415## What is in the YAML16Discovery/classify rules are drafted from the public URL scheme (`/data-centers/…`, newsroom) but **no extractor** is declared — the facility page template could not be inspected without spending credits, and shipping unverified extraction is not allowed. The connector is `enabled: false` so the scheduler ignores it.1718## How to re-enable191. Check whether the challenge has been lifted: `curl -A "DataCenterIndexBot/0.1" https://www.ironmountain.com/robots.txt` should return 200 text.202. Inspect 2–3 facility pages, confirm the URL pattern (`/data-centers/<region>/<facility>` is an assumption) and write `ironmountain_facility_v1` in `apps/worker/src/connectors/operators1/`.213. Verify with `try-connector` (`--limit 4`, spot-check two sites), then set `enabled: true`.22If the wall stays, an alternative is Iron Mountain's PeeringDB / registry presence via the registry connectors rather than the operator site.23