SPB Git forge
38commits 1branches 0releases
338.7 MBsize
maindefault branch
6 h agolast push
HTML 53.9% TypeScript 44.5% JavaScript 0.6% SQL 0.5%
8.9 KB · 130 lines tsx
Raw Blame History
1import type { Metadata } from "next";2import Link from "next/link";3import type { ReactNode } from "react";4import { Breadcrumbs } from "@/components/entity/breadcrumbs";5import { SectionNav } from "@/components/entity/section-nav";6import { Page, PageHeader, Section } from "@/components/ui/card";7import { TBody, THead, Table, Td, Th, Tr } from "@/components/ui/table";8import { routes } from "@/lib/routes";9import { pageMetadata } from "@/lib/seo";1011export const metadata: Metadata = pageMetadata({12  title: "About our crawler",13  description: "DataCenterIndexBot indexes public data center facility pages, press releases, planning filings and cloud region lists. How it identifies itself, how it behaves and how to opt out.",14  path: routes.bot(),15});1617/** Default identity of L1 fetches — packages/connectors/src/fetchers.ts (DEFAULT_UA, overridable via DCI_USER_AGENT). */18const UA = "DataCenterIndexBot/0.1 (+https://www.datacenterindex.io/bot; contact@spboucher.ai)";1920const NAV = [21  { id: "what", label: "What it is" },22  { id: "behaviour", label: "Behaviour" },23  { id: "identification", label: "Identification" },24  { id: "opt-out", label: "Opt out" },25  { id: "contact", label: "Contact" },26];2728const FREQUENCY: Array<[string, string]> = [29  ["Facility / site pages", "about weekly"],30  ["Newsroom, press releases, RSS", "daily"],31  ["Cloud provider region lists", "daily to weekly"],32  ["Index / listing pages and sitemaps", "weekly to monthly"],33  ["Planning filings, registries, datasets", "weekly to monthly, per the source's own cadence"],34];3536function P({ children }: { children: ReactNode }) {37  return <p className="mt-3 first:mt-0">{children}</p>;38}3940const linkCls = "text-ink underline decoration-hair-2 hover:text-accent";4142export default function BotPage() {43  return (44    <Page className="max-w-[880px]">45      <Breadcrumbs items={[{ name: "Crawler", href: routes.bot() }]} className="pt-3" />46      <PageHeader eyebrow="DataCenterIndexBot" title="About our crawler" description="If you found this page in your server logs: the requests come from the crawler that builds DataCenterIndex. Here is what it reads, how it behaves and how to make it stop." className="pt-3" />47      <SectionNav items={NAV} />4849      <div className="mt-6 flex flex-col gap-10 text-[14.5px] leading-relaxed text-ink-2">50        <Section id="what" className="scroll-mt-28" title="What it is">51          <P>52            DataCenterIndexBot collects <strong className="text-ink">public</strong> information about data center infrastructure: operators&rsquo; facility and campus pages, press releases and newsrooms, cloud providers&rsquo; published region and availability-zone lists, planning and permitting filings, utility and regulator publications, and open datasets and registries such as PeeringDB, OpenStreetMap, Wikidata and the World Bank.53          </P>54          <P>55            From these pages it extracts <em>facts</em> — a facility name, an address, an IT capacity in MW, an opening date, a status, a certification — and records where each fact was read, when, and with what confidence. The result is the structured, versioned graph published on this site and through the public API, with every value linked back to the page it came from and an attribution line for every source (see the <Link href={routes.sources()} className={linkCls}>source registry</Link>). It does not copy articles or reproduce pages; descriptions are short and attributed.56          </P>57        </Section>5859        <Section id="behaviour" className="scroll-mt-28" title="How it behaves">60          <ul className="list-disc space-y-1.5 pl-5">61            <li>62              <strong className="text-ink">robots.txt</strong> is fetched and honoured for every host, including <code className="figure text-[12.5px]">Disallow</code> rules and <code className="figure text-[12.5px]">Crawl-delay</code>; the file is re-read at least once a day. Both the group for <code className="figure text-[12.5px]">DataCenterIndexBot</code> and the <code className="figure text-[12.5px]">*</code> group are respected.63            </li>64            <li>65              <strong className="text-ink">Rate limits per host</strong>: by default at most 30 requests per minute and 2 concurrent connections, lower when the source asks for it or when responses slow down. Sources with published API limits (for example SEC EDGAR fair-access rules or the Wikidata query service) are read within those limits.66            </li>67            <li>68              <strong className="text-ink">Conditional requests</strong>: pages are requested with <code className="figure text-[12.5px]">If-None-Match</code> / <code className="figure text-[12.5px]">If-Modified-Since</code> when the server supports them, and every body is content-hashed so an unchanged page is never re-processed.69            </li>70            <li>71              <strong className="text-ink">Public pages only</strong>: no login, paywall, CAPTCHA or bot-wall is bypassed. When a page is only served as a JavaScript shell, a rendering service may be used to read it — the same page a visitor would see, at the same rate limits.72            </li>73            <li>74              <strong className="text-ink">Scoped discovery</strong>: only page types relevant to the source (facility pages, newsroom, region lists, filings) are followed; the crawler does not wander through an entire site.75            </li>76          </ul>77          <P>Typical revisit frequency by page type — actual schedules are set per source and are usually less frequent:</P>78          <div className="mt-3">79            <Table dense>80              <THead>81                <tr>82                  <Th>Page type</Th>83                  <Th>Revisit</Th>84                </tr>85              </THead>86              <TBody>87                {FREQUENCY.map(([t, f]) => (88                  <Tr key={t}>89                    <Td primary>{t}</Td>90                    <Td label="Revisit" className="text-ink-2">91                      {f}92                    </Td>93                  </Tr>94                ))}95              </TBody>96            </Table>97          </div>98        </Section>99100        <Section id="identification" className="scroll-mt-28" title="Identification">101          <P>Direct requests carry this User-Agent header:</P>102          <pre className="mt-3 overflow-x-auto rounded-sm border border-hair bg-surface px-3 py-2.5 text-[12.5px] leading-relaxed text-ink">103            <code className="figure">{UA}</code>104          </pre>105          <P>106            Requests originate from the infrastructure that hosts the index — the MacLustr cluster and dedicated servers at OVHcloud (Canada and France). Reverse DNS on those addresses is not guaranteed to resolve to a datacenterindex.io name, so please rely on the User-Agent string above, or write to us to confirm a specific address. When a source blocks the bot identity and a browser identity is used instead, the same rate limits apply and the source is listed in the <Link href={routes.sources()} className={linkCls}>registry</Link>.107          </P>108        </Section>109110        <Section id="opt-out" className="scroll-mt-28" title="How to opt out">111          <P>Add a group for the bot to your robots.txt. The change takes effect at the next robots.txt refresh (within 24 hours); pages already archived are not fetched again.</P>112          <pre className="mt-3 overflow-x-auto rounded-sm border border-hair bg-surface px-3 py-2.5 text-[12.5px] leading-relaxed text-ink">113            <code className="figure">{"User-agent: DataCenterIndexBot\nDisallow: /"}</code>114          </pre>115          <P>To exclude only part of a site, list the paths instead of <code className="figure text-[12.5px]">/</code>. A <code className="figure text-[12.5px]">Crawl-delay</code> directive slows the bot down without blocking it.</P>116          <P>117            You can also email <a href="mailto:contact@spboucher.ai" className={linkCls}>contact@spboucher.ai</a> to have a source removed from the index or to correct a value. Removal requests are honoured within 7 days: the connector is paused and archived documents are deleted. Facts already published (for example that a facility exists at an address, with its capacity) remain in the index with their attribution unless you ask for them to be removed as well; corrections are applied at the source and propagate through the normal pipeline, so the change is visible in the <Link href={routes.live()} className={linkCls}>live feed</Link>.118          </P>119        </Section>120121        <Section id="contact" className="scroll-mt-28" title="Contact">122          <P>123            DataCenterIndex is built and operated by Simon-Pierre Boucher. Questions about the crawler, data corrections, removal requests, data partnerships or API access: <a href="mailto:contact@spboucher.ai" className={linkCls}>contact@spboucher.ai</a>. See also the <Link href={routes.methodology()} className={linkCls}>methodology</Link> for the full crawling and reconciliation rules.124          </P>125        </Section>126      </div>127    </Page>128  );129}130