SPB Git forge
38commits 1branches 0releases
338.7 MBsize
maindefault branch
2 h agolast push
HTML 53.9% TypeScript 44.5% JavaScript 0.6% SQL 0.5%
20.6 KB

# Deployment — DataCenterIndex.io

Two OVHcloud dedicated servers in Beauharnois, joined by a private WireGuard link. Public traffic enters through the existing MacLustr Tunnel gateway (BHS64, Caddy + WireGuard). No new public port is opened on either server: the only inbound rules stay 22/tcp, 51820/udp, 51821/udp (ufw). Everything in deploy/ is driven from the laptop with the scripts in deploy/bin/.

# Topology

text
                     Internet
                        │  https://www.datacenterindex.io  (DNS A → 51.161.112.61)
                        ▼
   ┌─────────────────────────────────────────────┐
   │ BHS64 — gateway (Caddy TLS, WireGuard hub)  │  tunnelctl add www.datacenterindex.io BHS128:8300
   │ wg0 10.67.0.1                               │  tunnelctl redirect datacenterindex.io www.datacenterindex.io
   └───────────────┬─────────────────────────────┘
                   │ wg1  10.67.0.1 → 10.67.0.60:8300  (plain HTTP inside the tunnel)
                   ▼
   ┌───────────────────────────────────────────────────────────────────────────────┐
   │ BHS128 — DATA node  (ubuntu@51.161.112.69, 6c/12t, 128 GB, 2×960 GB NVMe R1)  │
   │  docker compose -p dci  (deploy/compose.data.yml)                             │
   │                                                                               │
   │  edge  caddy :8300 ── /api/* ──▶ api  (Fastify :8311)                         │
   │      published 10.67.0.60:8300  └── /* ──▶ web  (Next :8310)                  │
   │  scheduler (BullMQ producer)     [worker-maint: profile maintenance]          │
   │  postgres:17  clickhouse:25.8  redis:7  minio         ── published 10.68.0.1  │
   │  prometheus  grafana  node-exporter  cadvisor  pg/redis exporters              │
   │  /srv/dci/app (repo)   /srv/dci/backups   dci-backup.timer 03:30              │
   └───────────────┬───────────────────────────────────────────────────────────────┘
                   │ dci0  10.68.0.1 ⇄ 10.68.0.2  (private WireGuard /30)
                   ▼
   ┌───────────────────────────────────────────────────────────────────────────────┐
   │ BHS64b — CRAWL node (ubuntu@51.161.112.66, 8c/16t, 64 GB)                     │
   │  docker compose -p dci  (deploy/compose.crawl.yml)                            │
   │  worker (crawl+ingest, DCI_CRAWL_CONCURRENCY=6)  /healthz+/metrics 10.68.0.2:8320
   │  node-exporter 10.68.0.2:9100      [worker-b: profile scale → :8322]          │
   │  /srv/dci-backups  ← nightly rsync of the data node's backups                 │
   └───────────────────────────────────────────────────────────────────────────────┘

The crawl node fetches the web from its own public IP (51.161.112.66); the data node never crawls unless the maintenance profile is enabled (maintenance queue only). Both nodes are Linux and are not mld/agentd nodes.

# Ports

Where Bind Port Service Reachable from
BHS128 10.67.0.60 (wg1) 8300 edge Caddy → web/api BHS64 gateway only
BHS128 10.68.0.1 (dci0) 5432 Postgres 17 BHS64b
BHS128 10.68.0.1 8123 / 9009 ClickHouse HTTP / native (host 9009 → container 9000) BHS64b
BHS128 10.68.0.1 6379 Redis (requirepass) BHS64b
BHS128 10.68.0.1 9000 / 9001 MinIO S3 / console BHS64b, SSH tunnel
BHS128 10.68.0.1 3000 Grafana SSH tunnel (tunnel-grafana.sh)
BHS128 10.68.0.1 9090 Prometheus SSH tunnel
BHS128 docker net 8310 / 8311 / 9180 web / api / caddy metrics containers only
BHS64b 10.68.0.2 (dci0) 8320 (8322) worker /healthz /metrics Prometheus on BHS128
BHS64b 10.68.0.2 9100 node-exporter Prometheus on BHS128

MinIO S3 owns 10.68.0.1:9000; the ClickHouse native protocol is published on 10.68.0.1:9009 (container 9000). The application only uses the ClickHouse HTTP interface (8123); native is for clickhouse-client.

# Files

text
deploy/
  docker/Dockerfile.web|api|worker   multi-stage, node:22-bookworm-slim, pnpm 11 (pnpm fetch → offline filtered install), non-root, tini, HEALTHCHECK
  docker/entrypoint.sh               role by first CMD word: worker | scheduler | cli | migrate | api
  docker/ensure-clickhouse.ts        ClickHouse DB + DDL (run by `migrate`)
  compose.data.yml / compose.crawl.yml
  .env.data.example / .env.crawl.example   → copy to .env.data / .env.crawl (gitignored, never committed)
  edge/Caddyfile                     :8300 → api/web, security + cache headers, zstd/gzip, metrics :9180
  postgres/initdb/01-extensions.sql  pg_trgm, unaccent, pg_stat_statements (first init only)
  clickhouse/config.d, users.d       memory caps, backups path, prometheus endpoint; users default (loopback) + dci
  minio/init.sh, policy-dci-raw.json bucket dci-raw + least-privilege app user
  monitoring/prometheus.yml, alerts.yml, grafana/…   provisioned datasource + "DataCenterIndex — overview"
  systemd/dci-backup.{service,timer} daily 03:30 on the data node (installed by deploy.sh data)
  bin/                               build.sh · deploy.sh · status.sh · logs.sh · backup.sh · restore.sh · psql.sh · clickhouse.sh · tunnel-grafana.sh
  bin/remote/backup-run.sh           runs on BHS128 (timer + backup.sh)
  dev-compose.yml                    laptop only (ClickHouse + MinIO)

Images are built on the target host (docker compose build after an rsync of the repo to /srv/dci/app, excluding node_modules, .next, data, qa, .git, env files). The laptop never pushes images. BuildKit cache mounts keep the pnpm store and the Next cache between builds.

# First deploy (runbook)

Prerequisites already in place: Ubuntu 24.04, Docker 29 + compose v5, wg1 (10.67.0.60) and dci0 (10.68.0.1 / 10.68.0.2) up, ssh aliases BHS128 and BHS64b, sudo NOPASSWD.

  1. Secrets (laptop, once):
    bash
    cp deploy/.env.data.example deploy/.env.data
    cp deploy/.env.crawl.example deploy/.env.crawl
    for v in POSTGRES_PASSWORD CLICKHOUSE_PASSWORD REDIS_PASSWORD MINIO_ROOT_PASSWORD S3_SECRET_KEY DCI_ADMIN_TOKEN GRAFANA_ADMIN_PASSWORD; do echo "$v=$(openssl rand -hex 24)"; done
    Paste the values into deploy/.env.data; copy the same POSTGRES/REDIS/CLICKHOUSE/S3 values into deploy/.env.crawl, plus SCRAPFLY_API_KEY / FIRECRAWL_API_KEY there.
  2. DNS (GoDaddy): A datacenterindex.io → 51.161.112.61, A www → 51.161.112.61 (or CNAME to apex).
  3. Data node:
    bash
    deploy/bin/deploy.sh data
    Does: preflight (checks the WireGuard IPs, sets net.ipv4.ip_nonlocal_bind=1 so Docker can bind 10.67.0.60/10.68.0.1 even if dockerd starts before WireGuard after a reboot), /srv/dci layout, rsync, compose build, up stores, minio-init, run --rm migrate (migrations → seed → ClickHouse DDL), up -d --remove-orphans, installs dci-backup.timer, creates the ~/.ssh/dci-backup key on BHS128 and authorizes it on BHS64b (/srv/dci-backups), then waits for http://10.67.0.60:8300/api/health.
  4. Gateway route (on BHS64, as documented in cluster-skill/MACLUSTR-TUNNEL.md; BHS128 is already in the gateway ipmap as 10.67.0.60):
    bash
    ssh BHS64 'tunnelctl add www.datacenterindex.io BHS128:8300 && tunnelctl redirect datacenterindex.io www.datacenterindex.io'
    # or from the laptop: mlt add www.datacenterindex.io BHS128:8300
    curl -I https://www.datacenterindex.io/api/health
  5. Crawl node:
    bash
    deploy/bin/deploy.sh crawl
  6. Verify: deploy/bin/status.sh, deploy/bin/logs.sh worker, deploy/bin/tunnel-grafana.sh (Grafana → folder DataCenterIndex → overview; all Prometheus targets green).
  7. First connectors: ssh BHS64b then cd /srv/dci/app && docker compose -f deploy/compose.crawl.yml --env-file deploy/.env.crawl run --rm cli run peeringdb --dry-run --limit 5 (the cli service = pnpm dci …, running from the crawl IP with production credentials).

# Update (routine deploy)

bash
deploy/bin/deploy.sh data          # rsync → build → migrate → up -d (recreates only changed services) → health
deploy/bin/deploy.sh crawl         # rsync → build worker → up -d → /healthz
# faster variants
deploy/bin/deploy.sh data --skip-migrate --no-backup-setup
deploy/bin/build.sh data web       # just rebuild one image (then deploy.sh data --no-build)

Behaviour: compose build runs before up -d, so the old containers keep serving during the build. up -d recreates a service only when its image/config changed; web/api restart in a few seconds and the edge Caddy retries upstreams for 5 s (lb_try_duration), which absorbs the restart for most requests. Postgres, ClickHouse, Redis, MinIO are never recreated by an app deploy (their definitions do not change). Migrations are plain SQL, idempotent and tracked in _dci_migrations; the migrate one-shot is safe to re-run. Run deploy.sh data before deploy.sh crawl when a migration changes the schema the workers write to.

Roll back = redeploy the previous git revision (git checkout <rev> && deploy/bin/deploy.sh data --skip-migrate). Migrations are forward-only; a schema rollback goes through restore.sh.

# Secrets & rotation

All secrets live in deploy/.env.data and deploy/.env.crawl on the laptop (gitignored, .gitignore covers deploy/.env*) and are copied to /srv/dci/app/deploy/.env.<target> (mode 600) by the scripts. Never in git, never sent to the browser (NEXT_PUBLIC_* are the only build-time public values).

Secret Rotate
POSTGRES_PASSWORD psql.sh -c "ALTER USER dci PASSWORD 'new'" → update both env files → deploy.sh data --no-build --skip-migrate then deploy.sh crawl --no-build
REDIS_PASSWORD update env files → deploy.sh data --no-build --skip-migrate (redis restarts with the new requirepass; queued jobs persist in AOF) → deploy.sh crawl --no-build
CLICKHOUSE_PASSWORD update env files → same two deploys (users are XML from_env, applied at start)
MinIO root / S3_SECRET_KEY root: update env → deploy (MinIO reads root creds at start). App user: mc admin user add local dci <new> runs automatically in minio-init on the next up (it re-adds the user) → update .env.crawl → deploy.sh crawl --no-build
DCI_ADMIN_TOKEN update .env.data → deploy.sh data --no-build --skip-migrate (api restarts)
SCRAPFLY_API_KEY / FIRECRAWL_API_KEY .env.crawl → deploy.sh crawl --no-build
GRAFANA_ADMIN_PASSWORD .env.data → deploy; Grafana keeps the admin password in its DB — also change it in the UI or docker compose exec grafana grafana cli admin reset-admin-password <new>

# Backups & restore

Nightly dci-backup.timer (03:30 America/Toronto, Persistent=true) runs deploy/bin/remote/backup-run.sh on BHS128; deploy/bin/backup.sh runs the same thing on demand.

What Where (BHS128) How
Postgres (truth) /srv/dci/backups/pg/dci-<stamp>.dump pg_dump -Fc --compress=zstd
Connector configs, parsers, migrations, deploy config, git rev /srv/dci/backups/config/dci-config-<stamp>.tar.zst tar
ClickHouse dci /srv/dci/backups/clickhouse/dci-ch-<stamp>.zip BACKUP DATABASE … TO File() (best effort)
MinIO dci-raw (raw bodies) /srv/dci/backups/minio/dci-raw/ mc mirror --overwrite (incremental mirror)

Retention: 14 daily + 8 weekly (Sunday) for pg/config/clickhouse; the MinIO mirror is a single current copy. After pruning, the whole /srv/dci/backups/ is rsynced over the private link to ubuntu@10.68.0.2:/srv/dci-backups/ (BHS64b) with the dedicated key ~/.ssh/dci-backup (set up by deploy.sh data). Optional third copy on the laptop: deploy/bin/backup.sh --pull (latest dump + config → ~/Backups/dci/) or --pull-all.

Restore — always into a separate database first:

bash
deploy/bin/restore.sh latest                       # → database dci_restore on BHS128 (prod untouched)
deploy/bin/psql.sh -d dci_restore -c 'select count(*) from facilities'
deploy/bin/restore.sh latest --swap                # stops api/scheduler/workers, dci → dci_old_<stamp>, dci_restore → dci, restarts
deploy/bin/restore.sh ~/Backups/dci/dci-20260911-033000.dump --db dci_check

--swap is reversible (ALTER DATABASE … RENAME) until you drop dci_old_<stamp>. ClickHouse restore: clickhouse.sh --query "RESTORE DATABASE dci FROM File('/backups/dci-ch-<stamp>.zip')" (drop or rename the database first). MinIO: mc mirror /backup/dci-raw local/dci-raw from the minio-init image with /srv/dci/backups/minio mounted at /backup (see backup-run.sh step 4 for the exact compose run).

Before a risky change: deploy/bin/backup.sh (takes a fresh dump synchronously).

# Monitoring

Prometheus (90 d / 20 GB) scrapes: api:8311/api/metrics (dci_api_requests_total, dci_api_request_duration_seconds, cache counters), the crawl workers at 10.68.0.2:8320/metrics (+ :8322 for worker-b), the edge Caddy, node-exporter on both hosts, cAdvisor, postgres-, redis-exporter, ClickHouse (:9363) and MinIO. Rules in deploy/monitoring/alerts.yml (target down, API p95 > 1 s, 5xx > 5 %, crawl stalled, premium budget > 90 %, disk < 10 %, memory > 92 %, PG connections > 170) — no Alertmanager; they show as state in Grafana → Alerting.

Grafana is bound to 10.68.0.1:3000 (never public). Access:

bash
deploy/bin/tunnel-grafana.sh        # → http://localhost:3000 (also :9090 Prometheus, :9001 MinIO console)

Login GRAFANA_ADMIN_USER / GRAFANA_ADMIN_PASSWORD from .env.data. Home dashboard = DataCenterIndex — overview (crawl rate by outcome, premium credits/h by provider, queue depth, fetch and API latency, requests by status class, CPU/RAM/disk per node, container memory, PG connections/size). Dashboards are provisioned read-only from deploy/monitoring/grafana/dashboards/; edit the JSON in the repo and redeploy.

Worker metric contract expected by the dashboard/alerts (to be exported by apps/worker on /metrics): dci_crawl_fetches_total{connector,level,outcome=ok|not_modified|error|blocked|skipped}, dci_crawl_credits_total{provider=scrapfly|firecrawl}, dci_crawl_daily_budget{provider}, dci_queue_jobs{queue,state=waiting|active|delayed|failed}, dci_crawl_fetch_duration_seconds_bucket.

# Firewall statement

Nothing new is exposed publicly. ufw on both servers keeps only 22/tcp, 51820/udp, 51821/udp. Every Docker published port is bound to a WireGuard address (10.67.0.60 for the edge, 10.68.0.1 / 10.68.0.2 for stores and metrics), so the iptables DNAT rules Docker installs only match traffic already inside the tunnels; packets to 51.161.112.69:<port> or 51.161.112.66:<port> do not match and are dropped by ufw. Inter-container traffic stays on the compose network dci_default. The data node has no browser-facing service other than the edge, and the edge answers 404 to /api/metrics. TLS terminates on the gateway (BHS64, Let's Encrypt via Caddy); the edge trusts X-Forwarded-* from 10.67.0.0/24 only.

Docker bypasses ufw for published ports by design — that is why binding to the tunnel IPs (not 0.0.0.0) is the invariant to preserve when editing the compose files.

# Gateway (BHS64) commands

bash
# create the public route (once) — BHS128 = 10.67.0.60 in the gateway ipmap
ssh BHS64 tunnelctl add www.datacenterindex.io BHS128:8300
ssh BHS64 tunnelctl redirect datacenterindex.io www.datacenterindex.io
ssh BHS64 tunnelctl status                    # routes + WireGuard peers
# from the laptop wrapper
~/Desktop/cluster-skill/mlt add www.datacenterindex.io BHS128:8300

# Troubleshooting

Symptom Check / fix
deploy.sh data fails at up: cannot assign requested address on 10.67.0.60 or 10.68.0.1 WireGuard not up or IPs changed: ssh BHS128 ip -4 a show wg1; ip -4 a show dci0. deploy.sh sets net.ipv4.ip_nonlocal_bind=1 so this only happens if the sysctl was lost — sudo sysctl -p /etc/sysctl.d/90-dci.conf
502 from the public URL, edge healthy gateway route missing/stale → ssh BHS64 tunnelctl status; ssh BHS64 curl -sI http://10.67.0.60:8300/api/health
edge healthy, /api/health 503 api not ready: logs.sh api (DB URL, migrations). compose run --rm migrate again
web 500 / blank logs.sh web; the container must have API_URL_INTERNAL=http://api:8311; NEXT_PUBLIC_API_URL is baked at build → rebuild web if it changed
workers idle, queue growing logs.sh worker; from BHS64b: nc -zv 10.68.0.1 6379 5432 8123 9000 (dci0 down?), Redis password mismatch between the two env files
workers fail on S3 AccessDenied minio-init did not run (check compose ps -a minio-init, rerun up -d minio-init), or S3_ACCESS_KEY/S3_SECRET_KEY differ between env files
ClickHouse auth errors CLICKHOUSE_USER=dci + password in both env files; the default user is loopback-only by design
Prometheus target 10.68.0.2:8320 down crawl worker not running or WORKER_PORT changed; ssh BHS64b curl -s http://10.68.0.2:8320/healthz
Grafana shows no data tunnel-grafana.sh → http://localhost:9090/targets; metric names must match the contract above
Disk pressure on BHS128 status.sh data; largest consumers: /var/lib/docker/volumes/dci_miniodata, dci_pgdata, /srv/dci/backups/minio. Prune old images: docker image prune -af --filter until=168h
After a reboot nothing listens systemctl status docker; compose ps; containers are restart: unless-stopped so they return once dockerd is up (WireGuard first thanks to ip_nonlocal_bind)
Backup timer not running ssh BHS128 systemctl list-timers dci-backup.timer; journalctl -u dci-backup.service -n 50; off-node rsync needs ~/.ssh/dci-backup.pub in ubuntu@BHS64b:~/.ssh/authorized_keys
Need a shell in a container ssh BHS128 'cd /srv/dci/app && docker compose -f deploy/compose.data.yml --env-file deploy/.env.data exec api sh'
Rebuild from scratch (no cache) ssh BHS128 'cd /srv/dci/app && docker compose -f deploy/compose.data.yml --env-file deploy/.env.data build --no-cache web'

Local validation without servers: docker compose -f deploy/compose.data.yml --env-file deploy/.env.data.example config -q and docker build -f deploy/docker/Dockerfile.worker . from the repo root.

# 2026-09-12 upgrade — post-deploy data repair and gotchas

  • Deploy order when the API contract changes: the web image prerenders / and other static routes against the live API (NEXT_PUBLIC_API_URL build arg). Either keep pages tolerant of missing fields (done for the home page) or build/up api scheduler worker-maint first, then web (see /tmp/dci-redeploy.sh pattern: compose build api …, run --rm migrate, up -d api …, then compose build web, up -d web edge).
  • deploy/rsync-exclude.txt and .dockerignore must exclude only the root coverage/ report directory (/coverage): a bare coverage pattern silently dropped the /coverage web route from the image.
  • compose run --rm cli <cmd> passes <cmd> as the entrypoint ROLE — run dci commands as run --rm -T cli cli <cmd>.
  • scripts/ is not baked into the worker image; run repair scripts with bind mounts: run --rm -T -v /srv/dci/app/scripts:/app/scripts -v /srv/dci/app/packages/core/src:/app/packages/core/src cli node /app/node_modules/tsx/dist/cli.mjs /app/scripts/quality-apply.ts ….
  • deploy/bin/post-upgrade.sh [baseline|repair|reprocess|unbacked|refresh|audit] runs the whole repair sequence from the laptop: quality sweep + snapshot, hide vetoed / unconfirmed projects, fix slugs, link campuses, dci reprocess <id> --stale for every connector (offline, archived bodies), null figures no site-scoped claim backs, refresh stats + rankings, print the largest values.
  • API/edge health probes use /api/ready (Postgres-aware). DCI_WORKER_URL=http://worker-maint:8320 lets the API proxy the extraction debugger (/api/admin/documents/:id/trace) and /api/admin/data-gaps.