Deployment — DataCenterIndex.io
Two OVHcloud dedicated servers in Beauharnois, joined by a private WireGuard link. Public traffic enters
through the existing MacLustr Tunnel gateway (BHS64, Caddy + WireGuard). No new public port is opened
on either server: the only inbound rules stay 22/tcp, 51820/udp, 51821/udp (ufw). Everything in
deploy/ is driven from the laptop with the scripts in deploy/bin/.
Topology
Internet
│ https://www.datacenterindex.io (DNS A → 51.161.112.61)
▼
┌─────────────────────────────────────────────┐
│ BHS64 — gateway (Caddy TLS, WireGuard hub) │ tunnelctl add www.datacenterindex.io BHS128:8300
│ wg0 10.67.0.1 │ tunnelctl redirect datacenterindex.io www.datacenterindex.io
└───────────────┬─────────────────────────────┘
│ wg1 10.67.0.1 → 10.67.0.60:8300 (plain HTTP inside the tunnel)
▼
┌───────────────────────────────────────────────────────────────────────────────┐
│ BHS128 — DATA node (ubuntu@51.161.112.69, 6c/12t, 128 GB, 2×960 GB NVMe R1) │
│ docker compose -p dci (deploy/compose.data.yml) │
│ │
│ edge caddy :8300 ── /api/* ──▶ api (Fastify :8311) │
│ published 10.67.0.60:8300 └── /* ──▶ web (Next :8310) │
│ scheduler (BullMQ producer) [worker-maint: profile maintenance] │
│ postgres:17 clickhouse:25.8 redis:7 minio ── published 10.68.0.1 │
│ prometheus grafana node-exporter cadvisor pg/redis exporters │
│ /srv/dci/app (repo) /srv/dci/backups dci-backup.timer 03:30 │
└───────────────┬───────────────────────────────────────────────────────────────┘
│ dci0 10.68.0.1 ⇄ 10.68.0.2 (private WireGuard /30)
▼
┌───────────────────────────────────────────────────────────────────────────────┐
│ BHS64b — CRAWL node (ubuntu@51.161.112.66, 8c/16t, 64 GB) │
│ docker compose -p dci (deploy/compose.crawl.yml) │
│ worker (crawl+ingest, DCI_CRAWL_CONCURRENCY=6) /healthz+/metrics 10.68.0.2:8320
│ node-exporter 10.68.0.2:9100 [worker-b: profile scale → :8322] │
│ /srv/dci-backups ← nightly rsync of the data node's backups │
└───────────────────────────────────────────────────────────────────────────────┘The crawl node fetches the web from its own public IP (51.161.112.66); the data node never crawls unless the
maintenance profile is enabled (maintenance queue only). Both nodes are Linux and are not mld/agentd nodes.
Ports
| Where | Bind | Port | Service | Reachable from |
|---|---|---|---|---|
| BHS128 | 10.67.0.60 (wg1) | 8300 | edge Caddy → web/api | BHS64 gateway only |
| BHS128 | 10.68.0.1 (dci0) | 5432 | Postgres 17 | BHS64b |
| BHS128 | 10.68.0.1 | 8123 / 9009 | ClickHouse HTTP / native (host 9009 → container 9000) | BHS64b |
| BHS128 | 10.68.0.1 | 6379 | Redis (requirepass) | BHS64b |
| BHS128 | 10.68.0.1 | 9000 / 9001 | MinIO S3 / console | BHS64b, SSH tunnel |
| BHS128 | 10.68.0.1 | 3000 | Grafana | SSH tunnel (tunnel-grafana.sh) |
| BHS128 | 10.68.0.1 | 9090 | Prometheus | SSH tunnel |
| BHS128 | docker net | 8310 / 8311 / 9180 | web / api / caddy metrics | containers only |
| BHS64b | 10.68.0.2 (dci0) | 8320 (8322) | worker /healthz /metrics | Prometheus on BHS128 |
| BHS64b | 10.68.0.2 | 9100 | node-exporter | Prometheus on BHS128 |
MinIO S3 owns 10.68.0.1:9000; the ClickHouse native protocol is published on 10.68.0.1:9009 (container
9000). The application only uses the ClickHouse HTTP interface (8123); native is for clickhouse-client.
Files
deploy/
docker/Dockerfile.web|api|worker multi-stage, node:22-bookworm-slim, pnpm 11 (pnpm fetch → offline filtered install), non-root, tini, HEALTHCHECK
docker/entrypoint.sh role by first CMD word: worker | scheduler | cli | migrate | api
docker/ensure-clickhouse.ts ClickHouse DB + DDL (run by `migrate`)
compose.data.yml / compose.crawl.yml
.env.data.example / .env.crawl.example → copy to .env.data / .env.crawl (gitignored, never committed)
edge/Caddyfile :8300 → api/web, security + cache headers, zstd/gzip, metrics :9180
postgres/initdb/01-extensions.sql pg_trgm, unaccent, pg_stat_statements (first init only)
clickhouse/config.d, users.d memory caps, backups path, prometheus endpoint; users default (loopback) + dci
minio/init.sh, policy-dci-raw.json bucket dci-raw + least-privilege app user
monitoring/prometheus.yml, alerts.yml, grafana/… provisioned datasource + "DataCenterIndex — overview"
systemd/dci-backup.{service,timer} daily 03:30 on the data node (installed by deploy.sh data)
bin/ build.sh · deploy.sh · status.sh · logs.sh · backup.sh · restore.sh · psql.sh · clickhouse.sh · tunnel-grafana.sh
bin/remote/backup-run.sh runs on BHS128 (timer + backup.sh)
dev-compose.yml laptop only (ClickHouse + MinIO)Images are built on the target host (docker compose build after an rsync of the repo to
/srv/dci/app, excluding node_modules, .next, data, qa, .git, env files). The laptop never pushes
images. BuildKit cache mounts keep the pnpm store and the Next cache between builds.
First deploy (runbook)
Prerequisites already in place: Ubuntu 24.04, Docker 29 + compose v5, wg1 (10.67.0.60) and dci0
(10.68.0.1 / 10.68.0.2) up, ssh aliases BHS128 and BHS64b, sudo NOPASSWD.
- Secrets (laptop, once):Paste the values intobash
cp deploy/.env.data.example deploy/.env.data cp deploy/.env.crawl.example deploy/.env.crawl for v in POSTGRES_PASSWORD CLICKHOUSE_PASSWORD REDIS_PASSWORD MINIO_ROOT_PASSWORD S3_SECRET_KEY DCI_ADMIN_TOKEN GRAFANA_ADMIN_PASSWORD; do echo "$v=$(openssl rand -hex 24)"; donedeploy/.env.data; copy the same POSTGRES/REDIS/CLICKHOUSE/S3 values intodeploy/.env.crawl, plusSCRAPFLY_API_KEY/FIRECRAWL_API_KEYthere. - DNS (GoDaddy):
A datacenterindex.io → 51.161.112.61,A www → 51.161.112.61(or CNAME to apex). - Data node:Does: preflight (checks the WireGuard IPs, setsbash
deploy/bin/deploy.sh datanet.ipv4.ip_nonlocal_bind=1so Docker can bind 10.67.0.60/10.68.0.1 even if dockerd starts before WireGuard after a reboot),/srv/dcilayout, rsync,compose build,upstores,minio-init,run --rm migrate(migrations → seed → ClickHouse DDL),up -d --remove-orphans, installsdci-backup.timer, creates the~/.ssh/dci-backupkey on BHS128 and authorizes it on BHS64b (/srv/dci-backups), then waits forhttp://10.67.0.60:8300/api/health. - Gateway route (on BHS64, as documented in
cluster-skill/MACLUSTR-TUNNEL.md; BHS128 is already in the gateway ipmap as 10.67.0.60):bashssh BHS64 'tunnelctl add www.datacenterindex.io BHS128:8300 && tunnelctl redirect datacenterindex.io www.datacenterindex.io' # or from the laptop: mlt add www.datacenterindex.io BHS128:8300 curl -I https://www.datacenterindex.io/api/health - Crawl node:bash
deploy/bin/deploy.sh crawl - Verify:
deploy/bin/status.sh,deploy/bin/logs.sh worker,deploy/bin/tunnel-grafana.sh(Grafana → folder DataCenterIndex → overview; all Prometheus targets green). - First connectors:
ssh BHS64bthencd /srv/dci/app && docker compose -f deploy/compose.crawl.yml --env-file deploy/.env.crawl run --rm cli run peeringdb --dry-run --limit 5(thecliservice =pnpm dci …, running from the crawl IP with production credentials).
Update (routine deploy)
deploy/bin/deploy.sh data # rsync → build → migrate → up -d (recreates only changed services) → health
deploy/bin/deploy.sh crawl # rsync → build worker → up -d → /healthz
# faster variants
deploy/bin/deploy.sh data --skip-migrate --no-backup-setup
deploy/bin/build.sh data web # just rebuild one image (then deploy.sh data --no-build)Behaviour: compose build runs before up -d, so the old containers keep serving during the build. up -d
recreates a service only when its image/config changed; web/api restart in a few seconds and the edge
Caddy retries upstreams for 5 s (lb_try_duration), which absorbs the restart for most requests. Postgres,
ClickHouse, Redis, MinIO are never recreated by an app deploy (their definitions do not change). Migrations
are plain SQL, idempotent and tracked in _dci_migrations; the migrate one-shot is safe to re-run. Run
deploy.sh data before deploy.sh crawl when a migration changes the schema the workers write to.
Roll back = redeploy the previous git revision (git checkout <rev> && deploy/bin/deploy.sh data --skip-migrate).
Migrations are forward-only; a schema rollback goes through restore.sh.
Secrets & rotation
All secrets live in deploy/.env.data and deploy/.env.crawl on the laptop (gitignored, .gitignore
covers deploy/.env*) and are copied to /srv/dci/app/deploy/.env.<target> (mode 600) by the scripts.
Never in git, never sent to the browser (NEXT_PUBLIC_* are the only build-time public values).
| Secret | Rotate |
|---|---|
POSTGRES_PASSWORD |
psql.sh -c "ALTER USER dci PASSWORD 'new'" → update both env files → deploy.sh data --no-build --skip-migrate then deploy.sh crawl --no-build |
REDIS_PASSWORD |
update env files → deploy.sh data --no-build --skip-migrate (redis restarts with the new requirepass; queued jobs persist in AOF) → deploy.sh crawl --no-build |
CLICKHOUSE_PASSWORD |
update env files → same two deploys (users are XML from_env, applied at start) |
MinIO root / S3_SECRET_KEY |
root: update env → deploy (MinIO reads root creds at start). App user: mc admin user add local dci <new> runs automatically in minio-init on the next up (it re-adds the user) → update .env.crawl → deploy.sh crawl --no-build |
DCI_ADMIN_TOKEN |
update .env.data → deploy.sh data --no-build --skip-migrate (api restarts) |
SCRAPFLY_API_KEY / FIRECRAWL_API_KEY |
.env.crawl → deploy.sh crawl --no-build |
GRAFANA_ADMIN_PASSWORD |
.env.data → deploy; Grafana keeps the admin password in its DB — also change it in the UI or docker compose exec grafana grafana cli admin reset-admin-password <new> |
Backups & restore
Nightly dci-backup.timer (03:30 America/Toronto, Persistent=true) runs deploy/bin/remote/backup-run.sh
on BHS128; deploy/bin/backup.sh runs the same thing on demand.
| What | Where (BHS128) | How |
|---|---|---|
| Postgres (truth) | /srv/dci/backups/pg/dci-<stamp>.dump |
pg_dump -Fc --compress=zstd |
| Connector configs, parsers, migrations, deploy config, git rev | /srv/dci/backups/config/dci-config-<stamp>.tar.zst |
tar |
ClickHouse dci |
/srv/dci/backups/clickhouse/dci-ch-<stamp>.zip |
BACKUP DATABASE … TO File() (best effort) |
MinIO dci-raw (raw bodies) |
/srv/dci/backups/minio/dci-raw/ |
mc mirror --overwrite (incremental mirror) |
Retention: 14 daily + 8 weekly (Sunday) for pg/config/clickhouse; the MinIO mirror is a single current
copy. After pruning, the whole /srv/dci/backups/ is rsynced over the private link to
ubuntu@10.68.0.2:/srv/dci-backups/ (BHS64b) with the dedicated key ~/.ssh/dci-backup (set up by
deploy.sh data). Optional third copy on the laptop: deploy/bin/backup.sh --pull (latest dump + config
→ ~/Backups/dci/) or --pull-all.
Restore — always into a separate database first:
deploy/bin/restore.sh latest # → database dci_restore on BHS128 (prod untouched)
deploy/bin/psql.sh -d dci_restore -c 'select count(*) from facilities'
deploy/bin/restore.sh latest --swap # stops api/scheduler/workers, dci → dci_old_<stamp>, dci_restore → dci, restarts
deploy/bin/restore.sh ~/Backups/dci/dci-20260911-033000.dump --db dci_check--swap is reversible (ALTER DATABASE … RENAME) until you drop dci_old_<stamp>. ClickHouse restore:
clickhouse.sh --query "RESTORE DATABASE dci FROM File('/backups/dci-ch-<stamp>.zip')" (drop or rename the
database first). MinIO: mc mirror /backup/dci-raw local/dci-raw from the minio-init image with
/srv/dci/backups/minio mounted at /backup (see backup-run.sh step 4 for the exact compose run).
Before a risky change: deploy/bin/backup.sh (takes a fresh dump synchronously).
Monitoring
Prometheus (90 d / 20 GB) scrapes: api:8311/api/metrics (dci_api_requests_total,
dci_api_request_duration_seconds, cache counters), the crawl workers at 10.68.0.2:8320/metrics
(+ :8322 for worker-b), the edge Caddy, node-exporter on both hosts, cAdvisor, postgres-, redis-exporter,
ClickHouse (:9363) and MinIO. Rules in deploy/monitoring/alerts.yml (target down, API p95 > 1 s, 5xx > 5 %,
crawl stalled, premium budget > 90 %, disk < 10 %, memory > 92 %, PG connections > 170) — no Alertmanager;
they show as state in Grafana → Alerting.
Grafana is bound to 10.68.0.1:3000 (never public). Access:
deploy/bin/tunnel-grafana.sh # → http://localhost:3000 (also :9090 Prometheus, :9001 MinIO console)Login GRAFANA_ADMIN_USER / GRAFANA_ADMIN_PASSWORD from .env.data. Home dashboard = DataCenterIndex —
overview (crawl rate by outcome, premium credits/h by provider, queue depth, fetch and API latency,
requests by status class, CPU/RAM/disk per node, container memory, PG connections/size). Dashboards are
provisioned read-only from deploy/monitoring/grafana/dashboards/; edit the JSON in the repo and redeploy.
Worker metric contract expected by the dashboard/alerts (to be exported by apps/worker on /metrics):
dci_crawl_fetches_total{connector,level,outcome=ok|not_modified|error|blocked|skipped},
dci_crawl_credits_total{provider=scrapfly|firecrawl}, dci_crawl_daily_budget{provider},
dci_queue_jobs{queue,state=waiting|active|delayed|failed}, dci_crawl_fetch_duration_seconds_bucket.
Firewall statement
Nothing new is exposed publicly. ufw on both servers keeps only 22/tcp, 51820/udp, 51821/udp. Every
Docker published port is bound to a WireGuard address (10.67.0.60 for the edge, 10.68.0.1 / 10.68.0.2
for stores and metrics), so the iptables DNAT rules Docker installs only match traffic already inside the
tunnels; packets to 51.161.112.69:<port> or 51.161.112.66:<port> do not match and are dropped by ufw.
Inter-container traffic stays on the compose network dci_default. The data node has no browser-facing
service other than the edge, and the edge answers 404 to /api/metrics. TLS terminates on the gateway
(BHS64, Let's Encrypt via Caddy); the edge trusts X-Forwarded-* from 10.67.0.0/24 only.
Docker bypasses ufw for published ports by design — that is why binding to the tunnel IPs (not 0.0.0.0)
is the invariant to preserve when editing the compose files.
Gateway (BHS64) commands
# create the public route (once) — BHS128 = 10.67.0.60 in the gateway ipmap
ssh BHS64 tunnelctl add www.datacenterindex.io BHS128:8300
ssh BHS64 tunnelctl redirect datacenterindex.io www.datacenterindex.io
ssh BHS64 tunnelctl status # routes + WireGuard peers
# from the laptop wrapper
~/Desktop/cluster-skill/mlt add www.datacenterindex.io BHS128:8300Troubleshooting
| Symptom | Check / fix |
|---|---|
deploy.sh data fails at up: cannot assign requested address on 10.67.0.60 or 10.68.0.1 |
WireGuard not up or IPs changed: ssh BHS128 ip -4 a show wg1; ip -4 a show dci0. deploy.sh sets net.ipv4.ip_nonlocal_bind=1 so this only happens if the sysctl was lost — sudo sysctl -p /etc/sysctl.d/90-dci.conf |
| 502 from the public URL, edge healthy | gateway route missing/stale → ssh BHS64 tunnelctl status; ssh BHS64 curl -sI http://10.67.0.60:8300/api/health |
edge healthy, /api/health 503 |
api not ready: logs.sh api (DB URL, migrations). compose run --rm migrate again |
| web 500 / blank | logs.sh web; the container must have API_URL_INTERNAL=http://api:8311; NEXT_PUBLIC_API_URL is baked at build → rebuild web if it changed |
| workers idle, queue growing | logs.sh worker; from BHS64b: nc -zv 10.68.0.1 6379 5432 8123 9000 (dci0 down?), Redis password mismatch between the two env files |
workers fail on S3 AccessDenied |
minio-init did not run (check compose ps -a minio-init, rerun up -d minio-init), or S3_ACCESS_KEY/S3_SECRET_KEY differ between env files |
| ClickHouse auth errors | CLICKHOUSE_USER=dci + password in both env files; the default user is loopback-only by design |
Prometheus target 10.68.0.2:8320 down |
crawl worker not running or WORKER_PORT changed; ssh BHS64b curl -s http://10.68.0.2:8320/healthz |
| Grafana shows no data | tunnel-grafana.sh → http://localhost:9090/targets; metric names must match the contract above |
| Disk pressure on BHS128 | status.sh data; largest consumers: /var/lib/docker/volumes/dci_miniodata, dci_pgdata, /srv/dci/backups/minio. Prune old images: docker image prune -af --filter until=168h |
| After a reboot nothing listens | systemctl status docker; compose ps; containers are restart: unless-stopped so they return once dockerd is up (WireGuard first thanks to ip_nonlocal_bind) |
| Backup timer not running | ssh BHS128 systemctl list-timers dci-backup.timer; journalctl -u dci-backup.service -n 50; off-node rsync needs ~/.ssh/dci-backup.pub in ubuntu@BHS64b:~/.ssh/authorized_keys |
| Need a shell in a container | ssh BHS128 'cd /srv/dci/app && docker compose -f deploy/compose.data.yml --env-file deploy/.env.data exec api sh' |
| Rebuild from scratch (no cache) | ssh BHS128 'cd /srv/dci/app && docker compose -f deploy/compose.data.yml --env-file deploy/.env.data build --no-cache web' |
Local validation without servers: docker compose -f deploy/compose.data.yml --env-file deploy/.env.data.example config -q
and docker build -f deploy/docker/Dockerfile.worker . from the repo root.
2026-09-12 upgrade — post-deploy data repair and gotchas
- Deploy order when the API contract changes: the web image prerenders
/and other static routes against the live API (NEXT_PUBLIC_API_URLbuild arg). Either keep pages tolerant of missing fields (done for the home page) or build/upapi scheduler worker-maintfirst, thenweb(see/tmp/dci-redeploy.shpattern:compose build api …,run --rm migrate,up -d api …, thencompose build web,up -d web edge). deploy/rsync-exclude.txtand.dockerignoremust exclude only the rootcoverage/report directory (/coverage): a barecoveragepattern silently dropped the/coverageweb route from the image.compose run --rm cli <cmd>passes<cmd>as the entrypoint ROLE — rundcicommands asrun --rm -T cli cli <cmd>.scripts/is not baked into the worker image; run repair scripts with bind mounts:run --rm -T -v /srv/dci/app/scripts:/app/scripts -v /srv/dci/app/packages/core/src:/app/packages/core/src cli node /app/node_modules/tsx/dist/cli.mjs /app/scripts/quality-apply.ts ….deploy/bin/post-upgrade.sh [baseline|repair|reprocess|unbacked|refresh|audit]runs the whole repair sequence from the laptop: quality sweep + snapshot, hide vetoed / unconfirmed projects, fix slugs, link campuses,dci reprocess <id> --stalefor every connector (offline, archived bodies), null figures no site-scoped claim backs, refresh stats + rankings, print the largest values.- API/edge health probes use
/api/ready(Postgres-aware).DCI_WORKER_URL=http://worker-maint:8320lets the API proxy the extraction debugger (/api/admin/documents/:id/trace) and/api/admin/data-gaps.