# Deployment — DataCenterIndex.io Two OVHcloud dedicated servers in Beauharnois, joined by a private WireGuard link. Public traffic enters through the existing **MacLustr Tunnel** gateway (BHS64, Caddy + WireGuard). **No new public port is opened on either server**: the only inbound rules stay `22/tcp`, `51820/udp`, `51821/udp` (ufw). Everything in `deploy/` is driven from the laptop with the scripts in `deploy/bin/`. ## Topology ``` Internet │ https://www.datacenterindex.io (DNS A → 51.161.112.61) ▼ ┌─────────────────────────────────────────────┐ │ BHS64 — gateway (Caddy TLS, WireGuard hub) │ tunnelctl add www.datacenterindex.io BHS128:8300 │ wg0 10.67.0.1 │ tunnelctl redirect datacenterindex.io www.datacenterindex.io └───────────────┬─────────────────────────────┘ │ wg1 10.67.0.1 → 10.67.0.60:8300 (plain HTTP inside the tunnel) ▼ ┌───────────────────────────────────────────────────────────────────────────────┐ │ BHS128 — DATA node (ubuntu@51.161.112.69, 6c/12t, 128 GB, 2×960 GB NVMe R1) │ │ docker compose -p dci (deploy/compose.data.yml) │ │ │ │ edge caddy :8300 ── /api/* ──▶ api (Fastify :8311) │ │ published 10.67.0.60:8300 └── /* ──▶ web (Next :8310) │ │ scheduler (BullMQ producer) [worker-maint: profile maintenance] │ │ postgres:17 clickhouse:25.8 redis:7 minio ── published 10.68.0.1 │ │ prometheus grafana node-exporter cadvisor pg/redis exporters │ │ /srv/dci/app (repo) /srv/dci/backups dci-backup.timer 03:30 │ └───────────────┬───────────────────────────────────────────────────────────────┘ │ dci0 10.68.0.1 ⇄ 10.68.0.2 (private WireGuard /30) ▼ ┌───────────────────────────────────────────────────────────────────────────────┐ │ BHS64b — CRAWL node (ubuntu@51.161.112.66, 8c/16t, 64 GB) │ │ docker compose -p dci (deploy/compose.crawl.yml) │ │ worker (crawl+ingest, DCI_CRAWL_CONCURRENCY=6) /healthz+/metrics 10.68.0.2:8320 │ node-exporter 10.68.0.2:9100 [worker-b: profile scale → :8322] │ │ /srv/dci-backups ← nightly rsync of the data node's backups │ └───────────────────────────────────────────────────────────────────────────────┘ ``` The crawl node fetches the web from its own public IP (51.161.112.66); the data node never crawls unless the `maintenance` profile is enabled (maintenance queue only). Both nodes are Linux and are **not** mld/agentd nodes. ## Ports | Where | Bind | Port | Service | Reachable from | |---|---|---|---|---| | BHS128 | 10.67.0.60 (wg1) | 8300 | edge Caddy → web/api | BHS64 gateway only | | BHS128 | 10.68.0.1 (dci0) | 5432 | Postgres 17 | BHS64b | | BHS128 | 10.68.0.1 | 8123 / 9009 | ClickHouse HTTP / native (host 9009 → container 9000) | BHS64b | | BHS128 | 10.68.0.1 | 6379 | Redis (requirepass) | BHS64b | | BHS128 | 10.68.0.1 | 9000 / 9001 | MinIO S3 / console | BHS64b, SSH tunnel | | BHS128 | 10.68.0.1 | 3000 | Grafana | SSH tunnel (`tunnel-grafana.sh`) | | BHS128 | 10.68.0.1 | 9090 | Prometheus | SSH tunnel | | BHS128 | docker net | 8310 / 8311 / 9180 | web / api / caddy metrics | containers only | | BHS64b | 10.68.0.2 (dci0) | 8320 (8322) | worker /healthz /metrics | Prometheus on BHS128 | | BHS64b | 10.68.0.2 | 9100 | node-exporter | Prometheus on BHS128 | MinIO S3 owns `10.68.0.1:9000`; the ClickHouse native protocol is published on `10.68.0.1:9009` (container 9000). The application only uses the ClickHouse HTTP interface (8123); native is for `clickhouse-client`. ## Files ``` deploy/ docker/Dockerfile.web|api|worker multi-stage, node:22-bookworm-slim, pnpm 11 (pnpm fetch → offline filtered install), non-root, tini, HEALTHCHECK docker/entrypoint.sh role by first CMD word: worker | scheduler | cli | migrate | api docker/ensure-clickhouse.ts ClickHouse DB + DDL (run by `migrate`) compose.data.yml / compose.crawl.yml .env.data.example / .env.crawl.example → copy to .env.data / .env.crawl (gitignored, never committed) edge/Caddyfile :8300 → api/web, security + cache headers, zstd/gzip, metrics :9180 postgres/initdb/01-extensions.sql pg_trgm, unaccent, pg_stat_statements (first init only) clickhouse/config.d, users.d memory caps, backups path, prometheus endpoint; users default (loopback) + dci minio/init.sh, policy-dci-raw.json bucket dci-raw + least-privilege app user monitoring/prometheus.yml, alerts.yml, grafana/… provisioned datasource + "DataCenterIndex — overview" systemd/dci-backup.{service,timer} daily 03:30 on the data node (installed by deploy.sh data) bin/ build.sh · deploy.sh · status.sh · logs.sh · backup.sh · restore.sh · psql.sh · clickhouse.sh · tunnel-grafana.sh bin/remote/backup-run.sh runs on BHS128 (timer + backup.sh) dev-compose.yml laptop only (ClickHouse + MinIO) ``` Images are built **on the target host** (`docker compose build` after an rsync of the repo to `/srv/dci/app`, excluding `node_modules`, `.next`, `data`, `qa`, `.git`, env files). The laptop never pushes images. BuildKit cache mounts keep the pnpm store and the Next cache between builds. ## First deploy (runbook) Prerequisites already in place: Ubuntu 24.04, Docker 29 + compose v5, `wg1` (10.67.0.60) and `dci0` (10.68.0.1 / 10.68.0.2) up, ssh aliases `BHS128` and `BHS64b`, `sudo` NOPASSWD. 1. **Secrets** (laptop, once): ```bash cp deploy/.env.data.example deploy/.env.data cp deploy/.env.crawl.example deploy/.env.crawl for v in POSTGRES_PASSWORD CLICKHOUSE_PASSWORD REDIS_PASSWORD MINIO_ROOT_PASSWORD S3_SECRET_KEY DCI_ADMIN_TOKEN GRAFANA_ADMIN_PASSWORD; do echo "$v=$(openssl rand -hex 24)"; done ``` Paste the values into `deploy/.env.data`; copy the **same** POSTGRES/REDIS/CLICKHOUSE/S3 values into `deploy/.env.crawl`, plus `SCRAPFLY_API_KEY` / `FIRECRAWL_API_KEY` there. 2. **DNS** (GoDaddy): `A datacenterindex.io → 51.161.112.61`, `A www → 51.161.112.61` (or CNAME to apex). 3. **Data node**: ```bash deploy/bin/deploy.sh data ``` Does: preflight (checks the WireGuard IPs, sets `net.ipv4.ip_nonlocal_bind=1` so Docker can bind 10.67.0.60/10.68.0.1 even if dockerd starts before WireGuard after a reboot), `/srv/dci` layout, rsync, `compose build`, `up` stores, `minio-init`, `run --rm migrate` (migrations → seed → ClickHouse DDL), `up -d --remove-orphans`, installs `dci-backup.timer`, creates the `~/.ssh/dci-backup` key on BHS128 and authorizes it on BHS64b (`/srv/dci-backups`), then waits for `http://10.67.0.60:8300/api/health`. 4. **Gateway route** (on BHS64, as documented in `cluster-skill/MACLUSTR-TUNNEL.md`; BHS128 is already in the gateway ipmap as 10.67.0.60): ```bash ssh BHS64 'tunnelctl add www.datacenterindex.io BHS128:8300 && tunnelctl redirect datacenterindex.io www.datacenterindex.io' # or from the laptop: mlt add www.datacenterindex.io BHS128:8300 curl -I https://www.datacenterindex.io/api/health ``` 5. **Crawl node**: ```bash deploy/bin/deploy.sh crawl ``` 6. **Verify**: `deploy/bin/status.sh`, `deploy/bin/logs.sh worker`, `deploy/bin/tunnel-grafana.sh` (Grafana → folder DataCenterIndex → overview; all Prometheus targets green). 7. **First connectors**: `ssh BHS64b` then `cd /srv/dci/app && docker compose -f deploy/compose.crawl.yml --env-file deploy/.env.crawl run --rm cli run peeringdb --dry-run --limit 5` (the `cli` service = `pnpm dci …`, running from the crawl IP with production credentials). ## Update (routine deploy) ```bash deploy/bin/deploy.sh data # rsync → build → migrate → up -d (recreates only changed services) → health deploy/bin/deploy.sh crawl # rsync → build worker → up -d → /healthz # faster variants deploy/bin/deploy.sh data --skip-migrate --no-backup-setup deploy/bin/build.sh data web # just rebuild one image (then deploy.sh data --no-build) ``` Behaviour: `compose build` runs before `up -d`, so the old containers keep serving during the build. `up -d` recreates a service only when its image/config changed; `web`/`api` restart in a few seconds and the edge Caddy retries upstreams for 5 s (`lb_try_duration`), which absorbs the restart for most requests. Postgres, ClickHouse, Redis, MinIO are never recreated by an app deploy (their definitions do not change). Migrations are plain SQL, idempotent and tracked in `_dci_migrations`; the `migrate` one-shot is safe to re-run. Run `deploy.sh data` before `deploy.sh crawl` when a migration changes the schema the workers write to. Roll back = redeploy the previous git revision (`git checkout && deploy/bin/deploy.sh data --skip-migrate`). Migrations are forward-only; a schema rollback goes through `restore.sh`. ## Secrets & rotation All secrets live in `deploy/.env.data` and `deploy/.env.crawl` on the laptop (gitignored, `.gitignore` covers `deploy/.env*`) and are copied to `/srv/dci/app/deploy/.env.` (mode 600) by the scripts. Never in git, never sent to the browser (`NEXT_PUBLIC_*` are the only build-time public values). | Secret | Rotate | |---|---| | `POSTGRES_PASSWORD` | `psql.sh -c "ALTER USER dci PASSWORD 'new'"` → update both env files → `deploy.sh data --no-build --skip-migrate` then `deploy.sh crawl --no-build` | | `REDIS_PASSWORD` | update env files → `deploy.sh data --no-build --skip-migrate` (redis restarts with the new requirepass; queued jobs persist in AOF) → `deploy.sh crawl --no-build` | | `CLICKHOUSE_PASSWORD` | update env files → same two deploys (users are XML `from_env`, applied at start) | | MinIO root / `S3_SECRET_KEY` | root: update env → deploy (MinIO reads root creds at start). App user: `mc admin user add local dci ` runs automatically in `minio-init` on the next `up` (it re-adds the user) → update `.env.crawl` → `deploy.sh crawl --no-build` | | `DCI_ADMIN_TOKEN` | update `.env.data` → `deploy.sh data --no-build --skip-migrate` (api restarts) | | `SCRAPFLY_API_KEY` / `FIRECRAWL_API_KEY` | `.env.crawl` → `deploy.sh crawl --no-build` | | `GRAFANA_ADMIN_PASSWORD` | `.env.data` → deploy; Grafana keeps the admin password in its DB — also change it in the UI or `docker compose exec grafana grafana cli admin reset-admin-password ` | ## Backups & restore Nightly `dci-backup.timer` (03:30 America/Toronto, `Persistent=true`) runs `deploy/bin/remote/backup-run.sh` on BHS128; `deploy/bin/backup.sh` runs the same thing on demand. | What | Where (BHS128) | How | |---|---|---| | Postgres (truth) | `/srv/dci/backups/pg/dci-.dump` | `pg_dump -Fc --compress=zstd` | | Connector configs, parsers, migrations, deploy config, git rev | `/srv/dci/backups/config/dci-config-.tar.zst` | tar | | ClickHouse `dci` | `/srv/dci/backups/clickhouse/dci-ch-.zip` | `BACKUP DATABASE … TO File()` (best effort) | | MinIO `dci-raw` (raw bodies) | `/srv/dci/backups/minio/dci-raw/` | `mc mirror --overwrite` (incremental mirror) | Retention: **14 daily + 8 weekly** (Sunday) for pg/config/clickhouse; the MinIO mirror is a single current copy. After pruning, the whole `/srv/dci/backups/` is rsynced over the private link to **`ubuntu@10.68.0.2:/srv/dci-backups/`** (BHS64b) with the dedicated key `~/.ssh/dci-backup` (set up by `deploy.sh data`). Optional third copy on the laptop: `deploy/bin/backup.sh --pull` (latest dump + config → `~/Backups/dci/`) or `--pull-all`. Restore — always into a **separate database first**: ```bash deploy/bin/restore.sh latest # → database dci_restore on BHS128 (prod untouched) deploy/bin/psql.sh -d dci_restore -c 'select count(*) from facilities' deploy/bin/restore.sh latest --swap # stops api/scheduler/workers, dci → dci_old_, dci_restore → dci, restarts deploy/bin/restore.sh ~/Backups/dci/dci-20260911-033000.dump --db dci_check ``` `--swap` is reversible (`ALTER DATABASE … RENAME`) until you drop `dci_old_`. ClickHouse restore: `clickhouse.sh --query "RESTORE DATABASE dci FROM File('/backups/dci-ch-.zip')"` (drop or rename the database first). MinIO: `mc mirror /backup/dci-raw local/dci-raw` from the `minio-init` image with `/srv/dci/backups/minio` mounted at `/backup` (see `backup-run.sh` step 4 for the exact `compose run`). Before a risky change: `deploy/bin/backup.sh` (takes a fresh dump synchronously). ## Monitoring Prometheus (90 d / 20 GB) scrapes: `api:8311/api/metrics` (`dci_api_requests_total`, `dci_api_request_duration_seconds`, cache counters), the crawl workers at `10.68.0.2:8320/metrics` (+ `:8322` for `worker-b`), the edge Caddy, node-exporter on both hosts, cAdvisor, postgres-, redis-exporter, ClickHouse (`:9363`) and MinIO. Rules in `deploy/monitoring/alerts.yml` (target down, API p95 > 1 s, 5xx > 5 %, crawl stalled, premium budget > 90 %, disk < 10 %, memory > 92 %, PG connections > 170) — no Alertmanager; they show as state in Grafana → Alerting. Grafana is bound to **10.68.0.1:3000** (never public). Access: ```bash deploy/bin/tunnel-grafana.sh # → http://localhost:3000 (also :9090 Prometheus, :9001 MinIO console) ``` Login `GRAFANA_ADMIN_USER` / `GRAFANA_ADMIN_PASSWORD` from `.env.data`. Home dashboard = **DataCenterIndex — overview** (crawl rate by outcome, premium credits/h by provider, queue depth, fetch and API latency, requests by status class, CPU/RAM/disk per node, container memory, PG connections/size). Dashboards are provisioned read-only from `deploy/monitoring/grafana/dashboards/`; edit the JSON in the repo and redeploy. Worker metric contract expected by the dashboard/alerts (to be exported by `apps/worker` on `/metrics`): `dci_crawl_fetches_total{connector,level,outcome=ok|not_modified|error|blocked|skipped}`, `dci_crawl_credits_total{provider=scrapfly|firecrawl}`, `dci_crawl_daily_budget{provider}`, `dci_queue_jobs{queue,state=waiting|active|delayed|failed}`, `dci_crawl_fetch_duration_seconds_bucket`. ## Firewall statement Nothing new is exposed publicly. ufw on both servers keeps only `22/tcp`, `51820/udp`, `51821/udp`. Every Docker published port is bound to a WireGuard address (`10.67.0.60` for the edge, `10.68.0.1` / `10.68.0.2` for stores and metrics), so the iptables DNAT rules Docker installs only match traffic already inside the tunnels; packets to `51.161.112.69:` or `51.161.112.66:` do not match and are dropped by ufw. Inter-container traffic stays on the compose network `dci_default`. The data node has no browser-facing service other than the edge, and the edge answers `404` to `/api/metrics`. TLS terminates on the gateway (BHS64, Let's Encrypt via Caddy); the edge trusts `X-Forwarded-*` from `10.67.0.0/24` only. Docker bypasses ufw for published ports by design — that is why binding to the tunnel IPs (not `0.0.0.0`) is the invariant to preserve when editing the compose files. ## Gateway (BHS64) commands ```bash # create the public route (once) — BHS128 = 10.67.0.60 in the gateway ipmap ssh BHS64 tunnelctl add www.datacenterindex.io BHS128:8300 ssh BHS64 tunnelctl redirect datacenterindex.io www.datacenterindex.io ssh BHS64 tunnelctl status # routes + WireGuard peers # from the laptop wrapper ~/Desktop/cluster-skill/mlt add www.datacenterindex.io BHS128:8300 ``` ## Troubleshooting | Symptom | Check / fix | |---|---| | `deploy.sh data` fails at `up`: *cannot assign requested address* on 10.67.0.60 or 10.68.0.1 | WireGuard not up or IPs changed: `ssh BHS128 ip -4 a show wg1; ip -4 a show dci0`. deploy.sh sets `net.ipv4.ip_nonlocal_bind=1` so this only happens if the sysctl was lost — `sudo sysctl -p /etc/sysctl.d/90-dci.conf` | | 502 from the public URL, edge healthy | gateway route missing/stale → `ssh BHS64 tunnelctl status`; `ssh BHS64 curl -sI http://10.67.0.60:8300/api/health` | | edge healthy, `/api/health` 503 | api not ready: `logs.sh api` (DB URL, migrations). `compose run --rm migrate` again | | web 500 / blank | `logs.sh web`; the container must have `API_URL_INTERNAL=http://api:8311`; `NEXT_PUBLIC_API_URL` is baked at build → rebuild web if it changed | | workers idle, queue growing | `logs.sh worker`; from BHS64b: `nc -zv 10.68.0.1 6379 5432 8123 9000` (dci0 down?), Redis password mismatch between the two env files | | workers fail on S3 `AccessDenied` | `minio-init` did not run (check `compose ps -a minio-init`, rerun `up -d minio-init`), or `S3_ACCESS_KEY/S3_SECRET_KEY` differ between env files | | ClickHouse auth errors | `CLICKHOUSE_USER=dci` + password in both env files; the `default` user is loopback-only by design | | Prometheus target `10.68.0.2:8320` down | crawl worker not running or `WORKER_PORT` changed; `ssh BHS64b curl -s http://10.68.0.2:8320/healthz` | | Grafana shows no data | `tunnel-grafana.sh` → http://localhost:9090/targets; metric names must match the contract above | | Disk pressure on BHS128 | `status.sh data`; largest consumers: `/var/lib/docker/volumes/dci_miniodata`, `dci_pgdata`, `/srv/dci/backups/minio`. Prune old images: `docker image prune -af --filter until=168h` | | After a reboot nothing listens | `systemctl status docker`; `compose ps`; containers are `restart: unless-stopped` so they return once dockerd is up (WireGuard first thanks to `ip_nonlocal_bind`) | | Backup timer not running | `ssh BHS128 systemctl list-timers dci-backup.timer; journalctl -u dci-backup.service -n 50`; off-node rsync needs `~/.ssh/dci-backup.pub` in `ubuntu@BHS64b:~/.ssh/authorized_keys` | | Need a shell in a container | `ssh BHS128 'cd /srv/dci/app && docker compose -f deploy/compose.data.yml --env-file deploy/.env.data exec api sh'` | | Rebuild from scratch (no cache) | `ssh BHS128 'cd /srv/dci/app && docker compose -f deploy/compose.data.yml --env-file deploy/.env.data build --no-cache web'` | Local validation without servers: `docker compose -f deploy/compose.data.yml --env-file deploy/.env.data.example config -q` and `docker build -f deploy/docker/Dockerfile.worker .` from the repo root. ## 2026-09-12 upgrade — post-deploy data repair and gotchas - Deploy order when the API contract changes: the web image prerenders `/` and other static routes against the **live** API (`NEXT_PUBLIC_API_URL` build arg). Either keep pages tolerant of missing fields (done for the home page) or build/up `api scheduler worker-maint` first, then `web` (see `/tmp/dci-redeploy.sh` pattern: `compose build api …`, `run --rm migrate`, `up -d api …`, then `compose build web`, `up -d web edge`). - `deploy/rsync-exclude.txt` and `.dockerignore` must exclude only the root `coverage/` report directory (`/coverage`): a bare `coverage` pattern silently dropped the `/coverage` web route from the image. - `compose run --rm cli ` passes `` as the entrypoint ROLE — run `dci` commands as `run --rm -T cli cli `. - `scripts/` is not baked into the worker image; run repair scripts with bind mounts: `run --rm -T -v /srv/dci/app/scripts:/app/scripts -v /srv/dci/app/packages/core/src:/app/packages/core/src cli node /app/node_modules/tsx/dist/cli.mjs /app/scripts/quality-apply.ts …`. - `deploy/bin/post-upgrade.sh [baseline|repair|reprocess|unbacked|refresh|audit]` runs the whole repair sequence from the laptop: quality sweep + snapshot, hide vetoed / unconfirmed projects, fix slugs, link campuses, `dci reprocess --stale` for every connector (offline, archived bodies), null figures no site-scoped claim backs, refresh stats + rankings, print the largest values. - API/edge health probes use `/api/ready` (Postgres-aware). `DCI_WORKER_URL=http://worker-maint:8320` lets the API proxy the extraction debugger (`/api/admin/documents/:id/trace`) and `/api/admin/data-gaps`.