spb/datacenterindex
Public
HTML 53.9%
TypeScript 44.5%
JavaScript 0.6%
SQL 0.5%
1# Deployment — DataCenterIndex.io23Two OVHcloud dedicated servers in Beauharnois, joined by a private WireGuard link. Public traffic enters4through the existing **MacLustr Tunnel** gateway (BHS64, Caddy + WireGuard). **No new public port is opened5on either server**: the only inbound rules stay `22/tcp`, `51820/udp`, `51821/udp` (ufw). Everything in6`deploy/` is driven from the laptop with the scripts in `deploy/bin/`.78## Topology910```11 Internet12 │ https://www.datacenterindex.io (DNS A → 51.161.112.61)13 ▼14 ┌─────────────────────────────────────────────┐15 │ BHS64 — gateway (Caddy TLS, WireGuard hub) │ tunnelctl add www.datacenterindex.io BHS128:830016 │ wg0 10.67.0.1 │ tunnelctl redirect datacenterindex.io www.datacenterindex.io17 └───────────────┬─────────────────────────────┘18 │ wg1 10.67.0.1 → 10.67.0.60:8300 (plain HTTP inside the tunnel)19 ▼20 ┌───────────────────────────────────────────────────────────────────────────────┐21 │ BHS128 — DATA node (ubuntu@51.161.112.69, 6c/12t, 128 GB, 2×960 GB NVMe R1) │22 │ docker compose -p dci (deploy/compose.data.yml) │23 │ │24 │ edge caddy :8300 ── /api/* ──▶ api (Fastify :8311) │25 │ published 10.67.0.60:8300 └── /* ──▶ web (Next :8310) │26 │ scheduler (BullMQ producer) [worker-maint: profile maintenance] │27 │ postgres:17 clickhouse:25.8 redis:7 minio ── published 10.68.0.1 │28 │ prometheus grafana node-exporter cadvisor pg/redis exporters │29 │ /srv/dci/app (repo) /srv/dci/backups dci-backup.timer 03:30 │30 └───────────────┬───────────────────────────────────────────────────────────────┘31 │ dci0 10.68.0.1 ⇄ 10.68.0.2 (private WireGuard /30)32 ▼33 ┌───────────────────────────────────────────────────────────────────────────────┐34 │ BHS64b — CRAWL node (ubuntu@51.161.112.66, 8c/16t, 64 GB) │35 │ docker compose -p dci (deploy/compose.crawl.yml) │36 │ worker (crawl+ingest, DCI_CRAWL_CONCURRENCY=6) /healthz+/metrics 10.68.0.2:832037 │ node-exporter 10.68.0.2:9100 [worker-b: profile scale → :8322] │38 │ /srv/dci-backups ← nightly rsync of the data node's backups │39 └───────────────────────────────────────────────────────────────────────────────┘40```4142The crawl node fetches the web from its own public IP (51.161.112.66); the data node never crawls unless the43`maintenance` profile is enabled (maintenance queue only). Both nodes are Linux and are **not** mld/agentd nodes.4445## Ports4647| Where | Bind | Port | Service | Reachable from |48|---|---|---|---|---|49| BHS128 | 10.67.0.60 (wg1) | 8300 | edge Caddy → web/api | BHS64 gateway only |50| BHS128 | 10.68.0.1 (dci0) | 5432 | Postgres 17 | BHS64b |51| BHS128 | 10.68.0.1 | 8123 / 9009 | ClickHouse HTTP / native (host 9009 → container 9000) | BHS64b |52| BHS128 | 10.68.0.1 | 6379 | Redis (requirepass) | BHS64b |53| BHS128 | 10.68.0.1 | 9000 / 9001 | MinIO S3 / console | BHS64b, SSH tunnel |54| BHS128 | 10.68.0.1 | 3000 | Grafana | SSH tunnel (`tunnel-grafana.sh`) |55| BHS128 | 10.68.0.1 | 9090 | Prometheus | SSH tunnel |56| BHS128 | docker net | 8310 / 8311 / 9180 | web / api / caddy metrics | containers only |57| BHS64b | 10.68.0.2 (dci0) | 8320 (8322) | worker /healthz /metrics | Prometheus on BHS128 |58| BHS64b | 10.68.0.2 | 9100 | node-exporter | Prometheus on BHS128 |5960MinIO S3 owns `10.68.0.1:9000`; the ClickHouse native protocol is published on `10.68.0.1:9009` (container619000). The application only uses the ClickHouse HTTP interface (8123); native is for `clickhouse-client`.6263## Files6465```66deploy/67 docker/Dockerfile.web|api|worker multi-stage, node:22-bookworm-slim, pnpm 11 (pnpm fetch → offline filtered install), non-root, tini, HEALTHCHECK68 docker/entrypoint.sh role by first CMD word: worker | scheduler | cli | migrate | api69 docker/ensure-clickhouse.ts ClickHouse DB + DDL (run by `migrate`)70 compose.data.yml / compose.crawl.yml71 .env.data.example / .env.crawl.example → copy to .env.data / .env.crawl (gitignored, never committed)72 edge/Caddyfile :8300 → api/web, security + cache headers, zstd/gzip, metrics :918073 postgres/initdb/01-extensions.sql pg_trgm, unaccent, pg_stat_statements (first init only)74 clickhouse/config.d, users.d memory caps, backups path, prometheus endpoint; users default (loopback) + dci75 minio/init.sh, policy-dci-raw.json bucket dci-raw + least-privilege app user76 monitoring/prometheus.yml, alerts.yml, grafana/… provisioned datasource + "DataCenterIndex — overview"77 systemd/dci-backup.{service,timer} daily 03:30 on the data node (installed by deploy.sh data)78 bin/ build.sh · deploy.sh · status.sh · logs.sh · backup.sh · restore.sh · psql.sh · clickhouse.sh · tunnel-grafana.sh79 bin/remote/backup-run.sh runs on BHS128 (timer + backup.sh)80 dev-compose.yml laptop only (ClickHouse + MinIO)81```8283Images are built **on the target host** (`docker compose build` after an rsync of the repo to84`/srv/dci/app`, excluding `node_modules`, `.next`, `data`, `qa`, `.git`, env files). The laptop never pushes85images. BuildKit cache mounts keep the pnpm store and the Next cache between builds.8687## First deploy (runbook)8889Prerequisites already in place: Ubuntu 24.04, Docker 29 + compose v5, `wg1` (10.67.0.60) and `dci0`90(10.68.0.1 / 10.68.0.2) up, ssh aliases `BHS128` and `BHS64b`, `sudo` NOPASSWD.91921. **Secrets** (laptop, once):93 ```bash94 cp deploy/.env.data.example deploy/.env.data95 cp deploy/.env.crawl.example deploy/.env.crawl96 for v in POSTGRES_PASSWORD CLICKHOUSE_PASSWORD REDIS_PASSWORD MINIO_ROOT_PASSWORD S3_SECRET_KEY DCI_ADMIN_TOKEN GRAFANA_ADMIN_PASSWORD; do echo "$v=$(openssl rand -hex 24)"; done97 ```98 Paste the values into `deploy/.env.data`; copy the **same** POSTGRES/REDIS/CLICKHOUSE/S3 values into99 `deploy/.env.crawl`, plus `SCRAPFLY_API_KEY` / `FIRECRAWL_API_KEY` there.1002. **DNS** (GoDaddy): `A datacenterindex.io → 51.161.112.61`, `A www → 51.161.112.61` (or CNAME to apex).1013. **Data node**:102 ```bash103 deploy/bin/deploy.sh data104 ```105 Does: preflight (checks the WireGuard IPs, sets `net.ipv4.ip_nonlocal_bind=1` so Docker can bind106 10.67.0.60/10.68.0.1 even if dockerd starts before WireGuard after a reboot), `/srv/dci` layout, rsync,107 `compose build`, `up` stores, `minio-init`, `run --rm migrate` (migrations → seed → ClickHouse DDL),108 `up -d --remove-orphans`, installs `dci-backup.timer`, creates the `~/.ssh/dci-backup` key on BHS128 and109 authorizes it on BHS64b (`/srv/dci-backups`), then waits for `http://10.67.0.60:8300/api/health`.1104. **Gateway route** (on BHS64, as documented in `cluster-skill/MACLUSTR-TUNNEL.md`; BHS128 is already in111 the gateway ipmap as 10.67.0.60):112 ```bash113 ssh BHS64 'tunnelctl add www.datacenterindex.io BHS128:8300 && tunnelctl redirect datacenterindex.io www.datacenterindex.io'114 # or from the laptop: mlt add www.datacenterindex.io BHS128:8300115 curl -I https://www.datacenterindex.io/api/health116 ```1175. **Crawl node**:118 ```bash119 deploy/bin/deploy.sh crawl120 ```1216. **Verify**: `deploy/bin/status.sh`, `deploy/bin/logs.sh worker`, `deploy/bin/tunnel-grafana.sh`122 (Grafana → folder DataCenterIndex → overview; all Prometheus targets green).1237. **First connectors**: `ssh BHS64b` then124 `cd /srv/dci/app && docker compose -f deploy/compose.crawl.yml --env-file deploy/.env.crawl run --rm cli run peeringdb --dry-run --limit 5`125 (the `cli` service = `pnpm dci …`, running from the crawl IP with production credentials).126127## Update (routine deploy)128129```bash130deploy/bin/deploy.sh data # rsync → build → migrate → up -d (recreates only changed services) → health131deploy/bin/deploy.sh crawl # rsync → build worker → up -d → /healthz132# faster variants133deploy/bin/deploy.sh data --skip-migrate --no-backup-setup134deploy/bin/build.sh data web # just rebuild one image (then deploy.sh data --no-build)135```136137Behaviour: `compose build` runs before `up -d`, so the old containers keep serving during the build. `up -d`138recreates a service only when its image/config changed; `web`/`api` restart in a few seconds and the edge139Caddy retries upstreams for 5 s (`lb_try_duration`), which absorbs the restart for most requests. Postgres,140ClickHouse, Redis, MinIO are never recreated by an app deploy (their definitions do not change). Migrations141are plain SQL, idempotent and tracked in `_dci_migrations`; the `migrate` one-shot is safe to re-run. Run142`deploy.sh data` before `deploy.sh crawl` when a migration changes the schema the workers write to.143144Roll back = redeploy the previous git revision (`git checkout <rev> && deploy/bin/deploy.sh data --skip-migrate`).145Migrations are forward-only; a schema rollback goes through `restore.sh`.146147## Secrets & rotation148149All secrets live in `deploy/.env.data` and `deploy/.env.crawl` on the laptop (gitignored, `.gitignore`150covers `deploy/.env*`) and are copied to `/srv/dci/app/deploy/.env.<target>` (mode 600) by the scripts.151Never in git, never sent to the browser (`NEXT_PUBLIC_*` are the only build-time public values).152153| Secret | Rotate |154|---|---|155| `POSTGRES_PASSWORD` | `psql.sh -c "ALTER USER dci PASSWORD 'new'"` → update both env files → `deploy.sh data --no-build --skip-migrate` then `deploy.sh crawl --no-build` |156| `REDIS_PASSWORD` | update env files → `deploy.sh data --no-build --skip-migrate` (redis restarts with the new requirepass; queued jobs persist in AOF) → `deploy.sh crawl --no-build` |157| `CLICKHOUSE_PASSWORD` | update env files → same two deploys (users are XML `from_env`, applied at start) |158| MinIO root / `S3_SECRET_KEY` | root: update env → deploy (MinIO reads root creds at start). App user: `mc admin user add local dci <new>` runs automatically in `minio-init` on the next `up` (it re-adds the user) → update `.env.crawl` → `deploy.sh crawl --no-build` |159| `DCI_ADMIN_TOKEN` | update `.env.data` → `deploy.sh data --no-build --skip-migrate` (api restarts) |160| `SCRAPFLY_API_KEY` / `FIRECRAWL_API_KEY` | `.env.crawl` → `deploy.sh crawl --no-build` |161| `GRAFANA_ADMIN_PASSWORD` | `.env.data` → deploy; Grafana keeps the admin password in its DB — also change it in the UI or `docker compose exec grafana grafana cli admin reset-admin-password <new>` |162163## Backups & restore164165Nightly `dci-backup.timer` (03:30 America/Toronto, `Persistent=true`) runs `deploy/bin/remote/backup-run.sh`166on BHS128; `deploy/bin/backup.sh` runs the same thing on demand.167168| What | Where (BHS128) | How |169|---|---|---|170| Postgres (truth) | `/srv/dci/backups/pg/dci-<stamp>.dump` | `pg_dump -Fc --compress=zstd` |171| Connector configs, parsers, migrations, deploy config, git rev | `/srv/dci/backups/config/dci-config-<stamp>.tar.zst` | tar |172| ClickHouse `dci` | `/srv/dci/backups/clickhouse/dci-ch-<stamp>.zip` | `BACKUP DATABASE … TO File()` (best effort) |173| MinIO `dci-raw` (raw bodies) | `/srv/dci/backups/minio/dci-raw/` | `mc mirror --overwrite` (incremental mirror) |174175Retention: **14 daily + 8 weekly** (Sunday) for pg/config/clickhouse; the MinIO mirror is a single current176copy. After pruning, the whole `/srv/dci/backups/` is rsynced over the private link to177**`ubuntu@10.68.0.2:/srv/dci-backups/`** (BHS64b) with the dedicated key `~/.ssh/dci-backup` (set up by178`deploy.sh data`). Optional third copy on the laptop: `deploy/bin/backup.sh --pull` (latest dump + config179→ `~/Backups/dci/`) or `--pull-all`.180181Restore — always into a **separate database first**:182183```bash184deploy/bin/restore.sh latest # → database dci_restore on BHS128 (prod untouched)185deploy/bin/psql.sh -d dci_restore -c 'select count(*) from facilities'186deploy/bin/restore.sh latest --swap # stops api/scheduler/workers, dci → dci_old_<stamp>, dci_restore → dci, restarts187deploy/bin/restore.sh ~/Backups/dci/dci-20260911-033000.dump --db dci_check188```189190`--swap` is reversible (`ALTER DATABASE … RENAME`) until you drop `dci_old_<stamp>`. ClickHouse restore:191`clickhouse.sh --query "RESTORE DATABASE dci FROM File('/backups/dci-ch-<stamp>.zip')"` (drop or rename the192database first). MinIO: `mc mirror /backup/dci-raw local/dci-raw` from the `minio-init` image with193`/srv/dci/backups/minio` mounted at `/backup` (see `backup-run.sh` step 4 for the exact `compose run`).194195Before a risky change: `deploy/bin/backup.sh` (takes a fresh dump synchronously).196197## Monitoring198199Prometheus (90 d / 20 GB) scrapes: `api:8311/api/metrics` (`dci_api_requests_total`,200`dci_api_request_duration_seconds`, cache counters), the crawl workers at `10.68.0.2:8320/metrics`201(+ `:8322` for `worker-b`), the edge Caddy, node-exporter on both hosts, cAdvisor, postgres-, redis-exporter,202ClickHouse (`:9363`) and MinIO. Rules in `deploy/monitoring/alerts.yml` (target down, API p95 > 1 s, 5xx > 5 %,203crawl stalled, premium budget > 90 %, disk < 10 %, memory > 92 %, PG connections > 170) — no Alertmanager;204they show as state in Grafana → Alerting.205206Grafana is bound to **10.68.0.1:3000** (never public). Access:207208```bash209deploy/bin/tunnel-grafana.sh # → http://localhost:3000 (also :9090 Prometheus, :9001 MinIO console)210```211212Login `GRAFANA_ADMIN_USER` / `GRAFANA_ADMIN_PASSWORD` from `.env.data`. Home dashboard = **DataCenterIndex —213overview** (crawl rate by outcome, premium credits/h by provider, queue depth, fetch and API latency,214requests by status class, CPU/RAM/disk per node, container memory, PG connections/size). Dashboards are215provisioned read-only from `deploy/monitoring/grafana/dashboards/`; edit the JSON in the repo and redeploy.216217Worker metric contract expected by the dashboard/alerts (to be exported by `apps/worker` on `/metrics`):218`dci_crawl_fetches_total{connector,level,outcome=ok|not_modified|error|blocked|skipped}`,219`dci_crawl_credits_total{provider=scrapfly|firecrawl}`, `dci_crawl_daily_budget{provider}`,220`dci_queue_jobs{queue,state=waiting|active|delayed|failed}`, `dci_crawl_fetch_duration_seconds_bucket`.221222## Firewall statement223224Nothing new is exposed publicly. ufw on both servers keeps only `22/tcp`, `51820/udp`, `51821/udp`. Every225Docker published port is bound to a WireGuard address (`10.67.0.60` for the edge, `10.68.0.1` / `10.68.0.2`226for stores and metrics), so the iptables DNAT rules Docker installs only match traffic already inside the227tunnels; packets to `51.161.112.69:<port>` or `51.161.112.66:<port>` do not match and are dropped by ufw.228Inter-container traffic stays on the compose network `dci_default`. The data node has no browser-facing229service other than the edge, and the edge answers `404` to `/api/metrics`. TLS terminates on the gateway230(BHS64, Let's Encrypt via Caddy); the edge trusts `X-Forwarded-*` from `10.67.0.0/24` only.231232Docker bypasses ufw for published ports by design — that is why binding to the tunnel IPs (not `0.0.0.0`)233is the invariant to preserve when editing the compose files.234235## Gateway (BHS64) commands236237```bash238# create the public route (once) — BHS128 = 10.67.0.60 in the gateway ipmap239ssh BHS64 tunnelctl add www.datacenterindex.io BHS128:8300240ssh BHS64 tunnelctl redirect datacenterindex.io www.datacenterindex.io241ssh BHS64 tunnelctl status # routes + WireGuard peers242# from the laptop wrapper243~/Desktop/cluster-skill/mlt add www.datacenterindex.io BHS128:8300244```245246## Troubleshooting247248| Symptom | Check / fix |249|---|---|250| `deploy.sh data` fails at `up`: *cannot assign requested address* on 10.67.0.60 or 10.68.0.1 | WireGuard not up or IPs changed: `ssh BHS128 ip -4 a show wg1; ip -4 a show dci0`. deploy.sh sets `net.ipv4.ip_nonlocal_bind=1` so this only happens if the sysctl was lost — `sudo sysctl -p /etc/sysctl.d/90-dci.conf` |251| 502 from the public URL, edge healthy | gateway route missing/stale → `ssh BHS64 tunnelctl status`; `ssh BHS64 curl -sI http://10.67.0.60:8300/api/health` |252| edge healthy, `/api/health` 503 | api not ready: `logs.sh api` (DB URL, migrations). `compose run --rm migrate` again |253| web 500 / blank | `logs.sh web`; the container must have `API_URL_INTERNAL=http://api:8311`; `NEXT_PUBLIC_API_URL` is baked at build → rebuild web if it changed |254| workers idle, queue growing | `logs.sh worker`; from BHS64b: `nc -zv 10.68.0.1 6379 5432 8123 9000` (dci0 down?), Redis password mismatch between the two env files |255| workers fail on S3 `AccessDenied` | `minio-init` did not run (check `compose ps -a minio-init`, rerun `up -d minio-init`), or `S3_ACCESS_KEY/S3_SECRET_KEY` differ between env files |256| ClickHouse auth errors | `CLICKHOUSE_USER=dci` + password in both env files; the `default` user is loopback-only by design |257| Prometheus target `10.68.0.2:8320` down | crawl worker not running or `WORKER_PORT` changed; `ssh BHS64b curl -s http://10.68.0.2:8320/healthz` |258| Grafana shows no data | `tunnel-grafana.sh` → http://localhost:9090/targets; metric names must match the contract above |259| Disk pressure on BHS128 | `status.sh data`; largest consumers: `/var/lib/docker/volumes/dci_miniodata`, `dci_pgdata`, `/srv/dci/backups/minio`. Prune old images: `docker image prune -af --filter until=168h` |260| After a reboot nothing listens | `systemctl status docker`; `compose ps`; containers are `restart: unless-stopped` so they return once dockerd is up (WireGuard first thanks to `ip_nonlocal_bind`) |261| Backup timer not running | `ssh BHS128 systemctl list-timers dci-backup.timer; journalctl -u dci-backup.service -n 50`; off-node rsync needs `~/.ssh/dci-backup.pub` in `ubuntu@BHS64b:~/.ssh/authorized_keys` |262| Need a shell in a container | `ssh BHS128 'cd /srv/dci/app && docker compose -f deploy/compose.data.yml --env-file deploy/.env.data exec api sh'` |263| Rebuild from scratch (no cache) | `ssh BHS128 'cd /srv/dci/app && docker compose -f deploy/compose.data.yml --env-file deploy/.env.data build --no-cache web'` |264265Local validation without servers: `docker compose -f deploy/compose.data.yml --env-file deploy/.env.data.example config -q`266and `docker build -f deploy/docker/Dockerfile.worker .` from the repo root.267268## 2026-09-12 upgrade — post-deploy data repair and gotchas269270- Deploy order when the API contract changes: the web image prerenders `/` and other static routes against the **live** API271 (`NEXT_PUBLIC_API_URL` build arg). Either keep pages tolerant of missing fields (done for the home page) or build/up272 `api scheduler worker-maint` first, then `web` (see `/tmp/dci-redeploy.sh` pattern: `compose build api …`, `run --rm migrate`,273 `up -d api …`, then `compose build web`, `up -d web edge`).274- `deploy/rsync-exclude.txt` and `.dockerignore` must exclude only the root `coverage/` report directory (`/coverage`): a bare275 `coverage` pattern silently dropped the `/coverage` web route from the image.276- `compose run --rm cli <cmd>` passes `<cmd>` as the entrypoint ROLE — run `dci` commands as `run --rm -T cli cli <cmd>`.277- `scripts/` is not baked into the worker image; run repair scripts with bind mounts:278 `run --rm -T -v /srv/dci/app/scripts:/app/scripts -v /srv/dci/app/packages/core/src:/app/packages/core/src cli node /app/node_modules/tsx/dist/cli.mjs /app/scripts/quality-apply.ts …`.279- `deploy/bin/post-upgrade.sh [baseline|repair|reprocess|unbacked|refresh|audit]` runs the whole repair sequence from the laptop:280 quality sweep + snapshot, hide vetoed / unconfirmed projects, fix slugs, link campuses, `dci reprocess <id> --stale` for every281 connector (offline, archived bodies), null figures no site-scoped claim backs, refresh stats + rankings, print the largest values.282- API/edge health probes use `/api/ready` (Postgres-aware). `DCI_WORKER_URL=http://worker-maint:8320` lets the API proxy the283 extraction debugger (`/api/admin/documents/:id/trace`) and `/api/admin/data-gaps`.284