SPB Git forge
38commits 1branches 0releases
338.7 MBsize
maindefault branch
3 h agolast push
HTML 53.9% TypeScript 44.5% JavaScript 0.6% SQL 0.5%
16.0 KB · 166 lines markdown
Rendered Raw Blame History
1# Crawl operations — DataCenterIndex.io23> **2026-09-12 additions.** Connector health now has `blocked`, `schema_change`, `no_new_content` states and a `quarantine` flag (`dci quarantine <id> on|off`, auto-set after 5 consecutive failed/blocked runs or a non-full run creating ≥ 500 records — everything runs, nothing is published, `connector_runs.quarantined = true`). The effective extractor version is the YAML `parserVersion` plus the versions of the parsers actually used: bump a parser `version` and `dci reprocess <id> --stale` re-extracts only the documents extracted with an older version. `dci trace <doc>` shows every stage for one document; `dci quality` / `dci snapshot` / `dci gaps` run the data-quality jobs (also scheduled: quality every 6 h, snapshot + regression checks daily 00:45 UTC). 429 responses never escalate and honour `retry-after`; robots.txt failures fail closed to L1/L2; `noarchive` pages are never archived.45Runbook for the crawl runtime (`apps/worker`) in production: the crawl worker on **BHS64b**6(`compose.crawl.yml`, service `worker`, `/healthz` + `/metrics` on `10.68.0.2:8320`), the dedicated7**scheduler** and the maintenance-only worker on **BHS128** (`compose.data.yml`, services `scheduler` :8321 and8`worker-maint`), Postgres/Redis/ClickHouse/MinIO on BHS128. Deployment itself is in `docs/DEPLOY.md`; the9runtime internals in `apps/worker/README.md`.1011Shell shortcuts used below:1213```bash14alias dcd='ssh BHS128 "cd /srv/dci/app && docker compose -f deploy/compose.data.yml --env-file deploy/.env.data"'   # data node15alias dcc='ssh BHS64b "cd /srv/dci/app && docker compose -f deploy/compose.crawl.yml --env-file deploy/.env.crawl"'  # crawl node16alias dpsql='ssh BHS128 "cd /srv/dci/app && docker compose -f deploy/compose.data.yml --env-file deploy/.env.data exec -T postgres psql -U dci -d dci"'17```1819## 1. Daily checks (5 minutes)2021| # | Check | Command / where | Expected |22|---|---|---|---|23| 1 | Containers | `deploy/bin/status.sh` | everything `Up … (healthy)`; **scheduler healthy** (it was in a restart loop before 2026-09-11) |24| 2 | Worker liveness | `ssh BHS64b curl -s 10.68.0.2:8320/healthz \| jq '{ok,runningJobs,creditsToday,sharedBudgets,lastScheduler,queueDepth}'` | `ok: true`, `lastScheduler.at` < 2 min old |25| 3 | Scheduler liveness | `dcd exec -T scheduler node -e "fetch('http://127.0.0.1:8321/healthz').then(r=>r.text()).then(console.log)"` | `ok: true`, `lastTick.error: null` |26| 4 | Queue depth | Grafana → *Queue depth*, or `dci_queue_jobs{queue="crawl",state="waiting"}` | near 0 between ticks; a backlog > 50 for 30 min fires `QueueBacklog` |27| 5 | Premium credits | Grafana → *Premium budget used today*; `ssh BHS128 … redis-cli … MGET dci:budget:scrapfly:$(date -u +%F) dci:budget:firecrawl:$(date -u +%F)` | well below `DCI_*_DAILY_BUDGET` (400 / 200); know **which connector** spent them (§4) |28| 6 | Runs last 24 h | `dpsql -c "select status, count(*) from connector_runs where started_at > now() - interval '24 hours' group by 1"` | mostly `ok`; investigate `failed` / `aborted` |29| 7 | Failing / degraded connectors | `dpsql -c "select id, health, last_status, left(last_error,80) from connectors where enabled and not paused and (health <> 'ok' or last_status in ('failed','partial'))"` | empty or known |30| 8 | Doctor | `dcc run --rm cli doctor` | all critical checks pass; warn lines are triage input |31| 9 | Alerts | Grafana → Alerting | none firing (`TargetDown`, `CrawlStalled`, `SchedulerStale`, `WorkerDown`, `PremiumBudgetNearlyExhausted`, `FetchErrorRateHigh`, `CrawlRunsFailing`, `QueueBacklog`, disk/memory) |32| 10 | Backups | `ssh BHS128 systemctl list-timers dci-backup.timer`; `ssh BHS64b du -sh /srv/dci-backups/*` | timer waiting for 03:30, mirror grew last night |3334Weekly: `dpsql -c "select count(*) from documents where quarantined"` (should grow slowly, not jump), `select count(*)35from entity_matches where status='pending'`, disk on both nodes (`status.sh`).3637## 2. Reading the metrics3839Both processes export the same registry (`apps/worker/src/prom.ts`); Prometheus scrapes `dci-worker` (BHS64b40:8320/:8322 and `worker-maint:8320`) and `dci-scheduler` (`scheduler:8321`).4142| Metric | Meaning | Typical query |43|---|---|---|44| `dci_crawl_fetches_total{connector,level,outcome}` | one sample per fetch that returned (outcome `ok`, `not_modified`, `error`, `blocked`; level = final level `L1`–`L4`) | `sum by (outcome) (rate(…[5m]))`; per connector error ratio `sum by (connector) (rate(…{outcome=~"error\|blocked"}[1h])) / sum by (connector) (rate(…[1h]))` |45| `dci_crawl_fetch_duration_seconds` (histogram) | wall time of one fetch incl. escalation | `histogram_quantile(0.95, sum(rate(…_bucket[5m])) by (le))` |46| `dci_crawl_credits_total{provider}` | premium credits spent since process start | `sum by (provider) (increase(…[1h]))` |47| `dci_crawl_daily_budget{provider}` / `dci_crawl_daily_credits_used{provider}` | daily cap and **shared** usage today (Redis, survives restarts) | `used / budget` → the `PremiumBudgetNearlyExhausted` rule |48| `dci_ingest_entities_total{connector,result}` | entities handed to ingest: `created`, `updated`, `merged`, `unchanged`, `rejected` | `sum by (result) (rate(…[5m]))` |49| `dci_events_total` | change events emitted by ingest | `increase(…[24h])` |50| `dci_crawl_runs_total{status}` | runs finished by status (`ok`, `partial`, `failed`, `aborted`) | `CrawlRunsFailing` |51| `dci_worker_jobs_total{queue,result}` | BullMQ jobs: `completed`, `failed`, `stalled`, `deferred` (run lock held by another worker) | many `deferred` = two workers keep colliding on the same connectors |52| `dci_queue_jobs{queue,state}` | snapshot of queue depth per state at scrape time | `state="waiting"` / `"active"` / `"failed"` |53| `dci_worker_running_jobs`, `dci_worker_up`, `dci_worker_uptime_seconds` | per process | `WorkerDown` |54| `dci_scheduler_ticks_total{result}`, `dci_scheduler_enqueued_total`, `dci_scheduler_last_tick_timestamp_seconds` | scheduler health | `time() - max(last_tick) > 600` → `SchedulerStale` |5556Where the counters live outside Prometheus: `connector_runs.stats` (per run), `connectors.stats` (per connector,57refreshed after each run), ClickHouse `crawl_log` (one row per fetch, 400 days) and `page_changes`.5859## 3. Triaging a failing or degraded connector6061`connectors.health` is `degraded` when ≥ 20 % of a run's fetches failed, `failing` at ≥ 50 % or when the run62threw. Start from the last run:6364```bash65dpsql -c "select id, task, status, started_at, stats->>'fetched' f, stats->>'failed' fl, stats->>'credits' cr, left(error,120) from connector_runs where connector_id='<id>' order by started_at desc limit 5"66dpsql -c "select l->>'t' t, l->>'level' lvl, l->>'msg' from connector_runs r, jsonb_array_elements(r.log) l where r.id='<run_id>' and l->>'level' in ('warn','error') limit 40"67dpsql -c "select split_part(error,':',1) code, count(*) from documents where connector_id='<id>' and error_count>0 group by 1 order by 2 desc"68dcc run --rm cli docs <id> --due --limit 20        # what is about to be fetched69dcc run --rm cli run <id> --dry-run --limit 3 --verbose   # live reproduction from the crawl IP, nothing persisted70```7172Common signatures:7374| Symptom | Cause | Action |75|---|---|---|76| `HTTP 403` on every page, credits > 0 | bot wall; Scrapfly ASP also refused | after 2 consecutive failures the runtime stops escalating (`premiumAllowedAfterErrors`), so the cost stops by itself. Fix the connector: `fetch.userAgent: browser`, `renderJs: false` (cheaper), or find a feed / sitemap that is open |77| `rss … → 403` at discovery, `discovered 0 urls` | feed blocked for the bot identity (discovery is direct-only, never premium) | connector: `fetch.userAgent: browser` or another feed URL. The scheduler retries discovery on its cadence (≥ 6 h), never every tick |78| `discovered 0 urls` with a working feed (`filter: 0/20 items match`) | connector filter too narrow | connector code / YAML; not a runtime issue |79| `robots.txt disallows` | legal stop — leave it | remove the path from discovery |80| `fetch_failed`, `timeout`, `Headers Timeout` | upstream slow (Overpass, big PDFs) | raise `fetch.timeoutMs` for the connector or lower `fetch.concurrency`; documents back off ×2^(n−1) automatically |81| `partial` with validation errors only | extractor problem (`implausible MW`, `missing name`) | parser bug: file/line in the run log; data is **not** ingested for rejected entities |82| every refetch `CHANGED` with `diff_summary` null, versions 2 = fetch_count | two runs on the same connector overlapped (fixed 2026-09-11 by the per-connector run lock and the scheduler skipping group jobs while a `full` job is pending) | should not recur; if it does, check `dci_worker_jobs_total{result="deferred"}` and `dci:run-lock:*` in Redis |83| every refetch `CHANGED` with a small `ratio` (< 0.1) and sidebars in `added/removed` | page embeds "related articles" / ads: the fingerprint moves although the article did not | connector: extract main content only (`each`/`match` on the article container); the `change_frequency_score` will keep re-checking such pages twice as often until fixed |84| connector `never_run`, `enabled = f` | disabled in YAML | intentional (17 operator connectors are parked) |8586Pause / resume without redeploying: `dcc run --rm cli pause <id>` / `resume <id>` (sets `connectors.paused`).87Force one run now: `dcc run --rm cli enqueue <id> --task crawl --group newsroom` (a pending job with the same id is88coalesced; a run already in progress makes the job wait 60 s on the run lock).8990## 4. Budgets and cost control9192Four independent caps, all enforced in `apps/worker/src/context.ts` before a fetch is allowed to escalate:93941. **Per run** — `fetch.maxCreditsPerRun` (YAML, default 200): once spent, the rest of the run is direct-only.952. **Per connector per day** — `fetch.maxCreditsPerDay` (YAML, optional): Redis `dci:budget:connector:<id>:<day>`96   (UTC day, INCRBYFLOAT, 3-day TTL), shared by every worker. A warning is logged once per run when reached.973. **Per provider per day** — `DCI_SCRAPFLY_DAILY_BUDGET` / `DCI_FIRECRAWL_DAILY_BUDGET` (400 / 200): Redis98   `dci:budget:<provider>:<day>` shared by every worker (`RedisBudgetStore`, refreshed every99   `DCI_BUDGET_REFRESH_MS` = 15 s; the in-process counter never lets a worker undercount its own spend). When Redis100   is unreachable the in-process counter alone applies and one log line says so.1014. **Never premium** on discovery fetches (groups `sitemap`, `rss`: L1 → L2 only), never on documents that are not102   due (`dueDocuments` filters on `next_check`; only `--force` / `--url` runs bypass it), and only every 4th attempt103   for a document that already failed twice in a row.104105Reading the counters (the API `/api/admin/ops` reads the same keys):106107```bash108ssh BHS128 'cd /srv/dci/app && docker compose -f deploy/compose.data.yml --env-file deploy/.env.data exec -T redis sh -c "redis-cli -a \"\$REDIS_PASSWORD\" --no-auth-warning --scan --pattern \"dci:budget:*\" | sort | while read k; do echo \"\$k \$(redis-cli -a \"\$REDIS_PASSWORD\" --no-auth-warning get \$k)\"; done"'109dpsql -c "select connector_id, sum((stats->>'credits')::float) credits, sum((stats->>'fetched')::int) fetched from connector_runs where started_at >= date_trunc('day', now() at time zone 'utc') and (stats->>'credits')::float > 0 group by 1 order by 2 desc"110```111112Rules of thumb: Scrapfly with `asp + render_js` costs ~5–6 credits per page, `renderJs: false` ~1–2; a news113connector behind a bot wall should have `maxCreditsPerRun ≤ 40` and a `maxCreditsPerDay`. A connector that114needs premium for **every** page is a connector to rethink, not a budget to raise.115116Emergency stop: `dcc run --rm cli pause <id>`; or set `DCI_SCRAPFLY_DAILY_BUDGET=0` in `deploy/.env.crawl` and117`deploy/bin/deploy.sh crawl --no-build` (the fetcher reports itself unavailable, escalation stops at L2).118119## 5. Scheduling model (what the scheduler does every 60 s)120121* Per enabled, unpaused connector: a `full` job if it was never discovered, a `discover` job when the discovery122  cadence (`schedule.discovery` → `schedule.sitemap` → shortest group interval, minimum 6 h) has elapsed since123  `connector_state.lastDiscoverAt` — a discovery that legitimately finds nothing waits like any other.124* Per group with due documents: one `crawl` job `<connector>__<group>`, **unless** a `full`/`discover` job of that125  connector is pending. Job ids dedupe; `enqueueRun` re-adds only when the previous job finished.126* Workers take `dci:run-lock:<connector>` (TTL 4 h) before running; a colliding job is moved to delayed for 60 s127  (`dci_worker_jobs_total{result="deferred"}`).128* Maintenance schedulers (UTC): metrics 00:10, rankings 00:30, refresh-stats hourly, cleanup 01:00 (also129  reconciles orphaned runs). Orphans (`status = running` older than `DCI_ORPHAN_RUN_HOURS` = 6) are also aborted130  at every worker start.131* Priorities: manual enqueue 1, `full` 5, group crawl 10, `discover` 20 (lower = sooner).132133## 6. Resilience behaviour134135| Event | What happens |136|---|---|137| `docker compose stop worker` (SIGTERM) | scheduler loop stops, no new jobs are taken, in-flight documents finish, runs end `aborted` (their remaining documents stay due), heartbeat says `shuttingDown`. Deadline `DCI_SHUTDOWN_TIMEOUT_MS` (50 s) < `stop_grace_period` (90 s): past it, in-flight runs are marked `aborted` in Postgres and the process exits; BullMQ re-queues the active jobs (stalled → retried once). |138| Worker killed hard (OOM, host reboot) | the run lock expires (4 h TTL) or is bypassed by the orphan reconciliation at next start (`aborted`, 6 h); BullMQ's stalled check (60 s, `maxStalledCount: 1`) moves the job back to waiting once, then fails it. |139| Redis unreachable | ioredis reconnects with backoff (one log line per outage); BullMQ workers resume; the scheduler tick fails and retries next interval (`SchedulerStale` after 10 min); budgets fall back to in-process counters. |140| Postgres unreachable | per-document errors are caught (run ends `partial`/`failed`), the run row update is retried at the end; the worker keeps running. |141| ClickHouse down | `crawl_log` / `page_changes` / observations inserts are best effort (`DCI_CLICKHOUSE_OPTIONAL=1`): a warning per failed insert, never a failed run. |142| MinIO down | archive failures are logged per document; the document is still processed (no version diff without the old body). |143| Scheduler container down | the crawl worker's embedded loop keeps scheduling (`DCI_SCHEDULER` defaults to on for a crawl worker); the Redis lock `dci:scheduler:lock` prevents double enqueues when both run. |144145## 7. Scaling146147* **More throughput on BHS64b**: raise `DCI_CRAWL_CONCURRENCY` (jobs in parallel; each job honours its148  connector's `fetch.concurrency` and `rpm`), or start the second worker: `dcc --profile scale up -d worker-b`149  (metrics on :8322, already a Prometheus target). Two workers never run the same connector at once (run lock).150* **Watch**: `QueueBacklog`, `dci_worker_running_jobs` at the concurrency ceiling for long stretches, p95 fetch151  duration, memory of the worker container (`mem_limit` 12 g).152* **Per-host politeness is per process**: the token bucket (`rpm`, `concurrency`) is in-process, so two workers153  double the pressure on a host — keep `rpm` conservative when scaling out.154* **More connectors**: cost is dominated by discovery on the first run (`full`); `discovery.maxUrlsPerRun` now155  keeps the highest-priority groups (facility pages, seeds) rather than the first N sitemap entries.156* **Maintenance jobs** run on the data node (`worker-maint`, `DCI_MAINTENANCE_CONCURRENCY` 2) and on the crawl157  worker; heavy rankings/metrics can be moved off the crawl node with `DCI_QUEUES=crawl` in `.env.crawl`.158159## 8. Backups (data node)160161Nightly `dci-backup.timer` 03:30 → `deploy/bin/remote/backup-run.sh`; on demand `deploy/bin/backup.sh`162(≈ 1 min). Off-node mirror: `ubuntu@BHS64b:/srv/dci-backups` (rsync over dci0). Verify: `ssh BHS64b ls -la163/srv/dci-backups/pg`. Known defect (2026-09-11): the ClickHouse `BACKUP DATABASE` zip is written by the container164user (uid 101, mode 640) and is skipped by the rsync (`Permission denied`, exit 23) until `backup-run.sh` chowns it165— Postgres and config backups are unaffected.166