Crawl operations — DataCenterIndex.io
2026-09-12 additions. Connector health now has
blocked,schema_change,no_new_contentstates and aquarantineflag (dci quarantine <id> on|off, auto-set after 5 consecutive failed/blocked runs or a non-full run creating ≥ 500 records — everything runs, nothing is published,connector_runs.quarantined = true). The effective extractor version is the YAMLparserVersionplus the versions of the parsers actually used: bump a parserversionanddci reprocess <id> --stalere-extracts only the documents extracted with an older version.dci trace <doc>shows every stage for one document;dci quality/dci snapshot/dci gapsrun the data-quality jobs (also scheduled: quality every 6 h, snapshot + regression checks daily 00:45 UTC). 429 responses never escalate and honourretry-after; robots.txt failures fail closed to L1/L2;noarchivepages are never archived.
Runbook for the crawl runtime (apps/worker) in production: the crawl worker on BHS64b
(compose.crawl.yml, service worker, /healthz + /metrics on 10.68.0.2:8320), the dedicated
scheduler and the maintenance-only worker on BHS128 (compose.data.yml, services scheduler :8321 and
worker-maint), Postgres/Redis/ClickHouse/MinIO on BHS128. Deployment itself is in docs/DEPLOY.md; the
runtime internals in apps/worker/README.md.
Shell shortcuts used below:
alias dcd='ssh BHS128 "cd /srv/dci/app && docker compose -f deploy/compose.data.yml --env-file deploy/.env.data"' # data node
alias dcc='ssh BHS64b "cd /srv/dci/app && docker compose -f deploy/compose.crawl.yml --env-file deploy/.env.crawl"' # crawl node
alias dpsql='ssh BHS128 "cd /srv/dci/app && docker compose -f deploy/compose.data.yml --env-file deploy/.env.data exec -T postgres psql -U dci -d dci"'1. Daily checks (5 minutes)
| # | Check | Command / where | Expected |
|---|---|---|---|
| 1 | Containers | deploy/bin/status.sh |
everything Up … (healthy); scheduler healthy (it was in a restart loop before 2026-09-11) |
| 2 | Worker liveness | ssh BHS64b curl -s 10.68.0.2:8320/healthz | jq '{ok,runningJobs,creditsToday,sharedBudgets,lastScheduler,queueDepth}' |
ok: true, lastScheduler.at < 2 min old |
| 3 | Scheduler liveness | dcd exec -T scheduler node -e "fetch('http://127.0.0.1:8321/healthz').then(r=>r.text()).then(console.log)" |
ok: true, lastTick.error: null |
| 4 | Queue depth | Grafana → Queue depth, or dci_queue_jobs{queue="crawl",state="waiting"} |
near 0 between ticks; a backlog > 50 for 30 min fires QueueBacklog |
| 5 | Premium credits | Grafana → Premium budget used today; ssh BHS128 … redis-cli … MGET dci:budget:scrapfly:$(date -u +%F) dci:budget:firecrawl:$(date -u +%F) |
well below DCI_*_DAILY_BUDGET (400 / 200); know which connector spent them (§4) |
| 6 | Runs last 24 h | dpsql -c "select status, count(*) from connector_runs where started_at > now() - interval '24 hours' group by 1" |
mostly ok; investigate failed / aborted |
| 7 | Failing / degraded connectors | dpsql -c "select id, health, last_status, left(last_error,80) from connectors where enabled and not paused and (health <> 'ok' or last_status in ('failed','partial'))" |
empty or known |
| 8 | Doctor | dcc run --rm cli doctor |
all critical checks pass; warn lines are triage input |
| 9 | Alerts | Grafana → Alerting | none firing (TargetDown, CrawlStalled, SchedulerStale, WorkerDown, PremiumBudgetNearlyExhausted, FetchErrorRateHigh, CrawlRunsFailing, QueueBacklog, disk/memory) |
| 10 | Backups | ssh BHS128 systemctl list-timers dci-backup.timer; ssh BHS64b du -sh /srv/dci-backups/* |
timer waiting for 03:30, mirror grew last night |
Weekly: dpsql -c "select count(*) from documents where quarantined" (should grow slowly, not jump), select count(*) from entity_matches where status='pending', disk on both nodes (status.sh).
2. Reading the metrics
Both processes export the same registry (apps/worker/src/prom.ts); Prometheus scrapes dci-worker (BHS64b
:8320/:8322 and worker-maint:8320) and dci-scheduler (scheduler:8321).
| Metric | Meaning | Typical query |
|---|---|---|
dci_crawl_fetches_total{connector,level,outcome} |
one sample per fetch that returned (outcome ok, not_modified, error, blocked; level = final level L1–L4) |
sum by (outcome) (rate(…[5m])); per connector error ratio sum by (connector) (rate(…{outcome=~"error|blocked"}[1h])) / sum by (connector) (rate(…[1h])) |
dci_crawl_fetch_duration_seconds (histogram) |
wall time of one fetch incl. escalation | histogram_quantile(0.95, sum(rate(…_bucket[5m])) by (le)) |
dci_crawl_credits_total{provider} |
premium credits spent since process start | sum by (provider) (increase(…[1h])) |
dci_crawl_daily_budget{provider} / dci_crawl_daily_credits_used{provider} |
daily cap and shared usage today (Redis, survives restarts) | used / budget → the PremiumBudgetNearlyExhausted rule |
dci_ingest_entities_total{connector,result} |
entities handed to ingest: created, updated, merged, unchanged, rejected |
sum by (result) (rate(…[5m])) |
dci_events_total |
change events emitted by ingest | increase(…[24h]) |
dci_crawl_runs_total{status} |
runs finished by status (ok, partial, failed, aborted) |
CrawlRunsFailing |
dci_worker_jobs_total{queue,result} |
BullMQ jobs: completed, failed, stalled, deferred (run lock held by another worker) |
many deferred = two workers keep colliding on the same connectors |
dci_queue_jobs{queue,state} |
snapshot of queue depth per state at scrape time | state="waiting" / "active" / "failed" |
dci_worker_running_jobs, dci_worker_up, dci_worker_uptime_seconds |
per process | WorkerDown |
dci_scheduler_ticks_total{result}, dci_scheduler_enqueued_total, dci_scheduler_last_tick_timestamp_seconds |
scheduler health | time() - max(last_tick) > 600 → SchedulerStale |
Where the counters live outside Prometheus: connector_runs.stats (per run), connectors.stats (per connector,
refreshed after each run), ClickHouse crawl_log (one row per fetch, 400 days) and page_changes.
3. Triaging a failing or degraded connector
connectors.health is degraded when ≥ 20 % of a run's fetches failed, failing at ≥ 50 % or when the run
threw. Start from the last run:
dpsql -c "select id, task, status, started_at, stats->>'fetched' f, stats->>'failed' fl, stats->>'credits' cr, left(error,120) from connector_runs where connector_id='<id>' order by started_at desc limit 5"
dpsql -c "select l->>'t' t, l->>'level' lvl, l->>'msg' from connector_runs r, jsonb_array_elements(r.log) l where r.id='<run_id>' and l->>'level' in ('warn','error') limit 40"
dpsql -c "select split_part(error,':',1) code, count(*) from documents where connector_id='<id>' and error_count>0 group by 1 order by 2 desc"
dcc run --rm cli docs <id> --due --limit 20 # what is about to be fetched
dcc run --rm cli run <id> --dry-run --limit 3 --verbose # live reproduction from the crawl IP, nothing persistedCommon signatures:
| Symptom | Cause | Action |
|---|---|---|
HTTP 403 on every page, credits > 0 |
bot wall; Scrapfly ASP also refused | after 2 consecutive failures the runtime stops escalating (premiumAllowedAfterErrors), so the cost stops by itself. Fix the connector: fetch.userAgent: browser, renderJs: false (cheaper), or find a feed / sitemap that is open |
rss … → 403 at discovery, discovered 0 urls |
feed blocked for the bot identity (discovery is direct-only, never premium) | connector: fetch.userAgent: browser or another feed URL. The scheduler retries discovery on its cadence (≥ 6 h), never every tick |
discovered 0 urls with a working feed (filter: 0/20 items match) |
connector filter too narrow | connector code / YAML; not a runtime issue |
robots.txt disallows |
legal stop — leave it | remove the path from discovery |
fetch_failed, timeout, Headers Timeout |
upstream slow (Overpass, big PDFs) | raise fetch.timeoutMs for the connector or lower fetch.concurrency; documents back off ×2^(n−1) automatically |
partial with validation errors only |
extractor problem (implausible MW, missing name) |
parser bug: file/line in the run log; data is not ingested for rejected entities |
every refetch CHANGED with diff_summary null, versions 2 = fetch_count |
two runs on the same connector overlapped (fixed 2026-09-11 by the per-connector run lock and the scheduler skipping group jobs while a full job is pending) |
should not recur; if it does, check dci_worker_jobs_total{result="deferred"} and dci:run-lock:* in Redis |
every refetch CHANGED with a small ratio (< 0.1) and sidebars in added/removed |
page embeds "related articles" / ads: the fingerprint moves although the article did not | connector: extract main content only (each/match on the article container); the change_frequency_score will keep re-checking such pages twice as often until fixed |
connector never_run, enabled = f |
disabled in YAML | intentional (17 operator connectors are parked) |
Pause / resume without redeploying: dcc run --rm cli pause <id> / resume <id> (sets connectors.paused).
Force one run now: dcc run --rm cli enqueue <id> --task crawl --group newsroom (a pending job with the same id is
coalesced; a run already in progress makes the job wait 60 s on the run lock).
4. Budgets and cost control
Four independent caps, all enforced in apps/worker/src/context.ts before a fetch is allowed to escalate:
- Per run —
fetch.maxCreditsPerRun(YAML, default 200): once spent, the rest of the run is direct-only. - Per connector per day —
fetch.maxCreditsPerDay(YAML, optional): Redisdci:budget:connector:<id>:<day>(UTC day, INCRBYFLOAT, 3-day TTL), shared by every worker. A warning is logged once per run when reached. - Per provider per day —
DCI_SCRAPFLY_DAILY_BUDGET/DCI_FIRECRAWL_DAILY_BUDGET(400 / 200): Redisdci:budget:<provider>:<day>shared by every worker (RedisBudgetStore, refreshed everyDCI_BUDGET_REFRESH_MS= 15 s; the in-process counter never lets a worker undercount its own spend). When Redis is unreachable the in-process counter alone applies and one log line says so. - Never premium on discovery fetches (groups
sitemap,rss: L1 → L2 only), never on documents that are not due (dueDocumentsfilters onnext_check; only--force/--urlruns bypass it), and only every 4th attempt for a document that already failed twice in a row.
Reading the counters (the API /api/admin/ops reads the same keys):
ssh BHS128 'cd /srv/dci/app && docker compose -f deploy/compose.data.yml --env-file deploy/.env.data exec -T redis sh -c "redis-cli -a \"\$REDIS_PASSWORD\" --no-auth-warning --scan --pattern \"dci:budget:*\" | sort | while read k; do echo \"\$k \$(redis-cli -a \"\$REDIS_PASSWORD\" --no-auth-warning get \$k)\"; done"'
dpsql -c "select connector_id, sum((stats->>'credits')::float) credits, sum((stats->>'fetched')::int) fetched from connector_runs where started_at >= date_trunc('day', now() at time zone 'utc') and (stats->>'credits')::float > 0 group by 1 order by 2 desc"Rules of thumb: Scrapfly with asp + render_js costs ~5–6 credits per page, renderJs: false ~1–2; a news
connector behind a bot wall should have maxCreditsPerRun ≤ 40 and a maxCreditsPerDay. A connector that
needs premium for every page is a connector to rethink, not a budget to raise.
Emergency stop: dcc run --rm cli pause <id>; or set DCI_SCRAPFLY_DAILY_BUDGET=0 in deploy/.env.crawl and
deploy/bin/deploy.sh crawl --no-build (the fetcher reports itself unavailable, escalation stops at L2).
5. Scheduling model (what the scheduler does every 60 s)
- Per enabled, unpaused connector: a
fulljob if it was never discovered, adiscoverjob when the discovery cadence (schedule.discovery→schedule.sitemap→ shortest group interval, minimum 6 h) has elapsed sinceconnector_state.lastDiscoverAt— a discovery that legitimately finds nothing waits like any other. - Per group with due documents: one
crawljob<connector>__<group>, unless afull/discoverjob of that connector is pending. Job ids dedupe;enqueueRunre-adds only when the previous job finished. - Workers take
dci:run-lock:<connector>(TTL 4 h) before running; a colliding job is moved to delayed for 60 s (dci_worker_jobs_total{result="deferred"}). - Maintenance schedulers (UTC): metrics 00:10, rankings 00:30, refresh-stats hourly, cleanup 01:00 (also
reconciles orphaned runs). Orphans (
status = runningolder thanDCI_ORPHAN_RUN_HOURS= 6) are also aborted at every worker start. - Priorities: manual enqueue 1,
full5, group crawl 10,discover20 (lower = sooner).
6. Resilience behaviour
| Event | What happens |
|---|---|
docker compose stop worker (SIGTERM) |
scheduler loop stops, no new jobs are taken, in-flight documents finish, runs end aborted (their remaining documents stay due), heartbeat says shuttingDown. Deadline DCI_SHUTDOWN_TIMEOUT_MS (50 s) < stop_grace_period (90 s): past it, in-flight runs are marked aborted in Postgres and the process exits; BullMQ re-queues the active jobs (stalled → retried once). |
| Worker killed hard (OOM, host reboot) | the run lock expires (4 h TTL) or is bypassed by the orphan reconciliation at next start (aborted, 6 h); BullMQ's stalled check (60 s, maxStalledCount: 1) moves the job back to waiting once, then fails it. |
| Redis unreachable | ioredis reconnects with backoff (one log line per outage); BullMQ workers resume; the scheduler tick fails and retries next interval (SchedulerStale after 10 min); budgets fall back to in-process counters. |
| Postgres unreachable | per-document errors are caught (run ends partial/failed), the run row update is retried at the end; the worker keeps running. |
| ClickHouse down | crawl_log / page_changes / observations inserts are best effort (DCI_CLICKHOUSE_OPTIONAL=1): a warning per failed insert, never a failed run. |
| MinIO down | archive failures are logged per document; the document is still processed (no version diff without the old body). |
| Scheduler container down | the crawl worker's embedded loop keeps scheduling (DCI_SCHEDULER defaults to on for a crawl worker); the Redis lock dci:scheduler:lock prevents double enqueues when both run. |
7. Scaling
- More throughput on BHS64b: raise
DCI_CRAWL_CONCURRENCY(jobs in parallel; each job honours its connector'sfetch.concurrencyandrpm), or start the second worker:dcc --profile scale up -d worker-b(metrics on :8322, already a Prometheus target). Two workers never run the same connector at once (run lock). - Watch:
QueueBacklog,dci_worker_running_jobsat the concurrency ceiling for long stretches, p95 fetch duration, memory of the worker container (mem_limit12 g). - Per-host politeness is per process: the token bucket (
rpm,concurrency) is in-process, so two workers double the pressure on a host — keeprpmconservative when scaling out. - More connectors: cost is dominated by discovery on the first run (
full);discovery.maxUrlsPerRunnow keeps the highest-priority groups (facility pages, seeds) rather than the first N sitemap entries. - Maintenance jobs run on the data node (
worker-maint,DCI_MAINTENANCE_CONCURRENCY2) and on the crawl worker; heavy rankings/metrics can be moved off the crawl node withDCI_QUEUES=crawlin.env.crawl.
8. Backups (data node)
Nightly dci-backup.timer 03:30 → deploy/bin/remote/backup-run.sh; on demand deploy/bin/backup.sh
(≈ 1 min). Off-node mirror: ubuntu@BHS64b:/srv/dci-backups (rsync over dci0). Verify: ssh BHS64b ls -la /srv/dci-backups/pg. Known defect (2026-09-11): the ClickHouse BACKUP DATABASE zip is written by the container
user (uid 101, mode 640) and is skipped by the rsync (Permission denied, exit 23) until backup-run.sh chowns it
— Postgres and config backups are unaffected.