SPB Git forge
38commits 1branches 0releases
338.7 MBsize
maindefault branch
2 h agolast push
HTML 53.9% TypeScript 44.5% JavaScript 0.6% SQL 0.5%
16.0 KB

# Crawl operations — DataCenterIndex.io

2026-09-12 additions. Connector health now has blocked, schema_change, no_new_content states and a quarantine flag (dci quarantine <id> on|off, auto-set after 5 consecutive failed/blocked runs or a non-full run creating ≥ 500 records — everything runs, nothing is published, connector_runs.quarantined = true). The effective extractor version is the YAML parserVersion plus the versions of the parsers actually used: bump a parser version and dci reprocess <id> --stale re-extracts only the documents extracted with an older version. dci trace <doc> shows every stage for one document; dci quality / dci snapshot / dci gaps run the data-quality jobs (also scheduled: quality every 6 h, snapshot + regression checks daily 00:45 UTC). 429 responses never escalate and honour retry-after; robots.txt failures fail closed to L1/L2; noarchive pages are never archived.

Runbook for the crawl runtime (apps/worker) in production: the crawl worker on BHS64b (compose.crawl.yml, service worker, /healthz + /metrics on 10.68.0.2:8320), the dedicated scheduler and the maintenance-only worker on BHS128 (compose.data.yml, services scheduler :8321 and worker-maint), Postgres/Redis/ClickHouse/MinIO on BHS128. Deployment itself is in docs/DEPLOY.md; the runtime internals in apps/worker/README.md.

Shell shortcuts used below:

bash
alias dcd='ssh BHS128 "cd /srv/dci/app && docker compose -f deploy/compose.data.yml --env-file deploy/.env.data"'   # data node
alias dcc='ssh BHS64b "cd /srv/dci/app && docker compose -f deploy/compose.crawl.yml --env-file deploy/.env.crawl"'  # crawl node
alias dpsql='ssh BHS128 "cd /srv/dci/app && docker compose -f deploy/compose.data.yml --env-file deploy/.env.data exec -T postgres psql -U dci -d dci"'

# 1. Daily checks (5 minutes)

# Check Command / where Expected
1 Containers deploy/bin/status.sh everything Up … (healthy); scheduler healthy (it was in a restart loop before 2026-09-11)
2 Worker liveness ssh BHS64b curl -s 10.68.0.2:8320/healthz | jq '{ok,runningJobs,creditsToday,sharedBudgets,lastScheduler,queueDepth}' ok: true, lastScheduler.at < 2 min old
3 Scheduler liveness dcd exec -T scheduler node -e "fetch('http://127.0.0.1:8321/healthz').then(r=>r.text()).then(console.log)" ok: true, lastTick.error: null
4 Queue depth Grafana → Queue depth, or dci_queue_jobs{queue="crawl",state="waiting"} near 0 between ticks; a backlog > 50 for 30 min fires QueueBacklog
5 Premium credits Grafana → Premium budget used today; ssh BHS128 … redis-cli … MGET dci:budget:scrapfly:$(date -u +%F) dci:budget:firecrawl:$(date -u +%F) well below DCI_*_DAILY_BUDGET (400 / 200); know which connector spent them (§4)
6 Runs last 24 h dpsql -c "select status, count(*) from connector_runs where started_at > now() - interval '24 hours' group by 1" mostly ok; investigate failed / aborted
7 Failing / degraded connectors dpsql -c "select id, health, last_status, left(last_error,80) from connectors where enabled and not paused and (health <> 'ok' or last_status in ('failed','partial'))" empty or known
8 Doctor dcc run --rm cli doctor all critical checks pass; warn lines are triage input
9 Alerts Grafana → Alerting none firing (TargetDown, CrawlStalled, SchedulerStale, WorkerDown, PremiumBudgetNearlyExhausted, FetchErrorRateHigh, CrawlRunsFailing, QueueBacklog, disk/memory)
10 Backups ssh BHS128 systemctl list-timers dci-backup.timer; ssh BHS64b du -sh /srv/dci-backups/* timer waiting for 03:30, mirror grew last night

Weekly: dpsql -c "select count(*) from documents where quarantined" (should grow slowly, not jump), select count(*) from entity_matches where status='pending', disk on both nodes (status.sh).

# 2. Reading the metrics

Both processes export the same registry (apps/worker/src/prom.ts); Prometheus scrapes dci-worker (BHS64b :8320/:8322 and worker-maint:8320) and dci-scheduler (scheduler:8321).

Metric Meaning Typical query
dci_crawl_fetches_total{connector,level,outcome} one sample per fetch that returned (outcome ok, not_modified, error, blocked; level = final level L1–L4) sum by (outcome) (rate(…[5m])); per connector error ratio sum by (connector) (rate(…{outcome=~"error|blocked"}[1h])) / sum by (connector) (rate(…[1h]))
dci_crawl_fetch_duration_seconds (histogram) wall time of one fetch incl. escalation histogram_quantile(0.95, sum(rate(…_bucket[5m])) by (le))
dci_crawl_credits_total{provider} premium credits spent since process start sum by (provider) (increase(…[1h]))
dci_crawl_daily_budget{provider} / dci_crawl_daily_credits_used{provider} daily cap and shared usage today (Redis, survives restarts) used / budget → the PremiumBudgetNearlyExhausted rule
dci_ingest_entities_total{connector,result} entities handed to ingest: created, updated, merged, unchanged, rejected sum by (result) (rate(…[5m]))
dci_events_total change events emitted by ingest increase(…[24h])
dci_crawl_runs_total{status} runs finished by status (ok, partial, failed, aborted) CrawlRunsFailing
dci_worker_jobs_total{queue,result} BullMQ jobs: completed, failed, stalled, deferred (run lock held by another worker) many deferred = two workers keep colliding on the same connectors
dci_queue_jobs{queue,state} snapshot of queue depth per state at scrape time state="waiting" / "active" / "failed"
dci_worker_running_jobs, dci_worker_up, dci_worker_uptime_seconds per process WorkerDown
dci_scheduler_ticks_total{result}, dci_scheduler_enqueued_total, dci_scheduler_last_tick_timestamp_seconds scheduler health time() - max(last_tick) > 600 → SchedulerStale

Where the counters live outside Prometheus: connector_runs.stats (per run), connectors.stats (per connector, refreshed after each run), ClickHouse crawl_log (one row per fetch, 400 days) and page_changes.

# 3. Triaging a failing or degraded connector

connectors.health is degraded when ≥ 20 % of a run's fetches failed, failing at ≥ 50 % or when the run threw. Start from the last run:

bash
dpsql -c "select id, task, status, started_at, stats->>'fetched' f, stats->>'failed' fl, stats->>'credits' cr, left(error,120) from connector_runs where connector_id='<id>' order by started_at desc limit 5"
dpsql -c "select l->>'t' t, l->>'level' lvl, l->>'msg' from connector_runs r, jsonb_array_elements(r.log) l where r.id='<run_id>' and l->>'level' in ('warn','error') limit 40"
dpsql -c "select split_part(error,':',1) code, count(*) from documents where connector_id='<id>' and error_count>0 group by 1 order by 2 desc"
dcc run --rm cli docs <id> --due --limit 20        # what is about to be fetched
dcc run --rm cli run <id> --dry-run --limit 3 --verbose   # live reproduction from the crawl IP, nothing persisted

Common signatures:

Symptom Cause Action
HTTP 403 on every page, credits > 0 bot wall; Scrapfly ASP also refused after 2 consecutive failures the runtime stops escalating (premiumAllowedAfterErrors), so the cost stops by itself. Fix the connector: fetch.userAgent: browser, renderJs: false (cheaper), or find a feed / sitemap that is open
rss … → 403 at discovery, discovered 0 urls feed blocked for the bot identity (discovery is direct-only, never premium) connector: fetch.userAgent: browser or another feed URL. The scheduler retries discovery on its cadence (≥ 6 h), never every tick
discovered 0 urls with a working feed (filter: 0/20 items match) connector filter too narrow connector code / YAML; not a runtime issue
robots.txt disallows legal stop — leave it remove the path from discovery
fetch_failed, timeout, Headers Timeout upstream slow (Overpass, big PDFs) raise fetch.timeoutMs for the connector or lower fetch.concurrency; documents back off ×2^(n−1) automatically
partial with validation errors only extractor problem (implausible MW, missing name) parser bug: file/line in the run log; data is not ingested for rejected entities
every refetch CHANGED with diff_summary null, versions 2 = fetch_count two runs on the same connector overlapped (fixed 2026-09-11 by the per-connector run lock and the scheduler skipping group jobs while a full job is pending) should not recur; if it does, check dci_worker_jobs_total{result="deferred"} and dci:run-lock:* in Redis
every refetch CHANGED with a small ratio (< 0.1) and sidebars in added/removed page embeds "related articles" / ads: the fingerprint moves although the article did not connector: extract main content only (each/match on the article container); the change_frequency_score will keep re-checking such pages twice as often until fixed
connector never_run, enabled = f disabled in YAML intentional (17 operator connectors are parked)

Pause / resume without redeploying: dcc run --rm cli pause <id> / resume <id> (sets connectors.paused). Force one run now: dcc run --rm cli enqueue <id> --task crawl --group newsroom (a pending job with the same id is coalesced; a run already in progress makes the job wait 60 s on the run lock).

# 4. Budgets and cost control

Four independent caps, all enforced in apps/worker/src/context.ts before a fetch is allowed to escalate:

  1. Per run — fetch.maxCreditsPerRun (YAML, default 200): once spent, the rest of the run is direct-only.
  2. Per connector per day — fetch.maxCreditsPerDay (YAML, optional): Redis dci:budget:connector:<id>:<day> (UTC day, INCRBYFLOAT, 3-day TTL), shared by every worker. A warning is logged once per run when reached.
  3. Per provider per day — DCI_SCRAPFLY_DAILY_BUDGET / DCI_FIRECRAWL_DAILY_BUDGET (400 / 200): Redis dci:budget:<provider>:<day> shared by every worker (RedisBudgetStore, refreshed every DCI_BUDGET_REFRESH_MS = 15 s; the in-process counter never lets a worker undercount its own spend). When Redis is unreachable the in-process counter alone applies and one log line says so.
  4. Never premium on discovery fetches (groups sitemap, rss: L1 → L2 only), never on documents that are not due (dueDocuments filters on next_check; only --force / --url runs bypass it), and only every 4th attempt for a document that already failed twice in a row.

Reading the counters (the API /api/admin/ops reads the same keys):

bash
ssh BHS128 'cd /srv/dci/app && docker compose -f deploy/compose.data.yml --env-file deploy/.env.data exec -T redis sh -c "redis-cli -a \"\$REDIS_PASSWORD\" --no-auth-warning --scan --pattern \"dci:budget:*\" | sort | while read k; do echo \"\$k \$(redis-cli -a \"\$REDIS_PASSWORD\" --no-auth-warning get \$k)\"; done"'
dpsql -c "select connector_id, sum((stats->>'credits')::float) credits, sum((stats->>'fetched')::int) fetched from connector_runs where started_at >= date_trunc('day', now() at time zone 'utc') and (stats->>'credits')::float > 0 group by 1 order by 2 desc"

Rules of thumb: Scrapfly with asp + render_js costs ~5–6 credits per page, renderJs: false ~1–2; a news connector behind a bot wall should have maxCreditsPerRun ≤ 40 and a maxCreditsPerDay. A connector that needs premium for every page is a connector to rethink, not a budget to raise.

Emergency stop: dcc run --rm cli pause <id>; or set DCI_SCRAPFLY_DAILY_BUDGET=0 in deploy/.env.crawl and deploy/bin/deploy.sh crawl --no-build (the fetcher reports itself unavailable, escalation stops at L2).

# 5. Scheduling model (what the scheduler does every 60 s)

  • Per enabled, unpaused connector: a full job if it was never discovered, a discover job when the discovery cadence (schedule.discovery → schedule.sitemap → shortest group interval, minimum 6 h) has elapsed since connector_state.lastDiscoverAt — a discovery that legitimately finds nothing waits like any other.
  • Per group with due documents: one crawl job <connector>__<group>, unless a full/discover job of that connector is pending. Job ids dedupe; enqueueRun re-adds only when the previous job finished.
  • Workers take dci:run-lock:<connector> (TTL 4 h) before running; a colliding job is moved to delayed for 60 s (dci_worker_jobs_total{result="deferred"}).
  • Maintenance schedulers (UTC): metrics 00:10, rankings 00:30, refresh-stats hourly, cleanup 01:00 (also reconciles orphaned runs). Orphans (status = running older than DCI_ORPHAN_RUN_HOURS = 6) are also aborted at every worker start.
  • Priorities: manual enqueue 1, full 5, group crawl 10, discover 20 (lower = sooner).

# 6. Resilience behaviour

Event What happens
docker compose stop worker (SIGTERM) scheduler loop stops, no new jobs are taken, in-flight documents finish, runs end aborted (their remaining documents stay due), heartbeat says shuttingDown. Deadline DCI_SHUTDOWN_TIMEOUT_MS (50 s) < stop_grace_period (90 s): past it, in-flight runs are marked aborted in Postgres and the process exits; BullMQ re-queues the active jobs (stalled → retried once).
Worker killed hard (OOM, host reboot) the run lock expires (4 h TTL) or is bypassed by the orphan reconciliation at next start (aborted, 6 h); BullMQ's stalled check (60 s, maxStalledCount: 1) moves the job back to waiting once, then fails it.
Redis unreachable ioredis reconnects with backoff (one log line per outage); BullMQ workers resume; the scheduler tick fails and retries next interval (SchedulerStale after 10 min); budgets fall back to in-process counters.
Postgres unreachable per-document errors are caught (run ends partial/failed), the run row update is retried at the end; the worker keeps running.
ClickHouse down crawl_log / page_changes / observations inserts are best effort (DCI_CLICKHOUSE_OPTIONAL=1): a warning per failed insert, never a failed run.
MinIO down archive failures are logged per document; the document is still processed (no version diff without the old body).
Scheduler container down the crawl worker's embedded loop keeps scheduling (DCI_SCHEDULER defaults to on for a crawl worker); the Redis lock dci:scheduler:lock prevents double enqueues when both run.

# 7. Scaling

  • More throughput on BHS64b: raise DCI_CRAWL_CONCURRENCY (jobs in parallel; each job honours its connector's fetch.concurrency and rpm), or start the second worker: dcc --profile scale up -d worker-b (metrics on :8322, already a Prometheus target). Two workers never run the same connector at once (run lock).
  • Watch: QueueBacklog, dci_worker_running_jobs at the concurrency ceiling for long stretches, p95 fetch duration, memory of the worker container (mem_limit 12 g).
  • Per-host politeness is per process: the token bucket (rpm, concurrency) is in-process, so two workers double the pressure on a host — keep rpm conservative when scaling out.
  • More connectors: cost is dominated by discovery on the first run (full); discovery.maxUrlsPerRun now keeps the highest-priority groups (facility pages, seeds) rather than the first N sitemap entries.
  • Maintenance jobs run on the data node (worker-maint, DCI_MAINTENANCE_CONCURRENCY 2) and on the crawl worker; heavy rankings/metrics can be moved off the crawl node with DCI_QUEUES=crawl in .env.crawl.

# 8. Backups (data node)

Nightly dci-backup.timer 03:30 → deploy/bin/remote/backup-run.sh; on demand deploy/bin/backup.sh (≈ 1 min). Off-node mirror: ubuntu@BHS64b:/srv/dci-backups (rsync over dci0). Verify: ssh BHS64b ls -la /srv/dci-backups/pg. Known defect (2026-09-11): the ClickHouse BACKUP DATABASE zip is written by the container user (uid 101, mode 640) and is skipped by the rsync (Permission denied, exit 23) until backup-run.sh chowns it — Postgres and config backups are unaffected.