Architecture
8 probes (Go, launchd/systemd) RIPE RIS Live (wss) public status pages (optional)
Québec · Montréal · Atlanta · Gravelines 23 collectors, ~10k prefixes/s Statuspage / GCP / AWS RSS / status.io
Dublin · Istanbul · Uşak · Ayia Napa │ │
│ signed gzip batches (HTTPS) │ │
▼ ▼ ▼
┌──────────────── BHS64b (OVH Beauharnois, Docker Compose) ──────────────────────────────────────┐
│ edge Caddy :8350 (bound to wg1 10.67.0.61) ── /ingest,/api → api :8352 ── / → web :8351 │
│ api (FastAPI) ── ingest → ClickHouse writer + Redis live counters │
│ bgp (RIS Live) ── 10 s per-collector aggregates, per-origin minutes, sampled raw │
│ corroboration ── vendor_status (Postgres) │
│ engine (10 s) ── baselines (5 min) → robust z → components → scopes → global → events → fronts │
│ → Redis live state + pub/sub → SSE ; ClickHouse pressure_history/provenance │
│ Postgres 17 (registry, config, events) ClickHouse 25.8 (telemetry, history) Redis 7 (live) │
└──────────────────────────────────────────────────────────────────────────────────────────────────┘
▲ WireGuard 10.67.0.0/24
BHS64 gateway: Caddy TLS https://www.internetpressure.io → 10.67.0.61:8350 (DNS A → 51.161.112.61)Decisions
- Python (FastAPI) backend, Go probe. The spec mandates Go for the agent; the backend language is free. Python matches the other MacLustr platforms (AI Atlas, SatelliteIndex, CountryAtlas) and keeps the scoring code readable/testable. All stores are spoken to over their wire protocols (asyncpg, ClickHouse HTTP, redis) — no ORM.
- One backend image, four processes (
ip api,ip engine,ip bgp,ip corroboration). They share nothing but the stores. - ClickHouse HTTP interface via httpx (JSONEachRow) instead of a driver: async, dependency-free, trivially inspectable.
- Baselines from raw measurements (180 d TTL) recomputed every 5 minutes; MAD estimated from the IQR (0.7413 × IQR) in a
single pass; hour-of-day seasonality (±1 h) once ≥ 3 days of history exist.
measurements_1m(AggregatingMergeTree MV, 1 y) serves long-range charts. - Self-exclusion (
engine/health.py): < 2 fresh usable probes, every probe failing at once, or a store down ⇒internal_status = degraded, score frozen (stale: true), no history rows, no incidents. A probe failing ≥ 80 % of its HTTP targets while the median probe is fine is excluded (its uplink, not the Internet). - Calibration: until at least one signal has ≥ 20 % of
baseline_min_samples(12 five-minute buckets), the instrument publishespressure: null,level: calibrating. Nothing is invented. - Config-driven scoring:
packages/config/pressure.yaml, optionally overridden by the Postgresconfigrow (admin UI); the engine re-reads it every cycle. Weights must sum to 1 (validated). - Explainability: every cycle stores
signal_features(component, signal, scope, current, baseline, MAD, z, samples, stress, contribution) and the live payload carriesexplain[](attribution = component contribution × share of stress mass by region / signal / service) andcomponents[].drivers[].
Scoring pipeline (engine/compute.py, engine/scoring.py)
- Pair signals: per (probe, target, kind, resolver) in the last 120 s: latency z (ttfb, tcp, rtt), loss excess, failure-rate excess, 5xx/TLS/reset rates, DNS failure & latency; per target: corroborated unavailability (≥ 2 probes in ≥ 2 regions) and resolver disagreement (one resolver answers, another SERVFAILs); per (probe, target) traceroute: route fingerprint changed vs the 7-day dominant route, hop-count z, latency shift on changed paths.
- Pair stress 0..1:
z_to_stressramps from z = 1 to z = 6; rate signals userate_stress(excess / max(scale, 3 × baseline)); route churn usesratio_stress. - Signal stress = importance-weighted mean of pair stresses; the weight is damped by baseline coverage
(
coverage_factor), so thin baselines cannot move the index. - Component score =
saturate(Σ signal_weight × signal_stress),saturate(x) = 100·(1−e^{−kx})/(1−e^{−k}), k = 1.2. - Routing (global only, ASN level via
bgp_origin_1m): weighted-median robust z across collectors of per-minute withdrawals / announcements / origin changes vs their 7-day (seasonal) baseline; collector disagreement when a minority of collectors spike. - Corroboration: Σ impact(indicator) × service importance / 4, capped at 1 — weight 0.05, optional by design.
- Global = Σ weight × component over available components (weights renormalised). Regions/countries/services/ASNs reuse the same machinery over their signal subsets (routing null where not attributable).
- Velocity / acceleration / volatility from the last hour of
pressure_history. - Events (
engine/events.py): scope score ≥ 45 with confidence ≥ 0.45 for 2 cycles ⇒detected; →developing→active(6 cycles); < 30 ⇒recovering; 10 min under 30 ⇒resolved. Rule-based hypotheses with evidence. - Pressure Fronts (
engine/fronts.py): ≥ 3 stressed (probe, target) pairs on a source-region → destination-region corridor with saturated corridor stress ≥ 35.
Ethics & load
Per probe with ~27 700 targets (tiers 20 s / 45 s / 300 s / 900 s; the long tail is HTTP-only): ~60 HTTP/s, ~26 DNS/s (4 resolvers / 600 s, infrastructure only), ~20 ICMP runs/s (every 120 s, infrastructure only), ≤ 1 traceroute in flight (every 15 min
for the 30 tr: true targets). Keep-alive disabled (to measure TCP/TLS), body read capped at 16 KiB, no retries, no parallel
checks to one host, intervals never below 10 s. Total ≈ 105 measurements/s/probe → ~70 M rows/day for 8 probes (~3–4 GB/day
compressed in ClickHouse).
Replay
ip replay <from> <to> --weights '{…}' (and POST /api/admin/replay) recomputes the global index from stored component
history with alternative weights. Re-deriving components from raw measurements with alternative engine parameters is an
offline job to add when enough history exists (raw is kept 180 days for exactly that purpose).
Retention — we never erase (founder decision, 2026-09-12, migration 0002_keep_forever.sql)
Every table is kept indefinitely: measurements, measurements_1m, traceroutes, probe_health, bgp_stats_10s, bgp_origin_1m,
pressure_history, signal_features, engine_runs. The single exception is bgp_events, the raw RIS Live stream (≈ 10 k
prefixes/s at the source, already sampled 1/25 for announcements): 90 days. Nightly backups (deploy/bin/backup.sh) copy
Postgres and the irreplaceable ClickHouse tables off-node to BHS128. Disk growth ≈ 3–4 GB/day on BHS64b's 878 GB RAID-1;
plan a second ClickHouse volume (or a cold tier on BHS128) within a year.
Registry at scale (2026-09-13)
data/targets/targets.yaml (2 128 hand-curated) + data/targets/fragments/{government,finance,infrastructure,developer,media,commerce}.yaml
(26 700 more, every host verified by DNS + HTTP with the probe UA at authoring time). ip seed loads the main file then the
fragments, skipping duplicate ids or (host, url, ip) — first occurrence wins. Tier 4 (900 s) carries the long tail; the probes
receive the whole list (≈ 8 MB JSON, cached per config version) and schedule ≈ 90 000 jobs each with 48-way concurrency.
Engine: ≈ 35 000 (probe, target) pairs in each 120 s window, ≈ 100 000 signals, ~2 s per 15 s cycle. The sampling boost is
delivered inside batch responses and never bumps the config version (a bump makes all probes re-download the registry).