🔬 Prime Mystery Engine
A verification-first computational laboratory for prime gap statistics.
An autonomous research pipeline that studies gaps between consecutive primes under a strict
epistemic protocol: no claim without a certificate. Every statement is labelled
PROVEN / CERTIFIED / VERIFIED UP TO X / CONJECTURED / REFUTED, every scan is
deterministic and bit-reproducible across machines, and every headline object ships with a
standalone re-verification script.
🏆 Headline results
| Result | Status |
|---|---|
| Race conjecture refuted — "N(2,x) > N(4,x) for x ≥ 10⁶" is false: first tie at end-prime 80 966 861, first strict overtake at 80 966 933; the lead changes forever up to 10¹³ | REFUTED (minimal counterexample) |
Certified prime desert — gap of exactly 260 after the 25-digit prime 1116336781708038449369693 (merit 4.70), built from a hybrid covering system mod primes ≤ 59 |
CERTIFIED |
| Scaling law for consecutive-gap correlation — ρ(x)·ln x = c + d/ln x with c = −0.4835 ± 0.002, d = −1.777; out-of-sample predictions confirmed at 10¹² (−0.549 ± 0.005 → −0.54756) and 10¹³ (−0.5432 ± 0.004 → −0.54264) | CONJECTURED, 2/2 hits |
| Lag-2 law — ρ₂(x)·ln x → c₂ ≈ −0.278, prediction confirmed at 10¹³ (−0.2545 ± 0.004 → −0.25247) | CONJECTURED, 1/1 hit |
| First-order Hardy–Littlewood triple model rejected — right sign and drift form, but only ~32 % of the observed magnitude | MODEL REJECTED |
| All published tables reproduced — 53/53 maximal gaps below 10¹³ (OEIS A005250/A002386), twin counts, π(10ᵏ), CSG maxima | VERIFIED |
Predictions on record (falsifiable): ρ·ln x(10¹⁴) = −0.5387 ± 0.004 · ρ·ln x(10¹⁵) = −0.5350 ± 0.004 · ρ₂·ln x(10¹⁴) = −0.2553 ± 0.004
⚡ Verify the headline claims yourself (stock Python 3)
# the certified 260-gap desert (independent Miller–Rabin implementation)
python3 certs/verify_desert.py
# the race-conjecture refutation (stdlib-only sieve, ~30 s)
python3 certs/verify_c2_refutation.py
# the core-primitives validation gate (37 checks, needs numpy)
python3 src/test_core.py🗂 Repository layout
src/ core primitives, scans, distributed worker/merger, Metal GPU sieve, models
certs/ certificates + standalone verifiers (the "trust nothing" layer)
data/ raw checkpoint JSONs for every campaign (deterministic, re-mergeable)
paper/ LaTeX manuscript (main.tex → main.pdf, 13 pp., full math detail)
journal.md complete research journal: every cycle, decision, failure
conjectures.md the conjecture ledger with epistemic statuses
records.md state of the art (what would count as new; thresholds to beat)
report_cycle{1..4}.md per-cycle syntheses🖥 Compute
| Campaign | Range | Fleet | Wall time |
|---|---|---|---|
| Cycle 1 | 4×10⁹ | 1 laptop core | 15 s |
| Cycle 2 | 4×10¹⁰ | laptop + M3 Ultra (1 core each) | ~2 min |
| Cycle 3 | 10¹² | 12 Apple-silicon nodes, 162 CPU cores | ~5 min |
| Cycle 4 | 10¹³ | 3 nodes: 58 CPU cores + 3 GPUs (custom Metal kernel) | ~50 min |
The distributed design uses mergeable streaming statistics (exact integers < 2⁵³ in float64 → zero rounding, bit-identical across machines) with checkpoint-aligned chunking and a junction buffer, so a 10 000-chunk merge equals the single-machine scan exactly. The Metal sieve kernel was accepted only after producing byte-identical prime sets to the CPU sieve; it processed 11.7 % of cycle 4.
📄 Paper
Full write-up with all equations, proofs of the infrastructure lemmas, robustness checks and
threats-to-validity: paper/main.tex → paper/main.pdf.
👤 Author
Simon-Pierre Boucher · contact@spboucher.ai