spb/prime-mystery-engine
Public
Python 64.3%
TeX 35.7%
1<!--2journal.md — Prime Mystery Engine: research journal3Author: Simon-Pierre Boucher — contact@spboucher.ai4-->56# Research Journal78## 2026-08-06 — Phase 0: Setup910- Created structure: `/src`, `/data`, `/certs`, `journal.md`, `conjectures.md`, `records.md`.11- Implemented `src/core.py`: odd-only numpy sieve, segmented sieve, deterministic Miller–Rabin12 (12 bases, valid for n < 3.317×10^24, Sorenson–Webster), BPSW (strong Lucas, Selfridge params).13- Validation `src/test_core.py`: **37/37 PASS** in 0.3 s —14 π(10^2..10^8) exact (incl. π(10^8)=5761455 via segmented sieve), segment-vs-full sieve equality,15 full maximal-gap table below 10^6 reproduced, strong-pseudoprime & Carmichael traps,16 MR≡sieve on [0,40000), BPSW≡det-MR on 300 random 62-bit integers (seed 42).17- Gate passed → allowed to proceed.1819## 2026-08-06 — Phase 1: Axis selection2021**Chosen axis: #10 — Prime deserts (large gaps).** Justification (5 lines):221. Fully verifiable: a gap is certified by two primality certificates + compositeness of the interior (sieve or covering) — zero ambiguity.232. Measurable success criteria exist at every scale: merit g/ln p, CSG ratio g/ln²p, reproduction of the known maximal-gap table below our compute bound.243. Rich conjecture surface: Cramér/Granville corrections, residues of gap endpoints, jumping champions — testable to 10^9+ locally.254. State of the art is precisely documented (records.md) so novelty checks are cheap and honest.265. The PROVER role has a concrete deliverable even without records: covering-system desert constructions with independent re-verification scripts.2728## 2026-08-06 — Phase 2: State of the art2930- Web-checked current records (see `records.md`): largest maximal gap 1854 (2026), merit record 41.94 (2017),31 CSG record 0.9206 (Nyman 1999), exhaustive bound ≈ 2×10^19.32- Realistic targets: (i) independent verification of the maximal-gap table to ~4×10^9, (ii) massively tested33 statistical conjectures on gaps, (iii) certified covering-system desert + verification pipeline.3435## Cycle 1 — 2026-08-06 — COMPLETE (stopping criterion (a): a conjecture refuted + criterion (b): a construction certified)3637### EXPLORER (Phase 3)38- `src/explore_gaps.py 1e8` (0.6 s): all 25 maximal gaps below 10^8 reproduced = published table;39 jumping champion 6; gaps ≡ 0 mod 6 are 43.9% of all gaps; best merit 12.46 (gap 210 after 20831323).40- `src/analyze_gaps.py 1e8` with checkpoints 10^6/10^7/10^8 (bug found & fixed: segment boundaries41 must align on checkpoints — symptom: twin counts wrong at sub-segment checkpoints; after fix,42 N(2,x) = 8169 / 58980 / 440312 = literature values exactly).43- Key quantified observations: (O1) N2−N4 margins +26/+359/+55 — tiny and non-monotone;44 (O2) multiples of 6 strict local maxima up to g = 66 at all three scales; (O3) ρ(x)·ln x ≈ −0.6 stable;45 (O4) naive HL gap model exp(−g/ln t) shows structured deviations (deficit at g = 36, 72, 100, 108) —46 parked for a future cycle (needs inclusion–exclusion model before conjecturing).4748### CONJECTURER (Phase 4)49- C1 champion=6; C2 race N2>N4 for x ≥ 10^6; C3 G6(x) ≥ 66; C4 ρ·ln x ∈ [−0.65,−0.55] → c ≈ −0.6;50 C5 max CSG below 4×10^9 = 0.7395. Two trivial/parity statements discarded at birth. See conjectures.md.5152### ADVERSARY (Phase 5)53- `src/adversary_race.py 4e9` (15 s, deterministic): range = 40× discovery.54- **C2 REFUTED** — full-resolution scan (every gap event): first tie after 10^6 at end-prime 80966861,55 first strict overtake at 80966933; the race then changes leader forever in range (137.6M events with56 D ≤ 0; D(4×10^9) = −2270). Checkpoint-only verification had been misleading — lesson recorded.57- C1, C3, C5 survive (C3 strengthened: G6 jumps to 216 at 10^9). C4 survives with revision (drift of58 ρ·ln x toward −0.565; "≈ −0.6 limit" retracted).59- Cramér/Maier test applied: C2's refutation is exactly random-model behavior (fair random walk);60 C1/C3/C4 are structural (absent from a Cramér model).61- All 32 maximal gaps below 4×10^9 = published table (independent verification, `data/adversary_4e9.json`).6263### PROVER (Phase 6)64- Attempt 1 (`src/prover_desert.py`): pure covering system, primes ≤ 59, greedy residues + CRT →65 certified gap 90 at a 22-digit N, merit 1.78. Correct but weak.66- Attempt 2 (`src/prover_desert2.py`): **hybrid covering** — greedy covers 243/259 positions for every67 shift t (PROVEN); 16 holes certified composite for the chosen t by explicit trial factors (15) and one68 strong MR witness (1); endpoints deterministic-MR prime (< 3.317×10^24 bound). Result:69 **certified gap 260 after N = 1116336781708038449369693 (25 digits), merit 4.6955** —70 ×2.9 vs attempt 1, ×4.3 vs the classic primorial baseline (61) with the same primes. NOT a record71 (records.md: merit record 41.94) — the deliverable is the certified pipeline.72- Proven statement (elementary): for every t ≥ 0, the 243 covered positions of x + t·59# + i are73 composite — an explicit infinite family of long composite runs.74- Independent verifications: `certs/verify_desert.py` PASSED; `certs/verify_c2_refutation.py` PASSED.7576### Decisions77- Cycle closed. Next cycle: C4 at 10^10 (does ρ·ln x stabilize?) — see report_cycle1.md.7879## Cycle 2 — 2026-08-06 — COMPLETE (focus: C4 at 10^10/4×10^10, dual-machine)8081### Setup / budget82- User directive: run C4 at 10^10 on this computer AND on M3U96a. Strategy: identical deterministic83 scan on both — laptop to 10^10 (40 s), M3U96a (M3 Ultra) to 4×10^10 (128 s, single-core numpy).84 Deploy via `cluster-gateway.sh stage` to `~/cluster-projects/conjoncture/`.85- Sanity gate: `src/cycle2_c4.py 1e8` reproduces cycle-1 values exactly (ρ·ln x = −0.59604 / −0.61428 /86 −0.58221 at 10^6/10^7/10^8; D = 26/359/55). PASS.8788### EXPLORER/ADVERSARY (Phases 3+5 merged — C4 escalation)89- **Cross-machine validation:** every overlapping checkpoint (10^6…10^10) bit-for-bit identical90 between laptop and M3U96a. Independent-hardware reproducibility ✓.91- ρ₁·ln x continues drifting up: −0.58221 (10^8) → −0.56116 (10^10) → −0.55569 (4×10^10).92 Cycle-1 window [−0.65, −0.55] still holds at 4×10^10 — barely.93- New: lag-2 correlation ρ₂·ln x drifts DOWN: −0.235 (10^8) → −0.249 (4×10^10).94- Free continuations: champion = 6 up to 4×10^10 (C1 ✓); G6 ≥ 66 everywhere but NON-monotone95 (312 at 10^10 → 282 at 2×10^10 → 276 at 4×10^10 — tail noise, C3 statement unaffected);96 C2 race still swinging (D = +2681 at 10^10, +7160 at 2×10^10, −804 at 4×10^10);97 all 37 maximal gaps below 4×10^10 = published table incl. 384@20678048297, 394@22367084959,98 456@25056082087 (web-checked OEIS A002386 / t5k.org after a wrong hand-written checklist99 briefly suggested a mismatch — the SCAN was right, the from-memory list was wrong; lesson:100 never hand-write "known" tables from memory, always fetch).101- Max CSG below 4×10^10 = 0.79535 (gap 456 after 25056082087) — new in-range CSG high, known.102103### CONJECTURER (Phase 4)104- `src/fit_c4.py`: ρ₁·ln x = c + d/ln x fits 9 checkpoints (10^8…4×10^10) with max residual 0.0025:105 **c = −0.486 ± 0.010 (LOO), d = −1.72** → C4′ stated (see conjectures.md), with the sharp106 falsifiable prediction ρ·ln x(10^12) = −0.549 ± 0.005, and predicted exit of the cycle-1 window107 near x ≈ 5×10^11. C6 (lag-2, c₂ ≈ −0.29) stated with lower confidence.108109### PROVER (Phase 6)110- `src/model_c4.py`: first-order HL triple-correlation model P(g1,g2) ∝ S({0,g1,g1+g2})·e^{−(g1+g2)/λ}.111 Result: correct sign, correct c + d/λ drift form, magnitude ρ·λ ≈ −0.154 vs observed −0.49:112 **quantitatively REJECTED** (factor ~3.2). Honest conclusion: the anticorrelation is dominated by113 interior-compositeness constraints (inclusion–exclusion), not by the bare triple singular series.114 Building the full inclusion–exclusion model is the open PROVER problem for a later cycle.115116### Decisions117- Cycle closed (criterion (a)-adjacent: C4 superseded by refined C4′ with a falsifiable constant).118- Next precise action: test C4′ at 10^12 — predicted ρ·ln x = −0.549 ± 0.005. Needs ~25× more compute119 than 4×10^10 (~1 h on M3U96a single-core, or minutes core-proportionally across the cluster).120121## Cycle 3 — 2026-08-06 — COMPLETE (C4′ tested at 10^12, distributed on the cluster)122123### Setup / budget124- User directive: distribute; exclude M3U96b and M2U64. Design: 1000 chunks of 10^9 (boundaries =125 multiples of 10^9 so checkpoints fall on chunk edges), mergeable streaming sums, 5000-wide junction126 buffer per chunk (max gap < 10^12 is 540, so the buffer always contains the 2 preceding gaps).127- Small-scale gate first: 10 chunks × 10^9 merged locally reproduce the cycle-2 single-machine values128 at 10^10 EXACTLY (ρ, D, G6, champion). PASS.129- Fleet reality vs plan: 18 candidate nodes → m2m16 unreachable; numpy installed on 5 more nodes130 (Apple CLT Python 3.9); 5 nodes had NO usable python3 (missing Xcode CLT: m4ma, m4mb, m4mc,131 M2M32c, m2m8b) → dropped. Final fleet: 12 nodes / 162 cores, core-proportional quotas132 (174/99/99/99/86/86/74/74/62/49/49/49 chunks).133134### Incident (instructive failure)135- First launch: 9/12 nodes crashed instantly — `np.ndarray | None` annotation in core.py is136 Python ≥ 3.10 syntax; the 9 nodes run Apple's Python 3.9. Fix: `from __future__ import annotations`137 (annotations become lazy). Local test suite re-run (PASS), core.py re-pushed, import-checked on all138 nine, relaunched. Lesson: pin the LOWEST interpreter version in the fleet as the compatibility139 target, and import-check on every node class before a fleet launch.140- zsh gotchas hit twice in orchestration (`$NODES` not word-split → `${=NODES}`; `$n:c…` eaten as a141 history-style modifier in scp targets → quote remote paths).142143### ADVERSARY (Phase 5 — the run)144- 1000/1000 chunks collected; merge verified exact tiling of [2, 10^12).145- **Cross-validation:** merged distributed results equal the cycle-2 single-machine scan exactly at146 10^10, 2×10^10, 4×10^10 (|Δρ·ln x| = 0 at 5 decimals; D identical). Pipeline validated end-to-end.147- **C4′ VERDICT: CONFIRMED.** ρ·ln x(10^12) = −0.54756 ∈ predicted −0.549 ± 0.005.148 Old C4 window [−0.65, −0.55] exited between 4×10^11 and 10^12 as predicted (≈ 5×10^11).149- Checkpoints: −0.55381 (10^11), −0.55203 (2×10^11), −0.55027 (4×10^11), −0.54756 (10^12).150- Free continuations: champion = 6 to 10^12 (C1); G6 → 480 at 10^12 (C3, growing); C2 race STILL151 swinging at 10^12 (D = −238 out of ~1.5×10^10 gap events — astonishingly tight);152 all 48 maximal gaps below 10^12 = published table (A005250/A002386; tail spot-checked online,153 ends 540 @ 738832927927); max CSG below 10^12 unchanged in significance (known values).154- Compute: 17,719 core-seconds ≈ 4.9 core-hours; wall ≈ 5 min over two launches.155156### CONJECTURER (Phase 4 — refit)157- Combined 13-point fit (10^8…10^12): c = −0.48449 (LOO [−0.48793, −0.48196]), d = −1.7571,158 max residual 0.00245. New predictions: ρ·ln x(10^13) = −0.5432 ± 0.004, ρ·ln x(10^14) = −0.5390 ± 0.004.159- C6 firmed up: c₂ = −0.2806, d₂ = +0.7815 (opposite drift sign vs lag-1 confirmed);160 prediction ρ₂·ln x(10^13) = −0.2545 ± 0.004.161162### Decisions163- Cycle closed (criterion (b)-adjacent: prediction confirmed, law strengthened).164- **Fleet policy change (user directive, recorded):** future distributed runs use ONLY165 M3U96a + M4M64a + M4M64b, saturated, adding a Metal/GPU path (MLX or Metal kernels) for the sieve.166 Benchmark GPU vs numpy before committing a full run.167- Next precise action: cycle 4 — C4′ at 10^13 on the 3-node fleet with a Metal-accelerated sieve168 (10× cycle 3's work: ~50 core-hours CPU-only, so the GPU path is the enabler).169170## Cycle 4 — 2026-08-06 — COMPLETE (C4′ and C6 tested at 10^13; CPU+Metal GPU on 3 nodes)171172### Setup / gates173- Fleet per user directive: M3U96a (28c) + M4M64a (16c) + M4M64b (16c), saturated, plus Metal GPU.174- MLX installed where missing (PEP 668 → `--user --break-system-packages`).175- **GPU sieve** (`src/gpu_sieve.py`, `mx.fast.metal_kernel`): grid (blocks × base primes), byte marking176 from max(p², first multiple ≥ block), benign write races. Gates passed: byte-identical prime sets vs177 CPU sieve on 4 windows up to 10^13; full worker partial identical field-for-field on a 10^9 chunk.178 Bench: GPU 1.2–1.3 s vs CPU-core 2.7–3.6 s per 10^9 near 10^12 → GPU ≈ 2–3 CPU cores per node.179- 10,000 chunks of 10^9; measured-rate quotas 4000/3000/3000 with GPU sub-quotas 370/400/400180 (27/15/15 CPU workers, 1 core reserved per node for GPU extraction). Local 10-chunk merge gate181 reproduced cycle-2 values exactly before launch.182183### The run184- Wall ≈ 50 min; 147,738 core-seconds (~41 core-hours) including ~4% duplicated tail work from a185 late rebalance (M4 nodes finished their quotas and absorbed 500 tail chunks of M3U96a's queue;186 duplicates are byte-identical, tiling asserted at merge). GPU processed 1170 chunks (11.7%).187- Merge: 10,000 unique chunks tile [2, 10^13) exactly; **all 10 overlapping checkpoints identical to188 cycles 2/3** (|Δρ·ln x| = 0, D exact).189190### VERDICTS191- **C4′ CONFIRMED (2nd out-of-sample hit): ρ·ln x(10^13) = −0.54264 vs predicted −0.5432 ± 0.004.**192- **C6 CONFIRMED (1st hit): ρ₂·ln x(10^13) = −0.25247 vs predicted −0.2545 ± 0.004.**193- 16-point refits: c = −0.48354 (LOO ±0.002), d = −1.7772; c₂ = −0.27762, d₂ = +0.7187.194 Next: ρ·ln x(10^14) = −0.5387 ± 0.004; ρ₂·ln x(10^14) = −0.2553 ± 0.004.195- Continuations: champion 6 to 10^13; G6 = 576 at 10^13; C2 race STILL swinging (D: −98,967 at 2×10^12,196 +8,870 at 10^13); all 53 maximal gaps below 10^13 = published table (new in range: 582, 588, 602,197 652, 674 — tail web-checked t5k/Nicely); max CSG below 10^13 = 0.79754 (gap 652 after 2614941710599).198199### Decisions200- Cycle closed. Next precise action: cycle 5 — 10^14 (~10× work: ~400 CPU-core-hours equivalent;201 with the 3-node fleet ≈ 8–9 h wall, GPU included — schedule as an overnight run) to test202 ρ·ln x(10^14) = −0.5387 ± 0.004. Alternative next action if compute is deferred: the PROVER203 inclusion–exclusion model for the constant c (main open problem).204205## Cycle 5 — 2026-08-06 — ABORTED BY USER (no data retained)206207- C4′/C6 test at 10^14 launched on the 3-node fleet (10,000 chunks of 10^10, CPU+Metal GPU,208 measured-rate quotas 3990/3005/3005, ~35% complete at stop). User decision: stop here for now.209- All workers killed, partials and chunk lists removed from the nodes, local temp cleaned.210 Nodes back to idle (0–8% load). No partial results were merged — nothing to report, nothing claimed.211- **The predictions remain on record and untested**: ρ·ln x(10^14) = −0.5387 ± 0.004,212 ρ₂·ln x(10^14) = −0.2553 ± 0.004. Relaunch recipe: regenerate chunk lists, redeploy, run213 `worker_c4.py`/`gpu_runner.py` as in this journal's cycle-5 setup (workers are idempotent).214