# Research Journal ## 2026-08-06 — Phase 0: Setup - Created structure: `/src`, `/data`, `/certs`, `journal.md`, `conjectures.md`, `records.md`. - Implemented `src/core.py`: odd-only numpy sieve, segmented sieve, deterministic Miller–Rabin (12 bases, valid for n < 3.317×10^24, Sorenson–Webster), BPSW (strong Lucas, Selfridge params). - Validation `src/test_core.py`: **37/37 PASS** in 0.3 s — π(10^2..10^8) exact (incl. π(10^8)=5761455 via segmented sieve), segment-vs-full sieve equality, full maximal-gap table below 10^6 reproduced, strong-pseudoprime & Carmichael traps, MR≡sieve on [0,40000), BPSW≡det-MR on 300 random 62-bit integers (seed 42). - Gate passed → allowed to proceed. ## 2026-08-06 — Phase 1: Axis selection **Chosen axis: #10 — Prime deserts (large gaps).** Justification (5 lines): 1. Fully verifiable: a gap is certified by two primality certificates + compositeness of the interior (sieve or covering) — zero ambiguity. 2. Measurable success criteria exist at every scale: merit g/ln p, CSG ratio g/ln²p, reproduction of the known maximal-gap table below our compute bound. 3. Rich conjecture surface: Cramér/Granville corrections, residues of gap endpoints, jumping champions — testable to 10^9+ locally. 4. State of the art is precisely documented (records.md) so novelty checks are cheap and honest. 5. The PROVER role has a concrete deliverable even without records: covering-system desert constructions with independent re-verification scripts. ## 2026-08-06 — Phase 2: State of the art - Web-checked current records (see `records.md`): largest maximal gap 1854 (2026), merit record 41.94 (2017), CSG record 0.9206 (Nyman 1999), exhaustive bound ≈ 2×10^19. - Realistic targets: (i) independent verification of the maximal-gap table to ~4×10^9, (ii) massively tested statistical conjectures on gaps, (iii) certified covering-system desert + verification pipeline. ## Cycle 1 — 2026-08-06 — COMPLETE (stopping criterion (a): a conjecture refuted + criterion (b): a construction certified) ### EXPLORER (Phase 3) - `src/explore_gaps.py 1e8` (0.6 s): all 25 maximal gaps below 10^8 reproduced = published table; jumping champion 6; gaps ≡ 0 mod 6 are 43.9% of all gaps; best merit 12.46 (gap 210 after 20831323). - `src/analyze_gaps.py 1e8` with checkpoints 10^6/10^7/10^8 (bug found & fixed: segment boundaries must align on checkpoints — symptom: twin counts wrong at sub-segment checkpoints; after fix, N(2,x) = 8169 / 58980 / 440312 = literature values exactly). - Key quantified observations: (O1) N2−N4 margins +26/+359/+55 — tiny and non-monotone; (O2) multiples of 6 strict local maxima up to g = 66 at all three scales; (O3) ρ(x)·ln x ≈ −0.6 stable; (O4) naive HL gap model exp(−g/ln t) shows structured deviations (deficit at g = 36, 72, 100, 108) — parked for a future cycle (needs inclusion–exclusion model before conjecturing). ### CONJECTURER (Phase 4) - C1 champion=6; C2 race N2>N4 for x ≥ 10^6; C3 G6(x) ≥ 66; C4 ρ·ln x ∈ [−0.65,−0.55] → c ≈ −0.6; C5 max CSG below 4×10^9 = 0.7395. Two trivial/parity statements discarded at birth. See conjectures.md. ### ADVERSARY (Phase 5) - `src/adversary_race.py 4e9` (15 s, deterministic): range = 40× discovery. - **C2 REFUTED** — full-resolution scan (every gap event): first tie after 10^6 at end-prime 80966861, first strict overtake at 80966933; the race then changes leader forever in range (137.6M events with D ≤ 0; D(4×10^9) = −2270). Checkpoint-only verification had been misleading — lesson recorded. - C1, C3, C5 survive (C3 strengthened: G6 jumps to 216 at 10^9). C4 survives with revision (drift of ρ·ln x toward −0.565; "≈ −0.6 limit" retracted). - Cramér/Maier test applied: C2's refutation is exactly random-model behavior (fair random walk); C1/C3/C4 are structural (absent from a Cramér model). - All 32 maximal gaps below 4×10^9 = published table (independent verification, `data/adversary_4e9.json`). ### PROVER (Phase 6) - Attempt 1 (`src/prover_desert.py`): pure covering system, primes ≤ 59, greedy residues + CRT → certified gap 90 at a 22-digit N, merit 1.78. Correct but weak. - Attempt 2 (`src/prover_desert2.py`): **hybrid covering** — greedy covers 243/259 positions for every shift t (PROVEN); 16 holes certified composite for the chosen t by explicit trial factors (15) and one strong MR witness (1); endpoints deterministic-MR prime (< 3.317×10^24 bound). Result: **certified gap 260 after N = 1116336781708038449369693 (25 digits), merit 4.6955** — ×2.9 vs attempt 1, ×4.3 vs the classic primorial baseline (61) with the same primes. NOT a record (records.md: merit record 41.94) — the deliverable is the certified pipeline. - Proven statement (elementary): for every t ≥ 0, the 243 covered positions of x + t·59# + i are composite — an explicit infinite family of long composite runs. - Independent verifications: `certs/verify_desert.py` PASSED; `certs/verify_c2_refutation.py` PASSED. ### Decisions - Cycle closed. Next cycle: C4 at 10^10 (does ρ·ln x stabilize?) — see report_cycle1.md. ## Cycle 2 — 2026-08-06 — COMPLETE (focus: C4 at 10^10/4×10^10, dual-machine) ### Setup / budget - User directive: run C4 at 10^10 on this computer AND on M3U96a. Strategy: identical deterministic scan on both — laptop to 10^10 (40 s), M3U96a (M3 Ultra) to 4×10^10 (128 s, single-core numpy). Deploy via `cluster-gateway.sh stage` to `~/cluster-projects/conjoncture/`. - Sanity gate: `src/cycle2_c4.py 1e8` reproduces cycle-1 values exactly (ρ·ln x = −0.59604 / −0.61428 / −0.58221 at 10^6/10^7/10^8; D = 26/359/55). PASS. ### EXPLORER/ADVERSARY (Phases 3+5 merged — C4 escalation) - **Cross-machine validation:** every overlapping checkpoint (10^6…10^10) bit-for-bit identical between laptop and M3U96a. Independent-hardware reproducibility ✓. - ρ₁·ln x continues drifting up: −0.58221 (10^8) → −0.56116 (10^10) → −0.55569 (4×10^10). Cycle-1 window [−0.65, −0.55] still holds at 4×10^10 — barely. - New: lag-2 correlation ρ₂·ln x drifts DOWN: −0.235 (10^8) → −0.249 (4×10^10). - Free continuations: champion = 6 up to 4×10^10 (C1 ✓); G6 ≥ 66 everywhere but NON-monotone (312 at 10^10 → 282 at 2×10^10 → 276 at 4×10^10 — tail noise, C3 statement unaffected); C2 race still swinging (D = +2681 at 10^10, +7160 at 2×10^10, −804 at 4×10^10); all 37 maximal gaps below 4×10^10 = published table incl. 384@20678048297, 394@22367084959, 456@25056082087 (web-checked OEIS A002386 / t5k.org after a wrong hand-written checklist briefly suggested a mismatch — the SCAN was right, the from-memory list was wrong; lesson: never hand-write "known" tables from memory, always fetch). - Max CSG below 4×10^10 = 0.79535 (gap 456 after 25056082087) — new in-range CSG high, known. ### CONJECTURER (Phase 4) - `src/fit_c4.py`: ρ₁·ln x = c + d/ln x fits 9 checkpoints (10^8…4×10^10) with max residual 0.0025: **c = −0.486 ± 0.010 (LOO), d = −1.72** → C4′ stated (see conjectures.md), with the sharp falsifiable prediction ρ·ln x(10^12) = −0.549 ± 0.005, and predicted exit of the cycle-1 window near x ≈ 5×10^11. C6 (lag-2, c₂ ≈ −0.29) stated with lower confidence. ### PROVER (Phase 6) - `src/model_c4.py`: first-order HL triple-correlation model P(g1,g2) ∝ S({0,g1,g1+g2})·e^{−(g1+g2)/λ}. Result: correct sign, correct c + d/λ drift form, magnitude ρ·λ ≈ −0.154 vs observed −0.49: **quantitatively REJECTED** (factor ~3.2). Honest conclusion: the anticorrelation is dominated by interior-compositeness constraints (inclusion–exclusion), not by the bare triple singular series. Building the full inclusion–exclusion model is the open PROVER problem for a later cycle. ### Decisions - Cycle closed (criterion (a)-adjacent: C4 superseded by refined C4′ with a falsifiable constant). - Next precise action: test C4′ at 10^12 — predicted ρ·ln x = −0.549 ± 0.005. Needs ~25× more compute than 4×10^10 (~1 h on M3U96a single-core, or minutes core-proportionally across the cluster). ## Cycle 3 — 2026-08-06 — COMPLETE (C4′ tested at 10^12, distributed on the cluster) ### Setup / budget - User directive: distribute; exclude M3U96b and M2U64. Design: 1000 chunks of 10^9 (boundaries = multiples of 10^9 so checkpoints fall on chunk edges), mergeable streaming sums, 5000-wide junction buffer per chunk (max gap < 10^12 is 540, so the buffer always contains the 2 preceding gaps). - Small-scale gate first: 10 chunks × 10^9 merged locally reproduce the cycle-2 single-machine values at 10^10 EXACTLY (ρ, D, G6, champion). PASS. - Fleet reality vs plan: 18 candidate nodes → m2m16 unreachable; numpy installed on 5 more nodes (Apple CLT Python 3.9); 5 nodes had NO usable python3 (missing Xcode CLT: m4ma, m4mb, m4mc, M2M32c, m2m8b) → dropped. Final fleet: 12 nodes / 162 cores, core-proportional quotas (174/99/99/99/86/86/74/74/62/49/49/49 chunks). ### Incident (instructive failure) - First launch: 9/12 nodes crashed instantly — `np.ndarray | None` annotation in core.py is Python ≥ 3.10 syntax; the 9 nodes run Apple's Python 3.9. Fix: `from __future__ import annotations` (annotations become lazy). Local test suite re-run (PASS), core.py re-pushed, import-checked on all nine, relaunched. Lesson: pin the LOWEST interpreter version in the fleet as the compatibility target, and import-check on every node class before a fleet launch. - zsh gotchas hit twice in orchestration (`$NODES` not word-split → `${=NODES}`; `$n:c…` eaten as a history-style modifier in scp targets → quote remote paths). ### ADVERSARY (Phase 5 — the run) - 1000/1000 chunks collected; merge verified exact tiling of [2, 10^12). - **Cross-validation:** merged distributed results equal the cycle-2 single-machine scan exactly at 10^10, 2×10^10, 4×10^10 (|Δρ·ln x| = 0 at 5 decimals; D identical). Pipeline validated end-to-end. - **C4′ VERDICT: CONFIRMED.** ρ·ln x(10^12) = −0.54756 ∈ predicted −0.549 ± 0.005. Old C4 window [−0.65, −0.55] exited between 4×10^11 and 10^12 as predicted (≈ 5×10^11). - Checkpoints: −0.55381 (10^11), −0.55203 (2×10^11), −0.55027 (4×10^11), −0.54756 (10^12). - Free continuations: champion = 6 to 10^12 (C1); G6 → 480 at 10^12 (C3, growing); C2 race STILL swinging at 10^12 (D = −238 out of ~1.5×10^10 gap events — astonishingly tight); all 48 maximal gaps below 10^12 = published table (A005250/A002386; tail spot-checked online, ends 540 @ 738832927927); max CSG below 10^12 unchanged in significance (known values). - Compute: 17,719 core-seconds ≈ 4.9 core-hours; wall ≈ 5 min over two launches. ### CONJECTURER (Phase 4 — refit) - Combined 13-point fit (10^8…10^12): c = −0.48449 (LOO [−0.48793, −0.48196]), d = −1.7571, max residual 0.00245. New predictions: ρ·ln x(10^13) = −0.5432 ± 0.004, ρ·ln x(10^14) = −0.5390 ± 0.004. - C6 firmed up: c₂ = −0.2806, d₂ = +0.7815 (opposite drift sign vs lag-1 confirmed); prediction ρ₂·ln x(10^13) = −0.2545 ± 0.004. ### Decisions - Cycle closed (criterion (b)-adjacent: prediction confirmed, law strengthened). - **Fleet policy change (user directive, recorded):** future distributed runs use ONLY M3U96a + M4M64a + M4M64b, saturated, adding a Metal/GPU path (MLX or Metal kernels) for the sieve. Benchmark GPU vs numpy before committing a full run. - Next precise action: cycle 4 — C4′ at 10^13 on the 3-node fleet with a Metal-accelerated sieve (10× cycle 3's work: ~50 core-hours CPU-only, so the GPU path is the enabler). ## Cycle 4 — 2026-08-06 — COMPLETE (C4′ and C6 tested at 10^13; CPU+Metal GPU on 3 nodes) ### Setup / gates - Fleet per user directive: M3U96a (28c) + M4M64a (16c) + M4M64b (16c), saturated, plus Metal GPU. - MLX installed where missing (PEP 668 → `--user --break-system-packages`). - **GPU sieve** (`src/gpu_sieve.py`, `mx.fast.metal_kernel`): grid (blocks × base primes), byte marking from max(p², first multiple ≥ block), benign write races. Gates passed: byte-identical prime sets vs CPU sieve on 4 windows up to 10^13; full worker partial identical field-for-field on a 10^9 chunk. Bench: GPU 1.2–1.3 s vs CPU-core 2.7–3.6 s per 10^9 near 10^12 → GPU ≈ 2–3 CPU cores per node. - 10,000 chunks of 10^9; measured-rate quotas 4000/3000/3000 with GPU sub-quotas 370/400/400 (27/15/15 CPU workers, 1 core reserved per node for GPU extraction). Local 10-chunk merge gate reproduced cycle-2 values exactly before launch. ### The run - Wall ≈ 50 min; 147,738 core-seconds (~41 core-hours) including ~4% duplicated tail work from a late rebalance (M4 nodes finished their quotas and absorbed 500 tail chunks of M3U96a's queue; duplicates are byte-identical, tiling asserted at merge). GPU processed 1170 chunks (11.7%). - Merge: 10,000 unique chunks tile [2, 10^13) exactly; **all 10 overlapping checkpoints identical to cycles 2/3** (|Δρ·ln x| = 0, D exact). ### VERDICTS - **C4′ CONFIRMED (2nd out-of-sample hit): ρ·ln x(10^13) = −0.54264 vs predicted −0.5432 ± 0.004.** - **C6 CONFIRMED (1st hit): ρ₂·ln x(10^13) = −0.25247 vs predicted −0.2545 ± 0.004.** - 16-point refits: c = −0.48354 (LOO ±0.002), d = −1.7772; c₂ = −0.27762, d₂ = +0.7187. Next: ρ·ln x(10^14) = −0.5387 ± 0.004; ρ₂·ln x(10^14) = −0.2553 ± 0.004. - Continuations: champion 6 to 10^13; G6 = 576 at 10^13; C2 race STILL swinging (D: −98,967 at 2×10^12, +8,870 at 10^13); all 53 maximal gaps below 10^13 = published table (new in range: 582, 588, 602, 652, 674 — tail web-checked t5k/Nicely); max CSG below 10^13 = 0.79754 (gap 652 after 2614941710599). ### Decisions - Cycle closed. Next precise action: cycle 5 — 10^14 (~10× work: ~400 CPU-core-hours equivalent; with the 3-node fleet ≈ 8–9 h wall, GPU included — schedule as an overnight run) to test ρ·ln x(10^14) = −0.5387 ± 0.004. Alternative next action if compute is deferred: the PROVER inclusion–exclusion model for the constant c (main open problem). ## Cycle 5 — 2026-08-06 — ABORTED BY USER (no data retained) - C4′/C6 test at 10^14 launched on the 3-node fleet (10,000 chunks of 10^10, CPU+Metal GPU, measured-rate quotas 3990/3005/3005, ~35% complete at stop). User decision: stop here for now. - All workers killed, partials and chunk lists removed from the nodes, local temp cleaned. Nodes back to idle (0–8% load). No partial results were merged — nothing to report, nothing claimed. - **The predictions remain on record and untested**: ρ·ln x(10^14) = −0.5387 ± 0.004, ρ₂·ln x(10^14) = −0.2553 ± 0.004. Relaunch recipe: regenerate chunk lists, redeploy, run `worker_c4.py`/`gpu_runner.py` as in this journal's cycle-5 setup (workers are idempotent).