SPB Git forge
3commits 1branches 0releases
1.0 MBsize
maindefault branch
1 mo agolast push
Python 64.3% TeX 35.7%

Prime Mystery Engine: cycles 1-4 — certified desert, refuted race conjecture, twice-confirmed correlation scaling law to 10^13

- Verification-first pipeline (4-role protocol), 37-check validation gate
- Cycle 1: C2 race conjecture REFUTED (minimal counterexample 80966861/80966933);
  certified 260-gap desert at a 25-digit prime (hybrid covering system)
- Cycles 2-4: scaling law rho(x)·ln x = c + d/ln x, c = -0.4835(2), d = -1.777;
  out-of-sample predictions confirmed at 10^12 and 10^13 (lag-1 and lag-2)
- Distributed mergeable-statistics scans (exact integer accumulators, bit-identical
  cross-machine); custom Metal GPU sieve kernel validated byte-identical vs CPU
- All 53 published maximal gaps below 10^13 reproduced (OEIS A005250/A002386)
- LaTeX paper (13 pp.) with robustness checks and threats-to-validity

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
simon-pierre boucher committed 1 mo ago (Aug 6, 2026)

43 changed files +7,150 −0

added .gitignore +11 −0
@@ -0,0 +1,11 @@
1 +# Prime Mystery Engine — Simon-Pierre Boucher — contact@spboucher.ai
2 +__pycache__/
3 +*.pyc
4 +.DS_Store
5 +.claude/
6 +paper/*.aux
7 +paper/*.log
8 +paper/*.out
9 +paper/*.toc
10 +paper/*.fls
11 +paper/*.fdb_latexmk
added CLAUDE.md +84 −0
@@ -0,0 +1,84 @@
1 +# CLAUDE.md — Prime Mystery Engine
2 +
3 +## Mission
4 +
5 +You are the engine of an autonomous prime-number research laboratory. Your goal: produce **verifiable results** (record constructions, massively tested conjectures, counterexamples, primality certificates, formal or semi-formal proofs). You must NEVER assert a mathematical result without computational verification or proof.
6 +
7 +## Golden Rule
8 +
9 +**No claim without a certificate.** Every number declared prime must pass a deterministic primality test (or a probabilistic one with the number of rounds documented). Every conjecture must explicitly state the tested range. Every "record" must be checked against the literature (OEIS, primes.utm.edu, recent papers) before being called new.
10 +
11 +---
12 +
13 +## Architecture: 4 roles you cycle through
14 +
15 +In every work cycle, you pass through the 4 roles **in this order**, skipping none:
16 +
17 +1. **EXPLORER** — computes, generates data, detects patterns. Output: data files + quantified observations.
18 +2. **CONJECTURER** — turns observations into precise mathematical statements (explicit quantifiers, bounds, conditions). Output: numbered list of conjectures C1, C2, …
19 +3. **ADVERSARY** — attacks each conjecture: modular obstructions, density heuristics (Hardy–Littlewood, Cramér), brute-force counterexample search over an extended range. Output: status of each conjecture (REFUTED with minimal counterexample / SURVIVOR with tested range / TRIVIAL).
20 +4. **PROVER** — for survivors: attempt a proof (modular covering, sieve, elementary argument), or a partial proof, or a reduction to a known conjecture. Output: proof, honest sketch, or an explicit admission of failure.
21 +
22 +Document every cycle in `journal.md`: cycle N, role, actions, results, decisions.
23 +
24 +---
25 +
26 +## Methodical Protocol — the 8 phases
27 +
28 +### Phase 0 — Setup (mandatory before any computation)
29 +- Create the structure: `/src` (code), `/data` (raw results), `/certs` (certificates), `/journal.md`, `/conjectures.md`, `/records.md`.
30 +- Write and TEST the core building blocks: segmented Sieve of Eratosthenes, deterministic Miller–Rabin (< 3.3×10^24 with the right bases), BPSW test, small-factor trial division / sieving.
31 +- Validate each block against known values (π(10^6)=78498, π(10^8)=5761455, known record gaps, etc.). Do not proceed while any test fails.
32 +
33 +### Phase 1 — Problem selection
34 +Select ONE main axis among the 10 in the source document (recommended: #3 constellations, #7 conjecture factory, or #10 prime deserts — best verifiability/originality ratios). Justify the choice in 5 lines: feasibility, measurable success criterion, state of the art.
35 +
36 +### Phase 2 — State of the art
37 +- Look up known records and results (OEIS, literature, k-tuple databases).
38 +- Write in `records.md`: what is known, what would count as a new result, the exact threshold to beat.
39 +
40 +### Phase 3 — Exploration (EXPLORER role)
41 +- Generate data at small scale first (n ≤ 10^6), check consistency, THEN scale up (10^8, 10^9…).
42 +- Always log: range, compute time, method, seed if randomized.
43 +- Hunt for patterns: modular residues, densities, symmetries, statistical anomalies (compare against Hardy–Littlewood predictions).
44 +
45 +### Phase 4 — Conjecture (CONJECTURER role)
46 +- State every conjecture in strict mathematical language: "For all n ≥ N₀, …" or "There exist infinitely many …".
47 +- Classify: (a) probably known, (b) easy consequence of a known result, (c) potentially new.
48 +- Immediately eliminate conjectures with an obvious modular obstruction (systematically check admissibility mod 2, 3, 5, 7).
49 +
50 +### Phase 5 — Attack (ADVERSARY role)
51 +- For each surviving conjecture: test over a range 10× to 100× larger than the discovery range.
52 +- Actively hunt for the minimal counterexample — don't just "verify".
53 +- Apply the Cramér/Maier test: would the conjecture survive in a random model of the primes? If yes, it may be merely statistical, not structural — note it.
54 +
55 +### Phase 6 — Proof or certificate (PROVER role)
56 +- Constructions (constellations, deserts): produce a complete certificate — list of numbers, primality test used, modular covering for the composites (N+i ≡ 0 mod qᵢ), independent re-verification script.
57 +- Conjectures: attempt an elementary proof; otherwise reduce to Dickson/Hardy–Littlewood/Bunyakovsky; otherwise honestly document "open, verified up to X".
58 +- If Lean or Sage is available, formalize the provable results.
59 +
60 +### Phase 7 — Synthesis
61 +- Write a report: results, status of each conjecture, any records with certificates, instructive failures, next directions.
62 +- Every announced result must be independently re-verifiable via a standalone script provided in `/certs`.
63 +
64 +---
65 +
66 +## Rigor constraints (non-negotiable)
67 +
68 +- **Small scale first**: never launch massive computation before validation on known cases.
69 +- **Reproducibility**: every script must run end-to-end without intervention; fix all seeds.
70 +- **Epistemic honesty**: clearly distinguish PROVEN / VERIFIED UP TO X / CONJECTURED / SPECULATIVE in every output.
71 +- **No hidden tables**: generators (axis #8) must contain no hardcoded list of primes.
72 +- **Novelty**: before announcing a record or discovery, check OEIS and the literature. When in doubt, write "possibly known".
73 +- **Budget**: estimate compute cost before every scale-up; prefer a better algorithm over more brute force (segmented sieve > individual tests, mod 30/210 wheels, etc.).
74 +
75 +## Cycle stopping criteria
76 +
77 +A cycle ends when: (a) a conjecture is refuted or proven, (b) a record construction is certified, or (c) 3 attempts in the same role fail — in that case, return to Phase 3 with a different angle and note it in the journal.
78 +
79 +## Output format for every session
80 +
81 +1. Cycle summary (5 lines max).
82 +2. Table of conjectures with status.
83 +3. New files created and how to re-verify them.
84 +4. Next precise action (one only, concrete).
added README.md +92 −0
@@ -0,0 +1,92 @@
1 +<!--
2 +README.md — Prime Mystery Engine
3 +Author: Simon-Pierre Boucher — contact@spboucher.ai
4 +-->
5 +
6 +# 🔬 Prime Mystery Engine
7 +
8 +**A verification-first computational laboratory for prime gap statistics.**
9 +
10 +![Author](https://img.shields.io/badge/author-Simon--Pierre%20Boucher-1f6feb?style=for-the-badge)
11 +![Contact](https://img.shields.io/badge/contact-contact%40spboucher.ai-2ea44f?style=for-the-badge)
12 +
13 +![Python](https://img.shields.io/badge/Python-3.9%E2%80%933.14-3776AB?logo=python&logoColor=white)
14 +![NumPy](https://img.shields.io/badge/NumPy-vectorized%20sieves-013243?logo=numpy)
15 +![Metal](https://img.shields.io/badge/GPU-Metal%20%2F%20MLX-8A2BE2?logo=apple)
16 +![Scanned](https://img.shields.io/badge/range%20scanned-10%C2%B9%C2%B3-orange)
17 +![Predictions](https://img.shields.io/badge/out--of--sample%20predictions-3%2F3%20confirmed-brightgreen)
18 +![Certificates](https://img.shields.io/badge/certificates-independently%20re--verifiable-blue)
19 +![Paper](https://img.shields.io/badge/paper-LaTeX%20%2B%20PDF-b31b1b?logo=latex)
20 +
21 +An autonomous research pipeline that studies gaps between consecutive primes under a strict
22 +epistemic protocol: **no claim without a certificate**. Every statement is labelled
23 +`PROVEN` / `CERTIFIED` / `VERIFIED UP TO X` / `CONJECTURED` / `REFUTED`, every scan is
24 +deterministic and bit-reproducible across machines, and every headline object ships with a
25 +standalone re-verification script.
26 +
27 +---
28 +
29 +## 🏆 Headline results
30 +
31 +| Result | Status |
32 +|---|---|
33 +| **Race conjecture refuted** — "N(2,x) > N(4,x) for x ≥ 10⁶" is false: first tie at end-prime **80 966 861**, first strict overtake at **80 966 933**; the lead changes forever up to 10¹³ | `REFUTED` (minimal counterexample) |
34 +| **Certified prime desert** — gap of exactly **260** after the 25-digit prime `1116336781708038449369693` (merit 4.70), built from a hybrid covering system mod primes ≤ 59 | `CERTIFIED` |
35 +| **Scaling law for consecutive-gap correlation** — ρ(x)·ln x = c + d/ln x with **c = −0.4835 ± 0.002**, d = −1.777; out-of-sample predictions confirmed at 10¹² (−0.549 ± 0.005 → −0.54756) **and** 10¹³ (−0.5432 ± 0.004 → −0.54264) | `CONJECTURED`, 2/2 hits |
36 +| **Lag-2 law** — ρ₂(x)·ln x → c₂ ≈ −0.278, prediction confirmed at 10¹³ (−0.2545 ± 0.004 → −0.25247) | `CONJECTURED`, 1/1 hit |
37 +| **First-order Hardy–Littlewood triple model rejected** — right sign and drift form, but only ~32 % of the observed magnitude | `MODEL REJECTED` |
38 +| **All published tables reproduced** — 53/53 maximal gaps below 10¹³ (OEIS A005250/A002386), twin counts, π(10ᵏ), CSG maxima | `VERIFIED` |
39 +
40 +**Predictions on record (falsifiable):** ρ·ln x(10¹⁴) = −0.5387 ± 0.004 · ρ·ln x(10¹⁵) = −0.5350 ± 0.004 · ρ₂·ln x(10¹⁴) = −0.2553 ± 0.004
41 +
42 +---
43 +
44 +## ⚡ Verify the headline claims yourself (stock Python 3)
45 +
46 +```bash
47 +# the certified 260-gap desert (independent Miller–Rabin implementation)
48 +python3 certs/verify_desert.py
49 +
50 +# the race-conjecture refutation (stdlib-only sieve, ~30 s)
51 +python3 certs/verify_c2_refutation.py
52 +
53 +# the core-primitives validation gate (37 checks, needs numpy)
54 +python3 src/test_core.py
55 +```
56 +
57 +## 🗂 Repository layout
58 +
59 +```
60 +src/ core primitives, scans, distributed worker/merger, Metal GPU sieve, models
61 +certs/ certificates + standalone verifiers (the "trust nothing" layer)
62 +data/ raw checkpoint JSONs for every campaign (deterministic, re-mergeable)
63 +paper/ LaTeX manuscript (main.tex → main.pdf, 13 pp., full math detail)
64 +journal.md complete research journal: every cycle, decision, failure
65 +conjectures.md the conjecture ledger with epistemic statuses
66 +records.md state of the art (what would count as new; thresholds to beat)
67 +report_cycle{1..4}.md per-cycle syntheses
68 +```
69 +
70 +## 🖥 Compute
71 +
72 +| Campaign | Range | Fleet | Wall time |
73 +|---|---|---|---|
74 +| Cycle 1 | 4×10⁹ | 1 laptop core | 15 s |
75 +| Cycle 2 | 4×10¹⁰ | laptop + M3 Ultra (1 core each) | ~2 min |
76 +| Cycle 3 | 10¹² | 12 Apple-silicon nodes, 162 CPU cores | ~5 min |
77 +| Cycle 4 | 10¹³ | 3 nodes: 58 CPU cores + **3 GPUs (custom Metal kernel)** | ~50 min |
78 +
79 +The distributed design uses *mergeable* streaming statistics (exact integers < 2⁵³ in
80 +float64 → **zero rounding**, bit-identical across machines) with checkpoint-aligned chunking
81 +and a junction buffer, so a 10 000-chunk merge equals the single-machine scan exactly.
82 +The Metal sieve kernel was accepted only after producing byte-identical prime sets to the
83 +CPU sieve; it processed 11.7 % of cycle 4.
84 +
85 +## 📄 Paper
86 +
87 +Full write-up with all equations, proofs of the infrastructure lemmas, robustness checks and
88 +threats-to-validity: [`paper/main.tex`](paper/main.tex) → [`paper/main.pdf`](paper/main.pdf).
89 +
90 +## 👤 Author
91 +
92 +**Simon-Pierre Boucher** · [contact@spboucher.ai](mailto:contact@spboucher.ai)
added certs/desert_certificate.json +1098 −0
@@ -0,0 +1,1098 @@
1 +{
2 + "title": "Certified prime desert via hybrid covering system",
3 + "author": "Simon-Pierre Boucher \u2014 contact@spboucher.ai",
4 + "date": "2026-08-06",
5 + "claim": "N and N+260 are consecutive primes (gap exactly 260).",
6 + "N": "1116336781708038449369693",
7 + "gap": 260,
8 + "merit": 4.69551,
9 + "csg_ratio": 0.084799,
10 + "construction": {
11 + "prime_set_max": 59,
12 + "primorial_P": "1922760350154212639070",
13 + "crt_residue_x": "1135778618595118709093",
14 + "shift_t": 580,
15 + "classes_a_p": {
16 + "2": 1,
17 + "3": 1,
18 + "5": 2,
19 + "7": 6,
20 + "11": 8,
21 + "13": 1,
22 + "17": 16,
23 + "19": 6,
24 + "23": 3,
25 + "29": 2,
26 + "31": 18,
27 + "37": 24,
28 + "41": 4,
29 + "43": 9,
30 + "47": 9,
31 + "53": 28,
32 + "59": 38
33 + },
34 + "holes": [
35 + 36,
36 + 54,
37 + 68,
38 + 78,
39 + 108,
40 + 110,
41 + 114,
42 + 116,
43 + 126,
44 + 128,
45 + 180,
46 + 194,
47 + 198,
48 + 200,
49 + 218,
50 + 236
51 + ],
52 + "note": "positions with a covering factor are composite for EVERY shift t (proven); hole positions are certified composite for THIS t by explicit factor or strong MR witness"
53 + },
54 + "interior_certificates": {
55 + "1": {
56 + "type": "covering_factor",
57 + "q": 2
58 + },
59 + "2": {
60 + "type": "covering_factor",
61 + "q": 5
62 + },
63 + "3": {
64 + "type": "covering_factor",
65 + "q": 2
66 + },
67 + "4": {
68 + "type": "covering_factor",
69 + "q": 3
70 + },
71 + "5": {
72 + "type": "covering_factor",
73 + "q": 2
74 + },
75 + "6": {
76 + "type": "covering_factor",
77 + "q": 7
78 + },
79 + "7": {
80 + "type": "covering_factor",
81 + "q": 2
82 + },
83 + "8": {
84 + "type": "covering_factor",
85 + "q": 11
86 + },
87 + "9": {
88 + "type": "covering_factor",
89 + "q": 2
90 + },
91 + "10": {
92 + "type": "covering_factor",
93 + "q": 3
94 + },
95 + "11": {
96 + "type": "covering_factor",
97 + "q": 2
98 + },
99 + "12": {
100 + "type": "covering_factor",
101 + "q": 5
102 + },
103 + "13": {
104 + "type": "covering_factor",
105 + "q": 2
106 + },
107 + "14": {
108 + "type": "covering_factor",
109 + "q": 13
110 + },
111 + "15": {
112 + "type": "covering_factor",
113 + "q": 2
114 + },
115 + "16": {
116 + "type": "covering_factor",
117 + "q": 3
118 + },
119 + "17": {
120 + "type": "covering_factor",
121 + "q": 2
122 + },
123 + "18": {
124 + "type": "covering_factor",
125 + "q": 31
126 + },
127 + "19": {
128 + "type": "covering_factor",
129 + "q": 2
130 + },
131 + "20": {
132 + "type": "covering_factor",
133 + "q": 7
134 + },
135 + "21": {
136 + "type": "covering_factor",
137 + "q": 2
138 + },
139 + "22": {
140 + "type": "covering_factor",
141 + "q": 3
142 + },
143 + "23": {
144 + "type": "covering_factor",
145 + "q": 2
146 + },
147 + "24": {
148 + "type": "covering_factor",
149 + "q": 37
150 + },
151 + "25": {
152 + "type": "covering_factor",
153 + "q": 2
154 + },
155 + "26": {
156 + "type": "covering_factor",
157 + "q": 23
158 + },
159 + "27": {
160 + "type": "covering_factor",
161 + "q": 2
162 + },
163 + "28": {
164 + "type": "covering_factor",
165 + "q": 3
166 + },
167 + "29": {
168 + "type": "covering_factor",
169 + "q": 2
170 + },
171 + "30": {
172 + "type": "covering_factor",
173 + "q": 11
174 + },
175 + "31": {
176 + "type": "covering_factor",
177 + "q": 2
178 + },
179 + "32": {
180 + "type": "covering_factor",
181 + "q": 5
182 + },
183 + "33": {
184 + "type": "covering_factor",
185 + "q": 2
186 + },
187 + "34": {
188 + "type": "covering_factor",
189 + "q": 3
190 + },
191 + "35": {
192 + "type": "covering_factor",
193 + "q": 2
194 + },
195 + "36": {
196 + "type": "mr_witness",
197 + "a": 2
198 + },
199 + "37": {
200 + "type": "covering_factor",
201 + "q": 2
202 + },
203 + "38": {
204 + "type": "covering_factor",
205 + "q": 59
206 + },
207 + "39": {
208 + "type": "covering_factor",
209 + "q": 2
210 + },
211 + "40": {
212 + "type": "covering_factor",
213 + "q": 3
214 + },
215 + "41": {
216 + "type": "covering_factor",
217 + "q": 2
218 + },
219 + "42": {
220 + "type": "covering_factor",
221 + "q": 5
222 + },
223 + "43": {
224 + "type": "covering_factor",
225 + "q": 2
226 + },
227 + "44": {
228 + "type": "covering_factor",
229 + "q": 19
230 + },
231 + "45": {
232 + "type": "covering_factor",
233 + "q": 2
234 + },
235 + "46": {
236 + "type": "covering_factor",
237 + "q": 3
238 + },
239 + "47": {
240 + "type": "covering_factor",
241 + "q": 2
242 + },
243 + "48": {
244 + "type": "covering_factor",
245 + "q": 7
246 + },
247 + "49": {
248 + "type": "covering_factor",
249 + "q": 2
250 + },
251 + "50": {
252 + "type": "covering_factor",
253 + "q": 17
254 + },
255 + "51": {
256 + "type": "covering_factor",
257 + "q": 2
258 + },
259 + "52": {
260 + "type": "covering_factor",
261 + "q": 3
262 + },
263 + "53": {
264 + "type": "covering_factor",
265 + "q": 2
266 + },
267 + "54": {
268 + "type": "trial_factor",
269 + "q": 1753
270 + },
271 + "55": {
272 + "type": "covering_factor",
273 + "q": 2
274 + },
275 + "56": {
276 + "type": "covering_factor",
277 + "q": 47
278 + },
279 + "57": {
280 + "type": "covering_factor",
281 + "q": 2
282 + },
283 + "58": {
284 + "type": "covering_factor",
285 + "q": 3
286 + },
287 + "59": {
288 + "type": "covering_factor",
289 + "q": 2
290 + },
291 + "60": {
292 + "type": "covering_factor",
293 + "q": 29
294 + },
295 + "61": {
296 + "type": "covering_factor",
297 + "q": 2
298 + },
299 + "62": {
300 + "type": "covering_factor",
301 + "q": 5
302 + },
303 + "63": {
304 + "type": "covering_factor",
305 + "q": 2
306 + },
307 + "64": {
308 + "type": "covering_factor",
309 + "q": 3
310 + },
311 + "65": {
312 + "type": "covering_factor",
313 + "q": 2
314 + },
315 + "66": {
316 + "type": "covering_factor",
317 + "q": 13
318 + },
319 + "67": {
320 + "type": "covering_factor",
321 + "q": 2
322 + },
323 + "68": {
324 + "type": "trial_factor",
325 + "q": 18401
326 + },
327 + "69": {
328 + "type": "covering_factor",
329 + "q": 2
330 + },
331 + "70": {
332 + "type": "covering_factor",
333 + "q": 3
334 + },
335 + "71": {
336 + "type": "covering_factor",
337 + "q": 2
338 + },
339 + "72": {
340 + "type": "covering_factor",
341 + "q": 5
342 + },
343 + "73": {
344 + "type": "covering_factor",
345 + "q": 2
346 + },
347 + "74": {
348 + "type": "covering_factor",
349 + "q": 11
350 + },
351 + "75": {
352 + "type": "covering_factor",
353 + "q": 2
354 + },
355 + "76": {
356 + "type": "covering_factor",
357 + "q": 3
358 + },
359 + "77": {
360 + "type": "covering_factor",
361 + "q": 2
362 + },
363 + "78": {
364 + "type": "trial_factor",
365 + "q": 4211
366 + },
367 + "79": {
368 + "type": "covering_factor",
369 + "q": 2
370 + },
371 + "80": {
372 + "type": "covering_factor",
373 + "q": 31
374 + },
375 + "81": {
376 + "type": "covering_factor",
377 + "q": 2
378 + },
379 + "82": {
380 + "type": "covering_factor",
381 + "q": 3
382 + },
383 + "83": {
384 + "type": "covering_factor",
385 + "q": 2
386 + },
387 + "84": {
388 + "type": "covering_factor",
389 + "q": 17
390 + },
391 + "85": {
392 + "type": "covering_factor",
393 + "q": 2
394 + },
395 + "86": {
396 + "type": "covering_factor",
397 + "q": 41
398 + },
399 + "87": {
400 + "type": "covering_factor",
401 + "q": 2
402 + },
403 + "88": {
404 + "type": "covering_factor",
405 + "q": 3
406 + },
407 + "89": {
408 + "type": "covering_factor",
409 + "q": 2
410 + },
411 + "90": {
412 + "type": "covering_factor",
413 + "q": 7
414 + },
415 + "91": {
416 + "type": "covering_factor",
417 + "q": 2
418 + },
419 + "92": {
420 + "type": "covering_factor",
421 + "q": 5
422 + },
423 + "93": {
424 + "type": "covering_factor",
425 + "q": 2
426 + },
427 + "94": {
428 + "type": "covering_factor",
429 + "q": 3
430 + },
431 + "95": {
432 + "type": "covering_factor",
433 + "q": 2
434 + },
435 + "96": {
436 + "type": "covering_factor",
437 + "q": 11
438 + },
439 + "97": {
440 + "type": "covering_factor",
441 + "q": 2
442 + },
443 + "98": {
444 + "type": "covering_factor",
445 + "q": 37
446 + },
447 + "99": {
448 + "type": "covering_factor",
449 + "q": 2
450 + },
451 + "100": {
452 + "type": "covering_factor",
453 + "q": 3
454 + },
455 + "101": {
456 + "type": "covering_factor",
457 + "q": 2
458 + },
459 + "102": {
460 + "type": "covering_factor",
461 + "q": 5
462 + },
463 + "103": {
464 + "type": "covering_factor",
465 + "q": 2
466 + },
467 + "104": {
468 + "type": "covering_factor",
469 + "q": 7
470 + },
471 + "105": {
472 + "type": "covering_factor",
473 + "q": 2
474 + },
475 + "106": {
476 + "type": "covering_factor",
477 + "q": 3
478 + },
479 + "107": {
480 + "type": "covering_factor",
481 + "q": 2
482 + },
483 + "108": {
484 + "type": "trial_factor",
485 + "q": 97
486 + },
487 + "109": {
488 + "type": "covering_factor",
489 + "q": 2
490 + },
491 + "110": {
492 + "type": "trial_factor",
493 + "q": 73
494 + },
495 + "111": {
496 + "type": "covering_factor",
497 + "q": 2
498 + },
499 + "112": {
500 + "type": "covering_factor",
501 + "q": 3
502 + },
503 + "113": {
504 + "type": "covering_factor",
505 + "q": 2
506 + },
507 + "114": {
508 + "type": "trial_factor",
509 + "q": 139
510 + },
511 + "115": {
512 + "type": "covering_factor",
513 + "q": 2
514 + },
515 + "116": {
516 + "type": "trial_factor",
517 + "q": 62119
518 + },
519 + "117": {
520 + "type": "covering_factor",
521 + "q": 2
522 + },
523 + "118": {
524 + "type": "covering_factor",
525 + "q": 3
526 + },
527 + "119": {
528 + "type": "covering_factor",
529 + "q": 2
530 + },
531 + "120": {
532 + "type": "covering_factor",
533 + "q": 19
534 + },
535 + "121": {
536 + "type": "covering_factor",
537 + "q": 2
538 + },
539 + "122": {
540 + "type": "covering_factor",
541 + "q": 5
542 + },
543 + "123": {
544 + "type": "covering_factor",
545 + "q": 2
546 + },
547 + "124": {
548 + "type": "covering_factor",
549 + "q": 3
550 + },
551 + "125": {
552 + "type": "covering_factor",
553 + "q": 2
554 + },
555 + "126": {
556 + "type": "trial_factor",
557 + "q": 4813
558 + },
559 + "127": {
560 + "type": "covering_factor",
561 + "q": 2
562 + },
563 + "128": {
564 + "type": "trial_factor",
565 + "q": 167
566 + },
567 + "129": {
568 + "type": "covering_factor",
569 + "q": 2
570 + },
571 + "130": {
572 + "type": "covering_factor",
573 + "q": 3
574 + },
575 + "131": {
576 + "type": "covering_factor",
577 + "q": 2
578 + },
579 + "132": {
580 + "type": "covering_factor",
581 + "q": 5
582 + },
583 + "133": {
584 + "type": "covering_factor",
585 + "q": 2
586 + },
587 + "134": {
588 + "type": "covering_factor",
589 + "q": 53
590 + },
591 + "135": {
592 + "type": "covering_factor",
593 + "q": 2
594 + },
595 + "136": {
596 + "type": "covering_factor",
597 + "q": 3
598 + },
599 + "137": {
600 + "type": "covering_factor",
601 + "q": 2
602 + },
603 + "138": {
604 + "type": "covering_factor",
605 + "q": 43
606 + },
607 + "139": {
608 + "type": "covering_factor",
609 + "q": 2
610 + },
611 + "140": {
612 + "type": "covering_factor",
613 + "q": 11
614 + },
615 + "141": {
616 + "type": "covering_factor",
617 + "q": 2
618 + },
619 + "142": {
620 + "type": "covering_factor",
621 + "q": 3
622 + },
623 + "143": {
624 + "type": "covering_factor",
625 + "q": 2
626 + },
627 + "144": {
628 + "type": "covering_factor",
629 + "q": 13
630 + },
631 + "145": {
632 + "type": "covering_factor",
633 + "q": 2
634 + },
635 + "146": {
636 + "type": "covering_factor",
637 + "q": 7
638 + },
639 + "147": {
640 + "type": "covering_factor",
641 + "q": 2
642 + },
643 + "148": {
644 + "type": "covering_factor",
645 + "q": 3
646 + },
647 + "149": {
648 + "type": "covering_factor",
649 + "q": 2
650 + },
651 + "150": {
652 + "type": "covering_factor",
653 + "q": 47
654 + },
655 + "151": {
656 + "type": "covering_factor",
657 + "q": 2
658 + },
659 + "152": {
660 + "type": "covering_factor",
661 + "q": 5
662 + },
663 + "153": {
664 + "type": "covering_factor",
665 + "q": 2
666 + },
667 + "154": {
668 + "type": "covering_factor",
669 + "q": 3
670 + },
671 + "155": {
672 + "type": "covering_factor",
673 + "q": 2
674 + },
675 + "156": {
676 + "type": "covering_factor",
677 + "q": 59
678 + },
679 + "157": {
680 + "type": "covering_factor",
681 + "q": 2
682 + },
683 + "158": {
684 + "type": "covering_factor",
685 + "q": 19
686 + },
687 + "159": {
688 + "type": "covering_factor",
689 + "q": 2
690 + },
691 + "160": {
692 + "type": "covering_factor",
693 + "q": 3
694 + },
695 + "161": {
696 + "type": "covering_factor",
697 + "q": 2
698 + },
699 + "162": {
700 + "type": "covering_factor",
701 + "q": 5
702 + },
703 + "163": {
704 + "type": "covering_factor",
705 + "q": 2
706 + },
707 + "164": {
708 + "type": "covering_factor",
709 + "q": 23
710 + },
711 + "165": {
712 + "type": "covering_factor",
713 + "q": 2
714 + },
715 + "166": {
716 + "type": "covering_factor",
717 + "q": 3
718 + },
719 + "167": {
720 + "type": "covering_factor",
721 + "q": 2
722 + },
723 + "168": {
724 + "type": "covering_factor",
725 + "q": 41
726 + },
727 + "169": {
728 + "type": "covering_factor",
729 + "q": 2
730 + },
731 + "170": {
732 + "type": "covering_factor",
733 + "q": 13
734 + },
735 + "171": {
736 + "type": "covering_factor",
737 + "q": 2
738 + },
739 + "172": {
740 + "type": "covering_factor",
741 + "q": 3
742 + },
743 + "173": {
744 + "type": "covering_factor",
745 + "q": 2
746 + },
747 + "174": {
748 + "type": "covering_factor",
749 + "q": 7
750 + },
751 + "175": {
752 + "type": "covering_factor",
753 + "q": 2
754 + },
755 + "176": {
756 + "type": "covering_factor",
757 + "q": 29
758 + },
759 + "177": {
760 + "type": "covering_factor",
761 + "q": 2
762 + },
763 + "178": {
764 + "type": "covering_factor",
765 + "q": 3
766 + },
767 + "179": {
768 + "type": "covering_factor",
769 + "q": 2
770 + },
771 + "180": {
772 + "type": "trial_factor",
773 + "q": 131
774 + },
775 + "181": {
776 + "type": "covering_factor",
777 + "q": 2
778 + },
779 + "182": {
780 + "type": "covering_factor",
781 + "q": 5
782 + },
783 + "183": {
784 + "type": "covering_factor",
785 + "q": 2
786 + },
787 + "184": {
788 + "type": "covering_factor",
789 + "q": 3
790 + },
791 + "185": {
792 + "type": "covering_factor",
793 + "q": 2
794 + },
795 + "186": {
796 + "type": "covering_factor",
797 + "q": 17
798 + },
799 + "187": {
800 + "type": "covering_factor",
801 + "q": 2
802 + },
803 + "188": {
804 + "type": "covering_factor",
805 + "q": 7
806 + },
807 + "189": {
808 + "type": "covering_factor",
809 + "q": 2
810 + },
811 + "190": {
812 + "type": "covering_factor",
813 + "q": 3
814 + },
815 + "191": {
816 + "type": "covering_factor",
817 + "q": 2
818 + },
819 + "192": {
820 + "type": "covering_factor",
821 + "q": 5
822 + },
823 + "193": {
824 + "type": "covering_factor",
825 + "q": 2
826 + },
827 + "194": {
828 + "type": "trial_factor",
829 + "q": 1613
830 + },
831 + "195": {
832 + "type": "covering_factor",
833 + "q": 2
834 + },
835 + "196": {
836 + "type": "covering_factor",
837 + "q": 3
838 + },
839 + "197": {
840 + "type": "covering_factor",
841 + "q": 2
842 + },
843 + "198": {
844 + "type": "trial_factor",
845 + "q": 751
846 + },
847 + "199": {
848 + "type": "covering_factor",
849 + "q": 2
850 + },
851 + "200": {
852 + "type": "trial_factor",
853 + "q": 10427
854 + },
855 + "201": {
856 + "type": "covering_factor",
857 + "q": 2
858 + },
859 + "202": {
860 + "type": "covering_factor",
861 + "q": 3
862 + },
863 + "203": {
864 + "type": "covering_factor",
865 + "q": 2
866 + },
867 + "204": {
868 + "type": "covering_factor",
869 + "q": 31
870 + },
871 + "205": {
872 + "type": "covering_factor",
873 + "q": 2
874 + },
875 + "206": {
876 + "type": "covering_factor",
877 + "q": 11
878 + },
879 + "207": {
880 + "type": "covering_factor",
881 + "q": 2
882 + },
883 + "208": {
884 + "type": "covering_factor",
885 + "q": 3
886 + },
887 + "209": {
888 + "type": "covering_factor",
889 + "q": 2
890 + },
891 + "210": {
892 + "type": "covering_factor",
893 + "q": 23
894 + },
895 + "211": {
896 + "type": "covering_factor",
897 + "q": 2
898 + },
899 + "212": {
900 + "type": "covering_factor",
901 + "q": 5
902 + },
903 + "213": {
904 + "type": "covering_factor",
905 + "q": 2
906 + },
907 + "214": {
908 + "type": "covering_factor",
909 + "q": 3
910 + },
911 + "215": {
912 + "type": "covering_factor",
913 + "q": 2
914 + },
915 + "216": {
916 + "type": "covering_factor",
917 + "q": 7
918 + },
919 + "217": {
920 + "type": "covering_factor",
921 + "q": 2
922 + },
923 + "218": {
924 + "type": "trial_factor",
925 + "q": 3637
926 + },
927 + "219": {
928 + "type": "covering_factor",
929 + "q": 2
930 + },
931 + "220": {
932 + "type": "covering_factor",
933 + "q": 3
934 + },
935 + "221": {
936 + "type": "covering_factor",
937 + "q": 2
938 + },
939 + "222": {
940 + "type": "covering_factor",
941 + "q": 5
942 + },
943 + "223": {
944 + "type": "covering_factor",
945 + "q": 2
946 + },
947 + "224": {
948 + "type": "covering_factor",
949 + "q": 43
950 + },
951 + "225": {
952 + "type": "covering_factor",
953 + "q": 2
954 + },
955 + "226": {
956 + "type": "covering_factor",
957 + "q": 3
958 + },
959 + "227": {
960 + "type": "covering_factor",
961 + "q": 2
962 + },
963 + "228": {
964 + "type": "covering_factor",
965 + "q": 11
966 + },
967 + "229": {
968 + "type": "covering_factor",
969 + "q": 2
970 + },
971 + "230": {
972 + "type": "covering_factor",
973 + "q": 7
974 + },
975 + "231": {
976 + "type": "covering_factor",
977 + "q": 2
978 + },
979 + "232": {
980 + "type": "covering_factor",
981 + "q": 3
982 + },
983 + "233": {
984 + "type": "covering_factor",
985 + "q": 2
986 + },
987 + "234": {
988 + "type": "covering_factor",
989 + "q": 19
990 + },
991 + "235": {
992 + "type": "covering_factor",
993 + "q": 2
994 + },
995 + "236": {
996 + "type": "trial_factor",
997 + "q": 349
998 + },
999 + "237": {
1000 + "type": "covering_factor",
1001 + "q": 2
1002 + },
1003 + "238": {
1004 + "type": "covering_factor",
1005 + "q": 3
1006 + },
1007 + "239": {
1008 + "type": "covering_factor",
1009 + "q": 2
1010 + },
1011 + "240": {
1012 + "type": "covering_factor",
1013 + "q": 53
1014 + },
1015 + "241": {
1016 + "type": "covering_factor",
1017 + "q": 2
1018 + },
1019 + "242": {
1020 + "type": "covering_factor",
1021 + "q": 5
1022 + },
1023 + "243": {
1024 + "type": "covering_factor",
1025 + "q": 2
1026 + },
1027 + "244": {
1028 + "type": "covering_factor",
1029 + "q": 3
1030 + },
1031 + "245": {
1032 + "type": "covering_factor",
1033 + "q": 2
1034 + },
1035 + "246": {
1036 + "type": "covering_factor",
1037 + "q": 37
1038 + },
1039 + "247": {
1040 + "type": "covering_factor",
1041 + "q": 2
1042 + },
1043 + "248": {
1044 + "type": "covering_factor",
1045 + "q": 13
1046 + },
1047 + "249": {
1048 + "type": "covering_factor",
1049 + "q": 2
1050 + },
1051 + "250": {
1052 + "type": "covering_factor",
1053 + "q": 3
1054 + },
1055 + "251": {
1056 + "type": "covering_factor",
1057 + "q": 2
1058 + },
1059 + "252": {
1060 + "type": "covering_factor",
1061 + "q": 5
1062 + },
1063 + "253": {
1064 + "type": "covering_factor",
1065 + "q": 2
1066 + },
1067 + "254": {
1068 + "type": "covering_factor",
1069 + "q": 17
1070 + },
1071 + "255": {
1072 + "type": "covering_factor",
1073 + "q": 2
1074 + },
1075 + "256": {
1076 + "type": "covering_factor",
1077 + "q": 3
1078 + },
1079 + "257": {
1080 + "type": "covering_factor",
1081 + "q": 2
1082 + },
1083 + "258": {
1084 + "type": "covering_factor",
1085 + "q": 7
1086 + },
1087 + "259": {
1088 + "type": "covering_factor",
1089 + "q": 2
1090 + }
1091 + },
1092 + "endpoint_primality": {
1093 + "method": "deterministic Miller-Rabin, bases {2..37} (12 bases), valid for n < 3.317e24 (Sorenson-Webster 2015)",
1094 + "validity_bound": "3317044064679887385961981",
1095 + "both_endpoints_below_bound": true
1096 + },
1097 + "epistemic_status": "CERTIFIED construction; NOT a record (see records.md: merit record 41.94, largest maximal gap 1854). Baseline with same primes: classic primorial gives gap 61; pure covering (attempt 1) gave gap 90."
1098 +}
\ No newline at end of file
added certs/verify_c2_refutation.py +71 −0
@@ -0,0 +1,71 @@
1 +#!/usr/bin/env python3
2 +# =============================================================================
3 +# verify_c2_refutation.py — INDEPENDENT re-verification of the C2 counterexample
4 +# Author: Simon-Pierre Boucher — contact@spboucher.ai
5 +# =============================================================================
6 +# Claim refuted: "N(2,x) > N(4,x) for all x >= 10^6" (C2, cycle 1)
7 +# where N(g,x) counts consecutive prime pairs (p, p+g) with p+g <= x.
8 +# This script recomputes, stdlib-only (bytearray sieve, no numpy, no /src):
9 +# - D(p) = N2 - N4 after each prime p (gaps attributed to their END prime),
10 +# - checks D(p) > 0 for every end-prime p in (10^6, 80966861),
11 +# - checks D = 0 at end-prime 80966861 (first tie),
12 +# - checks D = -1 at end-prime 80966933 (first strict overtake by gap 4).
13 +# Runtime ~ 30 s pure Python. Exit 0 + "VERIFICATION PASSED" iff all hold.
14 +# =============================================================================
15 +
16 +import sys
17 +
18 +LIMIT = 81_000_000
19 +FIRST_TIE = 80_966_861
20 +FIRST_NEG = 80_966_933
21 +
22 +
23 +def main():
24 + sieve = bytearray([1]) * (LIMIT + 1)
25 + sieve[0] = sieve[1] = 0
26 + for i in range(2, int(LIMIT ** 0.5) + 1):
27 + if sieve[i]:
28 + sieve[i * i:: i] = bytearray(len(sieve[i * i:: i]))
29 + D = 0
30 + prev = 2
31 + ok_positive = True
32 + tie_val = neg_val = None
33 + min_after_1e6 = None
34 + for p in range(3, LIMIT + 1, 2):
35 + if not sieve[p]:
36 + continue
37 + g = p - prev
38 + if g == 2:
39 + D += 1
40 + elif g == 4:
41 + D -= 1
42 + prev = p
43 + if 10**6 < p < FIRST_TIE and D <= 0:
44 + ok_positive = False
45 + print(f"unexpected D<=0 at {p}: D={D}")
46 + if p == FIRST_TIE:
47 + tie_val = D
48 + if p == FIRST_NEG:
49 + neg_val = D
50 + failures = []
51 + if not ok_positive:
52 + failures.append("D not strictly positive on (10^6, 80966861)")
53 + if tie_val != 0:
54 + failures.append(f"D at {FIRST_TIE} is {tie_val}, expected 0")
55 + if neg_val != -1:
56 + failures.append(f"D at {FIRST_NEG} is {neg_val}, expected -1")
57 + if failures:
58 + print("VERIFICATION FAILED:")
59 + for f in failures:
60 + print(" -", f)
61 + sys.exit(1)
62 + print("VERIFICATION PASSED:")
63 + print(f" D(p) > 0 for every end-prime p in (10^6, {FIRST_TIE})")
64 + print(f" D({FIRST_TIE}) = 0 (first tie after 10^6)")
65 + print(f" D({FIRST_NEG}) = -1 (first strict N4 > N2 after 10^6)")
66 + print(" => C2 is REFUTED; minimal counterexample x = 80966861 (tie), "
67 + "80966933 (strict).")
68 +
69 +
70 +if __name__ == "__main__":
71 + main()
added certs/verify_desert.py +105 −0
@@ -0,0 +1,105 @@
1 +#!/usr/bin/env python3
2 +# =============================================================================
3 +# verify_desert.py — INDEPENDENT re-verification of desert_certificate.json
4 +# Author: Simon-Pierre Boucher — contact@spboucher.ai
5 +# =============================================================================
6 +# Standalone: stdlib only, NO imports from /src (independent implementation).
7 +# Verifies the claim "N and N+gap are consecutive primes":
8 +# (1) every interior position 1..gap-1 has a compositeness certificate and
9 +# the certificate checks out (divisor divides / MR witness is a witness);
10 +# (2) both endpoints pass deterministic Miller-Rabin (12 bases, valid
11 +# for n < 3.317e24) and lie below that bound;
12 +# (3) the covering classes are consistent with N (N+i = 0 mod q claimed).
13 +# Exit code 0 + "VERIFICATION PASSED" iff everything holds.
14 +# Usage: python3 verify_desert.py [certificate.json]
15 +# =============================================================================
16 +
17 +import json
18 +import sys
19 +from pathlib import Path
20 +
21 +MR_LIMIT = 3_317_044_064_679_887_385_961_981
22 +BASES = (2, 3, 5, 7, 11, 13, 17, 19, 23, 29, 31, 37)
23 +
24 +
25 +def mr_decompose(n):
26 + d, r = n - 1, 0
27 + while d % 2 == 0:
28 + d //= 2
29 + r += 1
30 + return d, r
31 +
32 +
33 +def is_strong_witness(n, a):
34 + """True iff a proves n composite (strong Miller-Rabin witness)."""
35 + d, r = mr_decompose(n)
36 + x = pow(a, d, n)
37 + if x == 1 or x == n - 1:
38 + return False
39 + for _ in range(r - 1):
40 + x = x * x % n
41 + if x == n - 1:
42 + return False
43 + return True
44 +
45 +
46 +def is_prime_det(n):
47 + """Deterministic MR for n < 3.317e24."""
48 + assert n < MR_LIMIT, "outside deterministic validity range"
49 + if n < 2:
50 + return False
51 + for p in BASES:
52 + if n % p == 0:
53 + return n == p
54 + return not any(is_strong_witness(n, a) for a in BASES)
55 +
56 +
57 +def main():
58 + path = sys.argv[1] if len(sys.argv) > 1 else Path(__file__).parent / "desert_certificate.json"
59 + cert = json.load(open(path))
60 + N = int(cert["N"])
61 + gap = int(cert["gap"])
62 + interior = cert["interior_certificates"]
63 + failures = []
64 +
65 + # (1) completeness: every interior position certified
66 + missing = [i for i in range(1, gap) if str(i) not in interior]
67 + if missing:
68 + failures.append(f"positions without certificate: {missing}")
69 +
70 + # (2) each interior certificate is valid
71 + for i_str, c in interior.items():
72 + i = int(i_str)
73 + v = N + i
74 + if c["type"] in ("covering_factor", "trial_factor"):
75 + q = int(c["q"])
76 + if v % q != 0 or v <= q:
77 + failures.append(f"position {i}: claimed factor {q} invalid")
78 + elif c["type"] == "mr_witness":
79 + if not is_strong_witness(v, int(c["a"])):
80 + failures.append(f"position {i}: base {c['a']} is NOT a witness")
81 + else:
82 + failures.append(f"position {i}: unknown certificate type {c['type']}")
83 +
84 + # (3) endpoints prime, within deterministic range
85 + if N + gap >= MR_LIMIT:
86 + failures.append("endpoint exceeds deterministic MR bound")
87 + else:
88 + if not is_prime_det(N):
89 + failures.append("N is not prime")
90 + if not is_prime_det(N + gap):
91 + failures.append(f"N+{gap} is not prime")
92 +
93 + if failures:
94 + print("VERIFICATION FAILED:")
95 + for f in failures:
96 + print(" -", f)
97 + sys.exit(1)
98 + print(f"VERIFICATION PASSED: {N} and {N}+{gap} are consecutive primes;")
99 + print(f"all {gap-1} interior numbers certified composite "
100 + f"(covering/trial factors + MR witnesses).")
101 + print(f"merit = {cert['merit']}, csg = {cert['csg_ratio']}")
102 +
103 +
104 +if __name__ == "__main__":
105 + main()
added conjectures.md +121 −0
@@ -0,0 +1,121 @@
1 +<!--
2 +conjectures.md — Prime Mystery Engine: conjecture ledger
3 +Author: Simon-Pierre Boucher — contact@spboucher.ai
4 +-->
5 +
6 +# Conjecture Ledger
7 +
8 +Status vocabulary: PROVEN / VERIFIED UP TO X / CONJECTURED / REFUTED (minimal counterexample) / SPECULATIVE.
9 +Notation: N(g,x) = #{consecutive prime pairs (p, p+g), p+g < x}. ρ(x) = Pearson correlation of
10 +(g_n, g_{n+1}) over consecutive gaps below x. CSG(p) = g(p)/ln²p. G6(x) = largest G such that every
11 +multiple of 6 ≤ G satisfies N(g,x) > N(g−2,x) and N(g,x) > N(g+2,x).
12 +
13 +## Cycle 1 (2026-08-06) — discovery range 10^8, deterministic scans (no seed)
14 +
15 +### C1 — Jumping champion
16 +**Statement.** For all x ∈ [10^5, 4×10^9], argmax_g N(g,x) = 6.
17 +**Classification.** (a) probably known — Odlyzko–Rubinstein–Wolf conjecture territory (champion 6 from ≈947 to ≈1.7×10^35).
18 +**Modular admissibility.** Consistent (6 = 2·3 maximizes the singular series among small gaps).
19 +**Status at discovery.** VERIFIED at checkpoints 10^6, 10^7, 10^8.
20 +**FINAL STATUS (adversary, 4×10^9).** SURVIVOR — VERIFIED UP TO 4×10^9 (champion = 6 at every checkpoint).
21 +Cramér/Maier note: NOT merely statistical — a Cramér random model has no mod-6 structure at all; this is
22 +structural (singular series). Known conjecture (ORW); we add nothing but an independent check.
23 +
24 +### C2 — The 2-vs-4 race never crosses again
25 +**Statement.** For all x ≥ 10^6, N(2,x) > N(4,x).
26 +**Classification.** (c) potentially interesting — asymptotically N(2,x) ~ N(4,x) under Hardy–Littlewood
27 +(identical singular series), so the sign of the difference is a second-order effect analogous to
28 +Chebyshev bias; margins observed are tiny (+26 at 10^6, +359 at 10^7, +55 at 10^8, i.e. ~10^-4 relative)
29 +and non-monotone, so a crossing is plausible. Possibly known in race literature.
30 +**Modular admissibility.** No obstruction (both gaps admissible).
31 +**Status at discovery.** VERIFIED at checkpoints 10^6, 10^7, 10^8. ⚠ Checkpoint-only: interior crossings not yet excluded.
32 +**FINAL STATUS (adversary, full resolution to 4×10^9).** **REFUTED.**
33 +Minimal counterexample above 10^6: first tie D = N2−N4 = 0 at end-prime **80966861**; first strict
34 +overtake (N4 > N2) at end-prime **80966933** (D = −1). The checkpoint verification was misleading:
35 +the race in fact changes leader forever in the tested range — D(10^9) = −173, D(2×10^9) = +1074,
36 +D(4×10^9) = −2270, with 137,574,763 gap-events where D ≤ 0 and the last one at 3999999979 (= the
37 +last prime scanned). Cramér/Maier note: consistent with a random-walk model of the difference
38 +(N(2) and N(4) have identical singular series) — the refutation is exactly what the random model
39 +predicts; the interesting open follow-up is whether the *logarithmic density* of the lead is biased
40 +(Chebyshev-style). Re-verify: `python3 certs/verify_c2_refutation.py` (stdlib-only, ~30 s).
41 +
42 +### C3 — Local dominance of multiples of 6
43 +**Statement.** G6(x) ≥ 66 for all x ≥ 10^6, and G6(x) → ∞.
44 +**Classification.** (b) easy consequence of Hardy–Littlewood heuristics (singular series of 6k beats
45 +neighbors); the finite claim is the testable part — the threshold 66 at 10^8 is limited by sample noise
46 +in the histogram tail, so G6 should grow with x.
47 +**Status at discovery.** G6 = 66 exactly at 10^6, 10^7 and 10^8 (fails at 72 each time — noise or structure? adversary must decide).
48 +**FINAL STATUS (adversary, 4×10^9).** SURVIVOR, strengthened — G6(10^9) = G6(4×10^9) = **216** ≥ 66.
49 +The failure at 72 below 10^8 was histogram-tail noise, not structure: G6 grows with x as predicted.
50 +VERIFIED UP TO 4×10^9; the G6(x) → ∞ part remains CONJECTURED (would follow from Hardy–Littlewood).
51 +
52 +### C4 — Scaled anticorrelation of consecutive gaps
53 +**Statement.** ρ(x) < 0 for all x ≥ 10^6, and ρ(x)·ln x → c with c ≈ −0.6; sharply: ρ(x)·ln x ∈ [−0.65, −0.55]
54 +for all x ∈ [10^6, 4×10^9].
55 +**Data.** ρ·ln x = −0.596 (10^6), −0.614 (10^7), −0.582 (10^8).
56 +**Classification.** (a/c) anticorrelation itself is known empirically; the precise scaling constant ≈ −0.6
57 +is a sharper claim, possibly known (gap correlation literature), possibly a clean new datum.
58 +**Status at discovery.** VERIFIED at 3 checkpoints.
59 +**FINAL STATUS (adversary, 4×10^9).** SURVIVOR with **revision**. ρ(x) < 0 and ρ·ln x ∈ [−0.65, −0.55]
60 +both VERIFIED UP TO 4×10^9 — but the sequence ρ·ln x = −0.596, −0.614, −0.582, −0.571, −0.567, −0.565
61 +(x = 10^6 … 4×10^9) drifts monotonically upward after 10^7, so the "limit ≈ −0.6" clause is retracted;
62 +revised claim: ρ(x)·ln x converges to some c ∈ [−0.60, −0.50] — CONJECTURED, next test at 10^10+.
63 +Cramér/Maier note: structural, NOT statistical — a Cramér model with independent gaps gives ρ = 0;
64 +the anticorrelation is sieve-induced. This is the most promising quantitative lead of cycle 1.
65 +**CYCLE 2 STATUS (laptop to 10^10 + M3U96a to 4×10^10, bit-for-bit identical on overlap).**
66 +Window [−0.65, −0.55]: VERIFIED UP TO 4×10^10 (ρ·ln x = −0.55569 there) — but now expected to fail:
67 +the 9-checkpoint fit (see C4′) predicts exit of the window near x ≈ 5×10^11. Superseded by C4′.
68 +
69 +### C4′ — refined scaling law for the gap anticorrelation (NEW, cycle 2)
70 +**Statement.** ρ(x)·ln x = c + d/ln x + o(1/ln x) with c = −0.486 ± 0.010 and d ≈ −1.72.
71 +Falsifiable near-term prediction: ρ(x)·ln x = −0.549 ± 0.005 at x = 10^12.
72 +**Evidence.** Least-squares on 9 checkpoints 10^8…4×10^10: c = −0.48618, d = −1.7226,
73 +max |residual| = 0.0025, leave-one-out c ∈ [−0.494, −0.482]. Data: `data/fit_c4.json`.
74 +**Model support (PROVER, cycle 2).** The first-order HL triple-correlation model
75 +(`src/model_c4.py`: P(g1,g2) ∝ S({0,g1,g1+g2})·exp(−(g1+g2)/λ)) predicts the correct SIGN and the
76 +same c + d/λ drift FORM, but magnitude ≈ −0.154 vs observed ≈ −0.49 (factor ~3.2 too small):
77 +first-order model quantitatively REJECTED — interior-compositeness (inclusion–exclusion) terms must
78 +carry most of the effect. Open problem for a future cycle.
79 +**Classification.** (c) potentially new as a precise constant; the phenomenon is known qualitatively.
80 +**Status.** CONJECTURED; VERIFIED-consistent up to 4×10^10.
81 +**CYCLE 3 STATUS (distributed scan to 10^12, 12 nodes / 1000 chunks / 162 cores).**
82 +**PREDICTION CONFIRMED**: observed ρ·ln x(10^12) = −0.54756, inside the predicted band −0.549 ± 0.005.
83 +The cycle-1 window [−0.65, −0.55] was exited between 4×10^11 (−0.55027) and 10^12 (−0.54756), matching
84 +the predicted exit ≈ 5×10^11 — old C4 now REFUTED-as-predicted. Combined 13-point fit (10^8…10^12):
85 +**c = −0.4845 ± 0.003 (LOO), d = −1.757**, max residual 0.0025. New falsifiable predictions:
86 +ρ·ln x = −0.5432 ± 0.004 at 10^13; −0.5390 ± 0.004 at 10^14.
87 +Cross-validation: distributed merge reproduces cycle-2 single-machine checkpoints exactly at
88 +10^10/2×10^10/4×10^10. Status: CONJECTURED and now twice-tested; VERIFIED-consistent up to 10^12.
89 +**CYCLE 4 STATUS (3 nodes, 58 CPU cores + 3 Metal GPUs, 10,000 chunks to 10^13).**
90 +**SECOND PREDICTION CONFIRMED**: observed ρ·ln x(10^13) = −0.54264, predicted −0.5432 ± 0.004.
91 +All 10 overlapping checkpoints identical to cycles 2/3 (exact). 16-point refit (10^8…10^13):
92 +**c = −0.48354 (LOO [−0.48570, −0.48165]), d = −1.7772**, max residual 0.00255.
93 +Next predictions: ρ·ln x(10^14) = −0.5387 ± 0.004; ρ·ln x(10^15) = −0.5350 ± 0.004.
94 +Status: CONJECTURED, twice-confirmed out of sample; VERIFIED-consistent up to 10^13.
95 +
96 +### C6 — lag-2 anticorrelation (NEW, cycle 2)
97 +**Statement.** ρ₂(x) < 0 (correlation of (g_n, g_{n+2})) with ρ₂(x)·ln x → c₂, c₂ ≈ −0.29 ± 0.03
98 +(2-point fit of the c + d/λ form: c₂ = −0.292, d₂ ≈ +1.05 — note the drift has OPPOSITE sign to C4′).
99 +**Data.** ρ₂·ln x: −0.235 (10^8) → −0.248 (10^10) → −0.249 (4×10^10), monotone decreasing.
100 +**Classification.** (c) same family as C4′; less tested (fit uses only endpoints — firm up in cycle 3).
101 +**Status.** CONJECTURED; ρ₂ < 0 VERIFIED UP TO 4×10^10.
102 +**CYCLE 3 STATUS.** ρ₂·ln x continues its monotone drift down: −0.25141 at 10^12. Proper 13-point fit:
103 +**c₂ = −0.2806, d₂ = +0.782** (opposite-sign drift vs lag-1 confirmed). ρ₂ < 0 VERIFIED UP TO 10^12;
104 +c₂ CONJECTURED ≈ −0.28 ± 0.02. Prediction: ρ₂·ln x(10^13) = −0.2545 ± 0.004.
105 +**CYCLE 4 STATUS. PREDICTION CONFIRMED**: observed ρ₂·ln x(10^13) = −0.25247 ∈ −0.2545 ± 0.004.
106 +16-point refit: c₂ = −0.27762, d₂ = +0.7187. Next prediction: ρ₂·ln x(10^14) = −0.2553 ± 0.004.
107 +ρ₂ < 0 VERIFIED UP TO 10^13.
108 +
109 +### C5 — Cramér–Shanks–Granville ratio below 4×10^9
110 +**Statement.** max_{p ≤ 4×10^9} CSG(p) = 210/ln²(20831323) ≈ 0.7395, attained at p = 20831323 (gap 210).
111 +**Classification.** (a) known — the maximal-gap table is exhaustively verified far beyond 4×10^9;
112 +this is an independent re-verification, not a discovery.
113 +**Status at discovery.** VERIFIED to 10^8 (max CSG 0.7395); the claim to 4×10^9 requires that
114 +no later maximal gap below 4×10^9 exceeds it (adversary run).
115 +**FINAL STATUS (adversary, 4×10^9).** VERIFIED UP TO 4×10^9 — max CSG = 0.73947 at p = 20831323;
116 +all 32 maximal gaps below 4×10^9 reproduced and equal to the published table (A005250/A002386),
117 +ending with gap 336 after 3842610773. Independent re-verification of known results, as intended.
118 +
119 +### Discarded at birth (Phase 4 filter)
120 +- "Odd gaps never occur beyond (2,3)" — trivial (parity), not logged as a conjecture.
121 +- Any conjecture about gaps ≡ 3 mod 6 etc. — trivial modular obstruction (all gaps beyond (2,3) are even).
added data/adversary_4e9.json +475 −0
@@ -0,0 +1,475 @@
1 +{
2 + "limit": 4000000000,
3 + "scan_seconds": 14.3,
4 + "C2_race": {
5 + "final_D": -2270,
6 + "n_events_D_nonpositive": 137574763,
7 + "last_prime_with_D_nonpositive": 3999999979,
8 + "first_zero_epochs_sample": [
9 + [
10 + 3,
11 + 0
12 + ],
13 + [
14 + 101,
15 + 0
16 + ],
17 + [
18 + 107,
19 + 0
20 + ],
21 + [
22 + 113,
23 + 0
24 + ],
25 + [
26 + 127,
27 + 0
28 + ],
29 + [
30 + 139,
31 + 0
32 + ],
33 + [
34 + 149,
35 + 0
36 + ],
37 + [
38 + 167,
39 + 0
40 + ],
41 + [
42 + 173,
43 + 0
44 + ],
45 + [
46 + 179,
47 + 0
48 + ],
49 + [
50 + 401,
51 + 0
52 + ],
53 + [
54 + 409,
55 + 0
56 + ],
57 + [
58 + 419,
59 + 0
60 + ],
61 + [
62 + 461,
63 + 0
64 + ],
65 + [
66 + 467,
67 + 0
68 + ],
69 + [
70 + 479,
71 + 0
72 + ],
73 + [
74 + 487,
75 + 0
76 + ],
77 + [
78 + 571,
79 + 0
80 + ],
81 + [
82 + 577,
83 + 0
84 + ],
85 + [
86 + 587,
87 + 0
88 + ],
89 + [
90 + 593,
91 + 0
92 + ],
93 + [
94 + 599,
95 + 0
96 + ],
97 + [
98 + 617,
99 + 0
100 + ],
101 + [
102 + 743,
103 + 0
104 + ],
105 + [
106 + 751,
107 + 0
108 + ],
109 + [
110 + 757,
111 + 0
112 + ],
113 + [
114 + 823,
115 + 0
116 + ],
117 + [
118 + 829,
119 + 0
120 + ],
121 + [
122 + 839,
123 + 0
124 + ],
125 + [
126 + 853,
127 + 0
128 + ],
129 + [
130 + 859,
131 + 0
132 + ],
133 + [
134 + 2113,
135 + 0
136 + ],
137 + [
138 + 2129,
139 + 0
140 + ],
141 + [
142 + 2141,
143 + 0
144 + ],
145 + [
146 + 2207,
147 + 0
148 + ],
149 + [
150 + 2213,
151 + 0
152 + ],
153 + [
154 + 2221,
155 + 0
156 + ],
157 + [
158 + 2237,
159 + 0
160 + ],
161 + [
162 + 2243,
163 + 0
164 + ],
165 + [
166 + 2251,
167 + 0
168 + ],
169 + [
170 + 2267,
171 + 0
172 + ],
173 + [
174 + 2273,
175 + 0
176 + ],
177 + [
178 + 2281,
179 + 0
180 + ],
181 + [
182 + 2287,
183 + 0
184 + ],
185 + [
186 + 2293,
187 + 0
188 + ],
189 + [
190 + 2311,
191 + 0
192 + ],
193 + [
194 + 2333,
195 + 0
196 + ],
197 + [
198 + 2339,
199 + 0
200 + ],
201 + [
202 + 2351,
203 + 0
204 + ],
205 + [
206 + 2357,
207 + 0
208 + ]
209 + ],
210 + "definition": "D(p) = N2 - N4 counting gaps whose END prime <= p"
211 + },
212 + "checkpoints": {
213 + "1000000": {
214 + "champion": 6,
215 + "N2": 8169,
216 + "N4": 8143,
217 + "D=N2-N4": 26,
218 + "G6": 66,
219 + "rho": -0.043143,
220 + "rho_times_lnx": -0.596
221 + },
222 + "10000000": {
223 + "champion": 6,
224 + "N2": 58980,
225 + "N4": 58621,
226 + "D=N2-N4": 359,
227 + "G6": 66,
228 + "rho": -0.038111,
229 + "rho_times_lnx": -0.6143
230 + },
231 + "100000000": {
232 + "champion": 6,
233 + "N2": 440312,
234 + "N4": 440257,
235 + "D=N2-N4": 55,
236 + "G6": 66,
237 + "rho": -0.031607,
238 + "rho_times_lnx": -0.5822
239 + },
240 + "1000000000": {
241 + "champion": 6,
242 + "N2": 3424506,
243 + "N4": 3424679,
244 + "D=N2-N4": -173,
245 + "G6": 216,
246 + "rho": -0.027532,
247 + "rho_times_lnx": -0.5705
248 + },
249 + "2000000000": {
250 + "champion": 6,
251 + "N2": 6388041,
252 + "N4": 6386967,
253 + "D=N2-N4": 1074,
254 + "G6": 216,
255 + "rho": -0.026453,
256 + "rho_times_lnx": -0.5665
257 + },
258 + "4000000000": {
259 + "champion": 6,
260 + "N2": 11944438,
261 + "N4": 11946708,
262 + "D=N2-N4": -2270,
263 + "G6": 216,
264 + "rho": -0.025561,
265 + "rho_times_lnx": -0.5651
266 + }
267 + },
268 + "maximal_gaps": [
269 + {
270 + "gap": 1,
271 + "after": 2,
272 + "merit": 1.4427,
273 + "csg": 2.08137
274 + },
275 + {
276 + "gap": 2,
277 + "after": 3,
278 + "merit": 1.8205,
279 + "csg": 1.65707
280 + },
281 + {
282 + "gap": 4,
283 + "after": 7,
284 + "merit": 2.0556,
285 + "csg": 1.05637
286 + },
287 + {
288 + "gap": 6,
289 + "after": 23,
290 + "merit": 1.9136,
291 + "csg": 0.61029
292 + },
293 + {
294 + "gap": 8,
295 + "after": 89,
296 + "merit": 1.7823,
297 + "csg": 0.39706
298 + },
299 + {
300 + "gap": 14,
301 + "after": 113,
302 + "merit": 2.9615,
303 + "csg": 0.62645
304 + },
305 + {
306 + "gap": 18,
307 + "after": 523,
308 + "merit": 2.8756,
309 + "csg": 0.45939
310 + },
311 + {
312 + "gap": 20,
313 + "after": 887,
314 + "merit": 2.9464,
315 + "csg": 0.43408
316 + },
317 + {
318 + "gap": 22,
319 + "after": 1129,
320 + "merit": 3.1299,
321 + "csg": 0.44527
322 + },
323 + {
324 + "gap": 34,
325 + "after": 1327,
326 + "merit": 4.7283,
327 + "csg": 0.65757
328 + },
329 + {
330 + "gap": 36,
331 + "after": 9551,
332 + "merit": 3.9282,
333 + "csg": 0.42864
334 + },
335 + {
336 + "gap": 44,
337 + "after": 15683,
338 + "merit": 4.5547,
339 + "csg": 0.47149
340 + },
341 + {
342 + "gap": 52,
343 + "after": 19609,
344 + "merit": 5.2612,
345 + "csg": 0.5323
346 + },
347 + {
348 + "gap": 72,
349 + "after": 31397,
350 + "merit": 6.9535,
351 + "csg": 0.67155
352 + },
353 + {
354 + "gap": 86,
355 + "after": 155921,
356 + "merit": 7.1924,
357 + "csg": 0.60151
358 + },
359 + {
360 + "gap": 96,
361 + "after": 360653,
362 + "merit": 7.5025,
363 + "csg": 0.58633
364 + },
365 + {
366 + "gap": 112,
367 + "after": 370261,
368 + "merit": 8.735,
369 + "csg": 0.68125
370 + },
371 + {
372 + "gap": 114,
373 + "after": 492113,
374 + "merit": 8.698,
375 + "csg": 0.66364
376 + },
377 + {
378 + "gap": 118,
379 + "after": 1349533,
380 + "merit": 8.3597,
381 + "csg": 0.59225
382 + },
383 + {
384 + "gap": 132,
385 + "after": 1357201,
386 + "merit": 9.3478,
387 + "csg": 0.66198
388 + },
389 + {
390 + "gap": 148,
391 + "after": 2010733,
392 + "merit": 10.197,
393 + "csg": 0.70257
394 + },
395 + {
396 + "gap": 154,
397 + "after": 4652353,
398 + "merit": 10.0307,
399 + "csg": 0.65334
400 + },
401 + {
402 + "gap": 180,
403 + "after": 17051707,
404 + "merit": 10.8097,
405 + "csg": 0.64916
406 + },
407 + {
408 + "gap": 210,
409 + "after": 20831323,
410 + "merit": 12.4615,
411 + "csg": 0.73947
412 + },
413 + {
414 + "gap": 220,
415 + "after": 47326693,
416 + "merit": 12.4487,
417 + "csg": 0.70441
418 + },
419 + {
420 + "gap": 222,
421 + "after": 122164747,
422 + "merit": 11.9221,
423 + "csg": 0.64025
424 + },
425 + {
426 + "gap": 234,
427 + "after": 189695659,
428 + "merit": 12.2764,
429 + "csg": 0.64406
430 + },
431 + {
432 + "gap": 248,
433 + "after": 191912783,
434 + "merit": 13.003,
435 + "csg": 0.68176
436 + },
437 + {
438 + "gap": 250,
439 + "after": 387096133,
440 + "merit": 12.6427,
441 + "csg": 0.63936
442 + },
443 + {
444 + "gap": 282,
445 + "after": 436273009,
446 + "merit": 14.1753,
447 + "csg": 0.71255
448 + },
449 + {
450 + "gap": 288,
451 + "after": 1294268491,
452 + "merit": 13.7266,
453 + "csg": 0.65423
454 + },
455 + {
456 + "gap": 292,
457 + "after": 1453168141,
458 + "merit": 13.8408,
459 + "csg": 0.65606
460 + },
461 + {
462 + "gap": 320,
463 + "after": 2300942549,
464 + "merit": 14.8447,
465 + "csg": 0.68864
466 + },
467 + {
468 + "gap": 336,
469 + "after": 3842610773,
470 + "merit": 15.2247,
471 + "csg": 0.68985
472 + }
473 + ],
474 + "max_csg": 0.73947
475 +}
\ No newline at end of file
added data/cycle2_c4_1e10_laptop.json +315 −0
@@ -0,0 +1,315 @@
1 +{
2 + "limit": 10000000000,
3 + "host": "laptop",
4 + "scan_seconds": 40.3,
5 + "checkpoints": {
6 + "1000000": {
7 + "rho1": -0.0431431,
8 + "rho1_lnx": -0.59604,
9 + "rho2": -0.011714,
10 + "rho2_lnx": -0.16184,
11 + "champion": 6,
12 + "G6": 66,
13 + "D_N2_minus_N4": 26,
14 + "n_gaps": 78497
15 + },
16 + "2000000": {
17 + "rho1": -0.0461202,
18 + "rho1_lnx": -0.66914,
19 + "rho2": -0.0146959,
20 + "rho2_lnx": -0.21322,
21 + "champion": 6,
22 + "G6": 66,
23 + "D_N2_minus_N4": 130,
24 + "n_gaps": 148932
25 + },
26 + "4000000": {
27 + "rho1": -0.0402461,
28 + "rho1_lnx": -0.61181,
29 + "rho2": -0.0160156,
30 + "rho2_lnx": -0.24347,
31 + "champion": 6,
32 + "G6": 66,
33 + "D_N2_minus_N4": 232,
34 + "n_gaps": 283145
35 + },
36 + "10000000": {
37 + "rho1": -0.0381114,
38 + "rho1_lnx": -0.61428,
39 + "rho2": -0.0135613,
40 + "rho2_lnx": -0.21858,
41 + "champion": 6,
42 + "G6": 66,
43 + "D_N2_minus_N4": 359,
44 + "n_gaps": 664578
45 + },
46 + "20000000": {
47 + "rho1": -0.0355623,
48 + "rho1_lnx": -0.59785,
49 + "rho2": -0.0147881,
50 + "rho2_lnx": -0.24861,
51 + "champion": 6,
52 + "G6": 66,
53 + "D_N2_minus_N4": 326,
54 + "n_gaps": 1270606
55 + },
56 + "40000000": {
57 + "rho1": -0.0332883,
58 + "rho1_lnx": -0.58269,
59 + "rho2": -0.0141482,
60 + "rho2_lnx": -0.24766,
61 + "champion": 6,
62 + "G6": 66,
63 + "D_N2_minus_N4": 520,
64 + "n_gaps": 2433653
65 + },
66 + "100000000": {
67 + "rho1": -0.0316066,
68 + "rho1_lnx": -0.58221,
69 + "rho2": -0.0127764,
70 + "rho2_lnx": -0.23535,
71 + "champion": 6,
72 + "G6": 66,
73 + "D_N2_minus_N4": 55,
74 + "n_gaps": 5761454
75 + },
76 + "200000000": {
77 + "rho1": -0.0300289,
78 + "rho1_lnx": -0.57397,
79 + "rho2": -0.0125457,
80 + "rho2_lnx": -0.2398,
81 + "champion": 6,
82 + "G6": 66,
83 + "D_N2_minus_N4": -343,
84 + "n_gaps": 11078936
85 + },
86 + "400000000": {
87 + "rho1": -0.0288227,
88 + "rho1_lnx": -0.57089,
89 + "rho2": -0.0123475,
90 + "rho2_lnx": -0.24457,
91 + "champion": 6,
92 + "G6": 66,
93 + "D_N2_minus_N4": -571,
94 + "n_gaps": 21336325
95 + },
96 + "1000000000": {
97 + "rho1": -0.0275318,
98 + "rho1_lnx": -0.57055,
99 + "rho2": -0.011703,
100 + "rho2_lnx": -0.24252,
101 + "champion": 6,
102 + "G6": 216,
103 + "D_N2_minus_N4": -173,
104 + "n_gaps": 50847533
105 + },
106 + "2000000000": {
107 + "rho1": -0.0264527,
108 + "rho1_lnx": -0.56652,
109 + "rho2": -0.0113653,
110 + "rho2_lnx": -0.2434,
111 + "champion": 6,
112 + "G6": 216,
113 + "D_N2_minus_N4": 1074,
114 + "n_gaps": 98222286
115 + },
116 + "4000000000": {
117 + "rho1": -0.0255606,
118 + "rho1_lnx": -0.56513,
119 + "rho2": -0.0110918,
120 + "rho2_lnx": -0.24524,
121 + "champion": 6,
122 + "G6": 216,
123 + "D_N2_minus_N4": -2270,
124 + "n_gaps": 189961811
125 + },
126 + "10000000000": {
127 + "rho1": -0.0243707,
128 + "rho1_lnx": -0.56116,
129 + "rho2": -0.0107629,
130 + "rho2_lnx": -0.24783,
131 + "champion": 6,
132 + "G6": 312,
133 + "D_N2_minus_N4": 2681,
134 + "n_gaps": 455052510
135 + }
136 + },
137 + "maximal_gaps": [
138 + {
139 + "gap": 1,
140 + "after": 2,
141 + "csg": 2.08137
142 + },
143 + {
144 + "gap": 2,
145 + "after": 3,
146 + "csg": 1.65707
147 + },
148 + {
149 + "gap": 4,
150 + "after": 7,
151 + "csg": 1.05637
152 + },
153 + {
154 + "gap": 6,
155 + "after": 23,
156 + "csg": 0.61029
157 + },
158 + {
159 + "gap": 8,
160 + "after": 89,
161 + "csg": 0.39706
162 + },
163 + {
164 + "gap": 14,
165 + "after": 113,
166 + "csg": 0.62645
167 + },
168 + {
169 + "gap": 18,
170 + "after": 523,
171 + "csg": 0.45939
172 + },
173 + {
174 + "gap": 20,
175 + "after": 887,
176 + "csg": 0.43408
177 + },
178 + {
179 + "gap": 22,
180 + "after": 1129,
181 + "csg": 0.44527
182 + },
183 + {
184 + "gap": 34,
185 + "after": 1327,
186 + "csg": 0.65757
187 + },
188 + {
189 + "gap": 36,
190 + "after": 9551,
191 + "csg": 0.42864
192 + },
193 + {
194 + "gap": 44,
195 + "after": 15683,
196 + "csg": 0.47149
197 + },
198 + {
199 + "gap": 52,
200 + "after": 19609,
201 + "csg": 0.5323
202 + },
203 + {
204 + "gap": 72,
205 + "after": 31397,
206 + "csg": 0.67155
207 + },
208 + {
209 + "gap": 86,
210 + "after": 155921,
211 + "csg": 0.60151
212 + },
213 + {
214 + "gap": 96,
215 + "after": 360653,
216 + "csg": 0.58633
217 + },
218 + {
219 + "gap": 112,
220 + "after": 370261,
221 + "csg": 0.68125
222 + },
223 + {
224 + "gap": 114,
225 + "after": 492113,
226 + "csg": 0.66364
227 + },
228 + {
229 + "gap": 118,
230 + "after": 1349533,
231 + "csg": 0.59225
232 + },
233 + {
234 + "gap": 132,
235 + "after": 1357201,
236 + "csg": 0.66198
237 + },
238 + {
239 + "gap": 148,
240 + "after": 2010733,
241 + "csg": 0.70257
242 + },
243 + {
244 + "gap": 154,
245 + "after": 4652353,
246 + "csg": 0.65334
247 + },
248 + {
249 + "gap": 180,
250 + "after": 17051707,
251 + "csg": 0.64916
252 + },
253 + {
254 + "gap": 210,
255 + "after": 20831323,
256 + "csg": 0.73947
257 + },
258 + {
259 + "gap": 220,
260 + "after": 47326693,
261 + "csg": 0.70441
262 + },
263 + {
264 + "gap": 222,
265 + "after": 122164747,
266 + "csg": 0.64025
267 + },
268 + {
269 + "gap": 234,
270 + "after": 189695659,
271 + "csg": 0.64406
272 + },
273 + {
274 + "gap": 248,
275 + "after": 191912783,
276 + "csg": 0.68176
277 + },
278 + {
279 + "gap": 250,
280 + "after": 387096133,
281 + "csg": 0.63936
282 + },
283 + {
284 + "gap": 282,
285 + "after": 436273009,
286 + "csg": 0.71255
287 + },
288 + {
289 + "gap": 288,
290 + "after": 1294268491,
291 + "csg": 0.65423
292 + },
293 + {
294 + "gap": 292,
295 + "after": 1453168141,
296 + "csg": 0.65606
297 + },
298 + {
299 + "gap": 320,
300 + "after": 2300942549,
301 + "csg": 0.68864
302 + },
303 + {
304 + "gap": 336,
305 + "after": 3842610773,
306 + "csg": 0.68985
307 + },
308 + {
309 + "gap": 354,
310 + "after": 4302407359,
311 + "csg": 0.71942
312 + }
313 + ],
314 + "method": "deterministic segmented sieve; streaming Pearson sums; gaps attributed to END prime; no randomness"
315 +}
\ No newline at end of file
added data/cycle2_c4_1e8_sanity.json +205 −0
@@ -0,0 +1,205 @@
1 +{
2 + "limit": 100000000,
3 + "host": "sanity",
4 + "scan_seconds": 0.3,
5 + "checkpoints": {
6 + "1000000": {
7 + "rho1": -0.0431431,
8 + "rho1_lnx": -0.59604,
9 + "rho2": -0.011714,
10 + "rho2_lnx": -0.16184,
11 + "champion": 6,
12 + "G6": 66,
13 + "D_N2_minus_N4": 26,
14 + "n_gaps": 78497
15 + },
16 + "2000000": {
17 + "rho1": -0.0461202,
18 + "rho1_lnx": -0.66914,
19 + "rho2": -0.0146959,
20 + "rho2_lnx": -0.21322,
21 + "champion": 6,
22 + "G6": 66,
23 + "D_N2_minus_N4": 130,
24 + "n_gaps": 148932
25 + },
26 + "4000000": {
27 + "rho1": -0.0402461,
28 + "rho1_lnx": -0.61181,
29 + "rho2": -0.0160156,
30 + "rho2_lnx": -0.24347,
31 + "champion": 6,
32 + "G6": 66,
33 + "D_N2_minus_N4": 232,
34 + "n_gaps": 283145
35 + },
36 + "10000000": {
37 + "rho1": -0.0381114,
38 + "rho1_lnx": -0.61428,
39 + "rho2": -0.0135613,
40 + "rho2_lnx": -0.21858,
41 + "champion": 6,
42 + "G6": 66,
43 + "D_N2_minus_N4": 359,
44 + "n_gaps": 664578
45 + },
46 + "20000000": {
47 + "rho1": -0.0355623,
48 + "rho1_lnx": -0.59785,
49 + "rho2": -0.0147881,
50 + "rho2_lnx": -0.24861,
51 + "champion": 6,
52 + "G6": 66,
53 + "D_N2_minus_N4": 326,
54 + "n_gaps": 1270606
55 + },
56 + "40000000": {
57 + "rho1": -0.0332883,
58 + "rho1_lnx": -0.58269,
59 + "rho2": -0.0141482,
60 + "rho2_lnx": -0.24766,
61 + "champion": 6,
62 + "G6": 66,
63 + "D_N2_minus_N4": 520,
64 + "n_gaps": 2433653
65 + },
66 + "100000000": {
67 + "rho1": -0.0316066,
68 + "rho1_lnx": -0.58221,
69 + "rho2": -0.0127764,
70 + "rho2_lnx": -0.23535,
71 + "champion": 6,
72 + "G6": 66,
73 + "D_N2_minus_N4": 55,
74 + "n_gaps": 5761454
75 + }
76 + },
77 + "maximal_gaps": [
78 + {
79 + "gap": 1,
80 + "after": 2,
81 + "csg": 2.08137
82 + },
83 + {
84 + "gap": 2,
85 + "after": 3,
86 + "csg": 1.65707
87 + },
88 + {
89 + "gap": 4,
90 + "after": 7,
91 + "csg": 1.05637
92 + },
93 + {
94 + "gap": 6,
95 + "after": 23,
96 + "csg": 0.61029
97 + },
98 + {
99 + "gap": 8,
100 + "after": 89,
101 + "csg": 0.39706
102 + },
103 + {
104 + "gap": 14,
105 + "after": 113,
106 + "csg": 0.62645
107 + },
108 + {
109 + "gap": 18,
110 + "after": 523,
111 + "csg": 0.45939
112 + },
113 + {
114 + "gap": 20,
115 + "after": 887,
116 + "csg": 0.43408
117 + },
118 + {
119 + "gap": 22,
120 + "after": 1129,
121 + "csg": 0.44527
122 + },
123 + {
124 + "gap": 34,
125 + "after": 1327,
126 + "csg": 0.65757
127 + },
128 + {
129 + "gap": 36,
130 + "after": 9551,
131 + "csg": 0.42864
132 + },
133 + {
134 + "gap": 44,
135 + "after": 15683,
136 + "csg": 0.47149
137 + },
138 + {
139 + "gap": 52,
140 + "after": 19609,
141 + "csg": 0.5323
142 + },
143 + {
144 + "gap": 72,
145 + "after": 31397,
146 + "csg": 0.67155
147 + },
148 + {
149 + "gap": 86,
150 + "after": 155921,
151 + "csg": 0.60151
152 + },
153 + {
154 + "gap": 96,
155 + "after": 360653,
156 + "csg": 0.58633
157 + },
158 + {
159 + "gap": 112,
160 + "after": 370261,
161 + "csg": 0.68125
162 + },
163 + {
164 + "gap": 114,
165 + "after": 492113,
166 + "csg": 0.66364
167 + },
168 + {
169 + "gap": 118,
170 + "after": 1349533,
171 + "csg": 0.59225
172 + },
173 + {
174 + "gap": 132,
175 + "after": 1357201,
176 + "csg": 0.66198
177 + },
178 + {
179 + "gap": 148,
180 + "after": 2010733,
181 + "csg": 0.70257
182 + },
183 + {
184 + "gap": 154,
185 + "after": 4652353,
186 + "csg": 0.65334
187 + },
188 + {
189 + "gap": 180,
190 + "after": 17051707,
191 + "csg": 0.64916
192 + },
193 + {
194 + "gap": 210,
195 + "after": 20831323,
196 + "csg": 0.73947
197 + },
198 + {
199 + "gap": 220,
200 + "after": 47326693,
201 + "csg": 0.70441
202 + }
203 + ],
204 + "method": "deterministic segmented sieve; streaming Pearson sums; gaps attributed to END prime; no randomness"
205 +}
\ No newline at end of file
added data/cycle2_c4_4e10_M3U96a.json +355 −0
@@ -0,0 +1,355 @@
1 +{
2 + "limit": 40000000000,
3 + "host": "M3U96a",
4 + "scan_seconds": 127.6,
5 + "checkpoints": {
6 + "1000000": {
7 + "rho1": -0.0431431,
8 + "rho1_lnx": -0.59604,
9 + "rho2": -0.011714,
10 + "rho2_lnx": -0.16184,
11 + "champion": 6,
12 + "G6": 66,
13 + "D_N2_minus_N4": 26,
14 + "n_gaps": 78497
15 + },
16 + "2000000": {
17 + "rho1": -0.0461202,
18 + "rho1_lnx": -0.66914,
19 + "rho2": -0.0146959,
20 + "rho2_lnx": -0.21322,
21 + "champion": 6,
22 + "G6": 66,
23 + "D_N2_minus_N4": 130,
24 + "n_gaps": 148932
25 + },
26 + "4000000": {
27 + "rho1": -0.0402461,
28 + "rho1_lnx": -0.61181,
29 + "rho2": -0.0160156,
30 + "rho2_lnx": -0.24347,
31 + "champion": 6,
32 + "G6": 66,
33 + "D_N2_minus_N4": 232,
34 + "n_gaps": 283145
35 + },
36 + "10000000": {
37 + "rho1": -0.0381114,
38 + "rho1_lnx": -0.61428,
39 + "rho2": -0.0135613,
40 + "rho2_lnx": -0.21858,
41 + "champion": 6,
42 + "G6": 66,
43 + "D_N2_minus_N4": 359,
44 + "n_gaps": 664578
45 + },
46 + "20000000": {
47 + "rho1": -0.0355623,
48 + "rho1_lnx": -0.59785,
49 + "rho2": -0.0147881,
50 + "rho2_lnx": -0.24861,
51 + "champion": 6,
52 + "G6": 66,
53 + "D_N2_minus_N4": 326,
54 + "n_gaps": 1270606
55 + },
56 + "40000000": {
57 + "rho1": -0.0332883,
58 + "rho1_lnx": -0.58269,
59 + "rho2": -0.0141482,
60 + "rho2_lnx": -0.24766,
61 + "champion": 6,
62 + "G6": 66,
63 + "D_N2_minus_N4": 520,
64 + "n_gaps": 2433653
65 + },
66 + "100000000": {
67 + "rho1": -0.0316066,
68 + "rho1_lnx": -0.58221,
69 + "rho2": -0.0127764,
70 + "rho2_lnx": -0.23535,
71 + "champion": 6,
72 + "G6": 66,
73 + "D_N2_minus_N4": 55,
74 + "n_gaps": 5761454
75 + },
76 + "200000000": {
77 + "rho1": -0.0300289,
78 + "rho1_lnx": -0.57397,
79 + "rho2": -0.0125457,
80 + "rho2_lnx": -0.2398,
81 + "champion": 6,
82 + "G6": 66,
83 + "D_N2_minus_N4": -343,
84 + "n_gaps": 11078936
85 + },
86 + "400000000": {
87 + "rho1": -0.0288227,
88 + "rho1_lnx": -0.57089,
89 + "rho2": -0.0123475,
90 + "rho2_lnx": -0.24457,
91 + "champion": 6,
92 + "G6": 66,
93 + "D_N2_minus_N4": -571,
94 + "n_gaps": 21336325
95 + },
96 + "1000000000": {
97 + "rho1": -0.0275318,
98 + "rho1_lnx": -0.57055,
99 + "rho2": -0.011703,
100 + "rho2_lnx": -0.24252,
101 + "champion": 6,
102 + "G6": 216,
103 + "D_N2_minus_N4": -173,
104 + "n_gaps": 50847533
105 + },
106 + "2000000000": {
107 + "rho1": -0.0264527,
108 + "rho1_lnx": -0.56652,
109 + "rho2": -0.0113653,
110 + "rho2_lnx": -0.2434,
111 + "champion": 6,
112 + "G6": 216,
113 + "D_N2_minus_N4": 1074,
114 + "n_gaps": 98222286
115 + },
116 + "4000000000": {
117 + "rho1": -0.0255606,
118 + "rho1_lnx": -0.56513,
119 + "rho2": -0.0110918,
120 + "rho2_lnx": -0.24524,
121 + "champion": 6,
122 + "G6": 216,
123 + "D_N2_minus_N4": -2270,
124 + "n_gaps": 189961811
125 + },
126 + "10000000000": {
127 + "rho1": -0.0243707,
128 + "rho1_lnx": -0.56116,
129 + "rho2": -0.0107629,
130 + "rho2_lnx": -0.24783,
131 + "champion": 6,
132 + "G6": 312,
133 + "D_N2_minus_N4": 2681,
134 + "n_gaps": 455052510
135 + },
136 + "20000000000": {
137 + "rho1": -0.0235904,
138 + "rho1_lnx": -0.55954,
139 + "rho2": -0.0104617,
140 + "rho2_lnx": -0.24814,
141 + "champion": 6,
142 + "G6": 282,
143 + "D_N2_minus_N4": 7160,
144 + "n_gaps": 882206715
145 + },
146 + "40000000000": {
147 + "rho1": -0.022763,
148 + "rho1_lnx": -0.55569,
149 + "rho2": -0.0102117,
150 + "rho2_lnx": -0.24929,
151 + "champion": 6,
152 + "G6": 276,
153 + "D_N2_minus_N4": -804,
154 + "n_gaps": 1711955432
155 + }
156 + },
157 + "maximal_gaps": [
158 + {
159 + "gap": 1,
160 + "after": 2,
161 + "csg": 2.08137
162 + },
163 + {
164 + "gap": 2,
165 + "after": 3,
166 + "csg": 1.65707
167 + },
168 + {
169 + "gap": 4,
170 + "after": 7,
171 + "csg": 1.05637
172 + },
173 + {
174 + "gap": 6,
175 + "after": 23,
176 + "csg": 0.61029
177 + },
178 + {
179 + "gap": 8,
180 + "after": 89,
181 + "csg": 0.39706
182 + },
183 + {
184 + "gap": 14,
185 + "after": 113,
186 + "csg": 0.62645
187 + },
188 + {
189 + "gap": 18,
190 + "after": 523,
191 + "csg": 0.45939
192 + },
193 + {
194 + "gap": 20,
195 + "after": 887,
196 + "csg": 0.43408
197 + },
198 + {
199 + "gap": 22,
200 + "after": 1129,
201 + "csg": 0.44527
202 + },
203 + {
204 + "gap": 34,
205 + "after": 1327,
206 + "csg": 0.65757
207 + },
208 + {
209 + "gap": 36,
210 + "after": 9551,
211 + "csg": 0.42864
212 + },
213 + {
214 + "gap": 44,
215 + "after": 15683,
216 + "csg": 0.47149
217 + },
218 + {
219 + "gap": 52,
220 + "after": 19609,
221 + "csg": 0.5323
222 + },
223 + {
224 + "gap": 72,
225 + "after": 31397,
226 + "csg": 0.67155
227 + },
228 + {
229 + "gap": 86,
230 + "after": 155921,
231 + "csg": 0.60151
232 + },
233 + {
234 + "gap": 96,
235 + "after": 360653,
236 + "csg": 0.58633
237 + },
238 + {
239 + "gap": 112,
240 + "after": 370261,
241 + "csg": 0.68125
242 + },
243 + {
244 + "gap": 114,
245 + "after": 492113,
246 + "csg": 0.66364
247 + },
248 + {
249 + "gap": 118,
250 + "after": 1349533,
251 + "csg": 0.59225
252 + },
253 + {
254 + "gap": 132,
255 + "after": 1357201,
256 + "csg": 0.66198
257 + },
258 + {
259 + "gap": 148,
260 + "after": 2010733,
261 + "csg": 0.70257
262 + },
263 + {
264 + "gap": 154,
265 + "after": 4652353,
266 + "csg": 0.65334
267 + },
268 + {
269 + "gap": 180,
270 + "after": 17051707,
271 + "csg": 0.64916
272 + },
273 + {
274 + "gap": 210,
275 + "after": 20831323,
276 + "csg": 0.73947
277 + },
278 + {
279 + "gap": 220,
280 + "after": 47326693,
281 + "csg": 0.70441
282 + },
283 + {
284 + "gap": 222,
285 + "after": 122164747,
286 + "csg": 0.64025
287 + },
288 + {
289 + "gap": 234,
290 + "after": 189695659,
291 + "csg": 0.64406
292 + },
293 + {
294 + "gap": 248,
295 + "after": 191912783,
296 + "csg": 0.68176
297 + },
298 + {
299 + "gap": 250,
300 + "after": 387096133,
301 + "csg": 0.63936
302 + },
303 + {
304 + "gap": 282,
305 + "after": 436273009,
306 + "csg": 0.71255
307 + },
308 + {
309 + "gap": 288,
310 + "after": 1294268491,
311 + "csg": 0.65423
312 + },
313 + {
314 + "gap": 292,
315 + "after": 1453168141,
316 + "csg": 0.65606
317 + },
318 + {
319 + "gap": 320,
320 + "after": 2300942549,
321 + "csg": 0.68864
322 + },
323 + {
324 + "gap": 336,
325 + "after": 3842610773,
326 + "csg": 0.68985
327 + },
328 + {
329 + "gap": 354,
330 + "after": 4302407359,
331 + "csg": 0.71942
332 + },
333 + {
334 + "gap": 382,
335 + "after": 10726904659,
336 + "csg": 0.71613
337 + },
338 + {
339 + "gap": 384,
340 + "after": 20678048297,
341 + "csg": 0.68064
342 + },
343 + {
344 + "gap": 394,
345 + "after": 22367084959,
346 + "csg": 0.69377
347 + },
348 + {
349 + "gap": 456,
350 + "after": 25056082087,
351 + "csg": 0.79535
352 + }
353 + ],
354 + "method": "deterministic segmented sieve; streaming Pearson sums; gaps attributed to END prime; no randomness"
355 +}
\ No newline at end of file
added data/cycle3_c4_1e12.json +350 −0
@@ -0,0 +1,350 @@
1 +{
2 + "limit": 1000000000000,
3 + "n_chunks": 1000,
4 + "hosts": [
5 + "MacBooknpierre2",
6 + "MacBookonpierre",
7 + "MacStudnpierre2",
8 + "MacStudnpierre3",
9 + "MacStudnpierre4",
10 + "MacStudnpierre5",
11 + "MacStudnpierre6",
12 + "MacStudnpierre8",
13 + "Minidesonpierre",
14 + "simonpiacStudio",
15 + "simonpierresMBP",
16 + "simonpirresMBP2"
17 + ],
18 + "total_core_seconds": 17718.7,
19 + "checkpoints": {
20 + "10000000000": {
21 + "rho1": -0.0243707,
22 + "rho1_lnx": -0.56116,
23 + "rho2": -0.0107629,
24 + "rho2_lnx": -0.24783,
25 + "champion": 6,
26 + "G6": 312,
27 + "D_N2_minus_N4": 2681
28 + },
29 + "20000000000": {
30 + "rho1": -0.0235904,
31 + "rho1_lnx": -0.55954,
32 + "rho2": -0.0104617,
33 + "rho2_lnx": -0.24814,
34 + "champion": 6,
35 + "G6": 282,
36 + "D_N2_minus_N4": 7160
37 + },
38 + "40000000000": {
39 + "rho1": -0.022763,
40 + "rho1_lnx": -0.55569,
41 + "rho2": -0.0102117,
42 + "rho2_lnx": -0.24929,
43 + "champion": 6,
44 + "G6": 276,
45 + "D_N2_minus_N4": -804
46 + },
47 + "100000000000": {
48 + "rho1": -0.021865,
49 + "rho1_lnx": -0.55381,
50 + "rho2": -0.0098503,
51 + "rho2_lnx": -0.24949,
52 + "champion": 6,
53 + "G6": 366,
54 + "D_N2_minus_N4": 2888
55 + },
56 + "200000000000": {
57 + "rho1": -0.0212141,
58 + "rho1_lnx": -0.55203,
59 + "rho2": -0.0096209,
60 + "rho2_lnx": -0.25035,
61 + "champion": 6,
62 + "G6": 402,
63 + "D_N2_minus_N4": 7550
64 + },
65 + "400000000000": {
66 + "rho1": -0.020598,
67 + "rho1_lnx": -0.55027,
68 + "rho2": -0.0093865,
69 + "rho2_lnx": -0.25076,
70 + "champion": 6,
71 + "G6": 450,
72 + "D_N2_minus_N4": 22767
73 + },
74 + "1000000000000": {
75 + "rho1": -0.0198167,
76 + "rho1_lnx": -0.54756,
77 + "rho2": -0.0090989,
78 + "rho2_lnx": -0.25141,
79 + "champion": 6,
80 + "G6": 480,
81 + "D_N2_minus_N4": -238
82 + }
83 + },
84 + "maximal_gaps": [
85 + {
86 + "gap": 1,
87 + "after": 2,
88 + "csg": 2.08137
89 + },
90 + {
91 + "gap": 2,
92 + "after": 3,
93 + "csg": 1.65707
94 + },
95 + {
96 + "gap": 4,
97 + "after": 7,
98 + "csg": 1.05637
99 + },
100 + {
101 + "gap": 6,
102 + "after": 23,
103 + "csg": 0.61029
104 + },
105 + {
106 + "gap": 8,
107 + "after": 89,
108 + "csg": 0.39706
109 + },
110 + {
111 + "gap": 14,
112 + "after": 113,
113 + "csg": 0.62645
114 + },
115 + {
116 + "gap": 18,
117 + "after": 523,
118 + "csg": 0.45939
119 + },
120 + {
121 + "gap": 20,
122 + "after": 887,
123 + "csg": 0.43408
124 + },
125 + {
126 + "gap": 22,
127 + "after": 1129,
128 + "csg": 0.44527
129 + },
130 + {
131 + "gap": 34,
132 + "after": 1327,
133 + "csg": 0.65757
134 + },
135 + {
136 + "gap": 36,
137 + "after": 9551,
138 + "csg": 0.42864
139 + },
140 + {
141 + "gap": 44,
142 + "after": 15683,
143 + "csg": 0.47149
144 + },
145 + {
146 + "gap": 52,
147 + "after": 19609,
148 + "csg": 0.5323
149 + },
150 + {
151 + "gap": 72,
152 + "after": 31397,
153 + "csg": 0.67155
154 + },
155 + {
156 + "gap": 86,
157 + "after": 155921,
158 + "csg": 0.60151
159 + },
160 + {
161 + "gap": 96,
162 + "after": 360653,
163 + "csg": 0.58633
164 + },
165 + {
166 + "gap": 112,
167 + "after": 370261,
168 + "csg": 0.68125
169 + },
170 + {
171 + "gap": 114,
172 + "after": 492113,
173 + "csg": 0.66364
174 + },
175 + {
176 + "gap": 118,
177 + "after": 1349533,
178 + "csg": 0.59225
179 + },
180 + {
181 + "gap": 132,
182 + "after": 1357201,
183 + "csg": 0.66198
184 + },
185 + {
186 + "gap": 148,
187 + "after": 2010733,
188 + "csg": 0.70257
189 + },
190 + {
191 + "gap": 154,
192 + "after": 4652353,
193 + "csg": 0.65334
194 + },
195 + {
196 + "gap": 180,
197 + "after": 17051707,
198 + "csg": 0.64916
199 + },
200 + {
201 + "gap": 210,
202 + "after": 20831323,
203 + "csg": 0.73947
204 + },
205 + {
206 + "gap": 220,
207 + "after": 47326693,
208 + "csg": 0.70441
209 + },
210 + {
211 + "gap": 222,
212 + "after": 122164747,
213 + "csg": 0.64025
214 + },
215 + {
216 + "gap": 234,
217 + "after": 189695659,
218 + "csg": 0.64406
219 + },
220 + {
221 + "gap": 248,
222 + "after": 191912783,
223 + "csg": 0.68176
224 + },
225 + {
226 + "gap": 250,
227 + "after": 387096133,
228 + "csg": 0.63936
229 + },
230 + {
231 + "gap": 282,
232 + "after": 436273009,
233 + "csg": 0.71255
234 + },
235 + {
236 + "gap": 288,
237 + "after": 1294268491,
238 + "csg": 0.65423
239 + },
240 + {
241 + "gap": 292,
242 + "after": 1453168141,
243 + "csg": 0.65606
244 + },
245 + {
246 + "gap": 320,
247 + "after": 2300942549,
248 + "csg": 0.68864
249 + },
250 + {
251 + "gap": 336,
252 + "after": 3842610773,
253 + "csg": 0.68985
254 + },
255 + {
256 + "gap": 354,
257 + "after": 4302407359,
258 + "csg": 0.71942
259 + },
260 + {
261 + "gap": 382,
262 + "after": 10726904659,
263 + "csg": 0.71613
264 + },
265 + {
266 + "gap": 384,
267 + "after": 20678048297,
268 + "csg": 0.68064
269 + },
270 + {
271 + "gap": 394,
272 + "after": 22367084959,
273 + "csg": 0.69377
274 + },
275 + {
276 + "gap": 456,
277 + "after": 25056082087,
278 + "csg": 0.79535
279 + },
280 + {
281 + "gap": 464,
282 + "after": 42652618343,
283 + "csg": 0.77451
284 + },
285 + {
286 + "gap": 468,
287 + "after": 127976334671,
288 + "csg": 0.7155
289 + },
290 + {
291 + "gap": 474,
292 + "after": 182226896239,
293 + "csg": 0.70505
294 + },
295 + {
296 + "gap": 486,
297 + "after": 241160624143,
298 + "csg": 0.70753
299 + },
300 + {
301 + "gap": 490,
302 + "after": 297501075799,
303 + "csg": 0.70206
304 + },
305 + {
306 + "gap": 500,
307 + "after": 303371455241,
308 + "csg": 0.71533
309 + },
310 + {
311 + "gap": 514,
312 + "after": 304599508537,
313 + "csg": 0.73513
314 + },
315 + {
316 + "gap": 516,
317 + "after": 416608695821,
318 + "csg": 0.72082
319 + },
320 + {
321 + "gap": 532,
322 + "after": 461690510011,
323 + "csg": 0.7375
324 + },
325 + {
326 + "gap": 534,
327 + "after": 614487453523,
328 + "csg": 0.72476
329 + },
330 + {
331 + "gap": 540,
332 + "after": 738832927927,
333 + "csg": 0.72305
334 + }
335 + ],
336 + "cross_validation_vs_cycle2": {
337 + "10000000000": {
338 + "rho1_lnx_diff": 0.0,
339 + "exact_D_match": true
340 + },
341 + "20000000000": {
342 + "rho1_lnx_diff": 0.0,
343 + "exact_D_match": true
344 + },
345 + "40000000000": {
346 + "rho1_lnx_diff": 0.0,
347 + "exact_D_match": true
348 + }
349 + }
350 +}
\ No newline at end of file
added data/cycle4_c4_1e13.json +419 −0
@@ -0,0 +1,419 @@
1 +{
2 + "limit": 10000000000000,
3 + "n_chunks": 10000,
4 + "hosts": [
5 + "MacStudnpierre2",
6 + "MacStudnpierre2/gpu",
7 + "MacStudnpierre4",
8 + "MacStudnpierre4/gpu",
9 + "MacStudnpierre6",
10 + "MacStudnpierre6/gpu"
11 + ],
12 + "total_core_seconds": 147737.5,
13 + "checkpoints": {
14 + "10000000000": {
15 + "rho1": -0.0243707,
16 + "rho1_lnx": -0.56116,
17 + "rho2": -0.0107629,
18 + "rho2_lnx": -0.24783,
19 + "champion": 6,
20 + "G6": 312,
21 + "D_N2_minus_N4": 2681
22 + },
23 + "20000000000": {
24 + "rho1": -0.0235904,
25 + "rho1_lnx": -0.55954,
26 + "rho2": -0.0104617,
27 + "rho2_lnx": -0.24814,
28 + "champion": 6,
29 + "G6": 282,
30 + "D_N2_minus_N4": 7160
31 + },
32 + "40000000000": {
33 + "rho1": -0.022763,
34 + "rho1_lnx": -0.55569,
35 + "rho2": -0.0102117,
36 + "rho2_lnx": -0.24929,
37 + "champion": 6,
38 + "G6": 276,
39 + "D_N2_minus_N4": -804
40 + },
41 + "100000000000": {
42 + "rho1": -0.021865,
43 + "rho1_lnx": -0.55381,
44 + "rho2": -0.0098503,
45 + "rho2_lnx": -0.24949,
46 + "champion": 6,
47 + "G6": 366,
48 + "D_N2_minus_N4": 2888
49 + },
50 + "200000000000": {
51 + "rho1": -0.0212141,
52 + "rho1_lnx": -0.55203,
53 + "rho2": -0.0096209,
54 + "rho2_lnx": -0.25035,
55 + "champion": 6,
56 + "G6": 402,
57 + "D_N2_minus_N4": 7550
58 + },
59 + "400000000000": {
60 + "rho1": -0.020598,
61 + "rho1_lnx": -0.55027,
62 + "rho2": -0.0093865,
63 + "rho2_lnx": -0.25076,
64 + "champion": 6,
65 + "G6": 450,
66 + "D_N2_minus_N4": 22767
67 + },
68 + "1000000000000": {
69 + "rho1": -0.0198167,
70 + "rho1_lnx": -0.54756,
71 + "rho2": -0.0090989,
72 + "rho2_lnx": -0.25141,
73 + "champion": 6,
74 + "G6": 480,
75 + "D_N2_minus_N4": -238
76 + },
77 + "2000000000000": {
78 + "rho1": -0.019277,
79 + "rho1_lnx": -0.54601,
80 + "rho2": -0.0088835,
81 + "rho2_lnx": -0.25162,
82 + "champion": 6,
83 + "G6": 486,
84 + "D_N2_minus_N4": -98967
85 + },
86 + "4000000000000": {
87 + "rho1": -0.0187639,
88 + "rho1_lnx": -0.54448,
89 + "rho2": -0.0086864,
90 + "rho2_lnx": -0.25205,
91 + "champion": 6,
92 + "G6": 486,
93 + "D_N2_minus_N4": -37509
94 + },
95 + "10000000000000": {
96 + "rho1": -0.0181281,
97 + "rho1_lnx": -0.54264,
98 + "rho2": -0.0084342,
99 + "rho2_lnx": -0.25247,
100 + "champion": 6,
101 + "G6": 576,
102 + "D_N2_minus_N4": 8870
103 + }
104 + },
105 + "maximal_gaps": [
106 + {
107 + "gap": 1,
108 + "after": 2,
109 + "csg": 2.08137
110 + },
111 + {
112 + "gap": 2,
113 + "after": 3,
114 + "csg": 1.65707
115 + },
116 + {
117 + "gap": 4,
118 + "after": 7,
119 + "csg": 1.05637
120 + },
121 + {
122 + "gap": 6,
123 + "after": 23,
124 + "csg": 0.61029
125 + },
126 + {
127 + "gap": 8,
128 + "after": 89,
129 + "csg": 0.39706
130 + },
131 + {
132 + "gap": 14,
133 + "after": 113,
134 + "csg": 0.62645
135 + },
136 + {
137 + "gap": 18,
138 + "after": 523,
139 + "csg": 0.45939
140 + },
141 + {
142 + "gap": 20,
143 + "after": 887,
144 + "csg": 0.43408
145 + },
146 + {
147 + "gap": 22,
148 + "after": 1129,
149 + "csg": 0.44527
150 + },
151 + {
152 + "gap": 34,
153 + "after": 1327,
154 + "csg": 0.65757
155 + },
156 + {
157 + "gap": 36,
158 + "after": 9551,
159 + "csg": 0.42864
160 + },
161 + {
162 + "gap": 44,
163 + "after": 15683,
164 + "csg": 0.47149
165 + },
166 + {
167 + "gap": 52,
168 + "after": 19609,
169 + "csg": 0.5323
170 + },
171 + {
172 + "gap": 72,
173 + "after": 31397,
174 + "csg": 0.67155
175 + },
176 + {
177 + "gap": 86,
178 + "after": 155921,
179 + "csg": 0.60151
180 + },
181 + {
182 + "gap": 96,
183 + "after": 360653,
184 + "csg": 0.58633
185 + },
186 + {
187 + "gap": 112,
188 + "after": 370261,
189 + "csg": 0.68125
190 + },
191 + {
192 + "gap": 114,
193 + "after": 492113,
194 + "csg": 0.66364
195 + },
196 + {
197 + "gap": 118,
198 + "after": 1349533,
199 + "csg": 0.59225
200 + },
201 + {
202 + "gap": 132,
203 + "after": 1357201,
204 + "csg": 0.66198
205 + },
206 + {
207 + "gap": 148,
208 + "after": 2010733,
209 + "csg": 0.70257
210 + },
211 + {
212 + "gap": 154,
213 + "after": 4652353,
214 + "csg": 0.65334
215 + },
216 + {
217 + "gap": 180,
218 + "after": 17051707,
219 + "csg": 0.64916
220 + },
221 + {
222 + "gap": 210,
223 + "after": 20831323,
224 + "csg": 0.73947
225 + },
226 + {
227 + "gap": 220,
228 + "after": 47326693,
229 + "csg": 0.70441
230 + },
231 + {
232 + "gap": 222,
233 + "after": 122164747,
234 + "csg": 0.64025
235 + },
236 + {
237 + "gap": 234,
238 + "after": 189695659,
239 + "csg": 0.64406
240 + },
241 + {
242 + "gap": 248,
243 + "after": 191912783,
244 + "csg": 0.68176
245 + },
246 + {
247 + "gap": 250,
248 + "after": 387096133,
249 + "csg": 0.63936
250 + },
251 + {
252 + "gap": 282,
253 + "after": 436273009,
254 + "csg": 0.71255
255 + },
256 + {
257 + "gap": 288,
258 + "after": 1294268491,
259 + "csg": 0.65423
260 + },
261 + {
262 + "gap": 292,
263 + "after": 1453168141,
264 + "csg": 0.65606
265 + },
266 + {
267 + "gap": 320,
268 + "after": 2300942549,
269 + "csg": 0.68864
270 + },
271 + {
272 + "gap": 336,
273 + "after": 3842610773,
274 + "csg": 0.68985
275 + },
276 + {
277 + "gap": 354,
278 + "after": 4302407359,
279 + "csg": 0.71942
280 + },
281 + {
282 + "gap": 382,
283 + "after": 10726904659,
284 + "csg": 0.71613
285 + },
286 + {
287 + "gap": 384,
288 + "after": 20678048297,
289 + "csg": 0.68064
290 + },
291 + {
292 + "gap": 394,
293 + "after": 22367084959,
294 + "csg": 0.69377
295 + },
296 + {
297 + "gap": 456,
298 + "after": 25056082087,
299 + "csg": 0.79535
300 + },
301 + {
302 + "gap": 464,
303 + "after": 42652618343,
304 + "csg": 0.77451
305 + },
306 + {
307 + "gap": 468,
308 + "after": 127976334671,
309 + "csg": 0.7155
310 + },
311 + {
312 + "gap": 474,
313 + "after": 182226896239,
314 + "csg": 0.70505
315 + },
316 + {
317 + "gap": 486,
318 + "after": 241160624143,
319 + "csg": 0.70753
320 + },
321 + {
322 + "gap": 490,
323 + "after": 297501075799,
324 + "csg": 0.70206
325 + },
326 + {
327 + "gap": 500,
328 + "after": 303371455241,
329 + "csg": 0.71533
330 + },
331 + {
332 + "gap": 514,
333 + "after": 304599508537,
334 + "csg": 0.73513
335 + },
336 + {
337 + "gap": 516,
338 + "after": 416608695821,
339 + "csg": 0.72082
340 + },
341 + {
342 + "gap": 532,
343 + "after": 461690510011,
344 + "csg": 0.7375
345 + },
346 + {
347 + "gap": 534,
348 + "after": 614487453523,
349 + "csg": 0.72476
350 + },
351 + {
352 + "gap": 540,
353 + "after": 738832927927,
354 + "csg": 0.72305
355 + },
356 + {
357 + "gap": 582,
358 + "after": 1346294310749,
359 + "csg": 0.74616
360 + },
361 + {
362 + "gap": 588,
363 + "after": 1408695493609,
364 + "csg": 0.75141
365 + },
366 + {
367 + "gap": 602,
368 + "after": 1968188556461,
369 + "csg": 0.75123
370 + },
371 + {
372 + "gap": 652,
373 + "after": 2614941710599,
374 + "csg": 0.79754
375 + },
376 + {
377 + "gap": 674,
378 + "after": 7177162611713,
379 + "csg": 0.76917
380 + }
381 + ],
382 + "cross_validation_vs_cycle2": {
383 + "10000000000": {
384 + "ref": "cycle3_c4_1e12.json",
385 + "rho1_lnx_diff": 0.0,
386 + "exact_D_match": true
387 + },
388 + "20000000000": {
389 + "ref": "cycle3_c4_1e12.json",
390 + "rho1_lnx_diff": 0.0,
391 + "exact_D_match": true
392 + },
393 + "40000000000": {
394 + "ref": "cycle3_c4_1e12.json",
395 + "rho1_lnx_diff": 0.0,
396 + "exact_D_match": true
397 + },
398 + "100000000000": {
399 + "ref": "cycle3_c4_1e12.json",
400 + "rho1_lnx_diff": 0.0,
401 + "exact_D_match": true
402 + },
403 + "200000000000": {
404 + "ref": "cycle3_c4_1e12.json",
405 + "rho1_lnx_diff": 0.0,
406 + "exact_D_match": true
407 + },
408 + "400000000000": {
409 + "ref": "cycle3_c4_1e12.json",
410 + "rho1_lnx_diff": 0.0,
411 + "exact_D_match": true
412 + },
413 + "1000000000000": {
414 + "ref": "cycle3_c4_1e12.json",
415 + "rho1_lnx_diff": 0.0,
416 + "exact_D_match": true
417 + }
418 + }
419 +}
\ No newline at end of file
added data/fit_c4.json +11 −0
@@ -0,0 +1,11 @@
1 +{
2 + "source": "cycle2_c4_4e10_M3U96a.json",
3 + "n_points": 9,
4 + "c_extrapolated": -0.48618,
5 + "d": -1.7226,
6 + "c_loo_range": [
7 + -0.49422,
8 + -0.48193
9 + ],
10 + "max_abs_residual": 0.00252
11 +}
\ No newline at end of file
added data/gap_histogram_1e8.csv +98 −0
@@ -0,0 +1,98 @@
1 +gap,count
2 +1,1
3 +2,440312
4 +4,440257
5 +6,768752
6 +8,334180
7 +10,430016
8 +12,538382
9 +14,293201
10 +16,215804
11 +18,384738
12 +20,202922
13 +22,175945
14 +24,257548
15 +26,119465
16 +28,129567
17 +30,222847
18 +32,68291
19 +34,71248
20 +36,114028
21 +38,51756
22 +40,60761
23 +42,86637
24 +44,34881
25 +46,29327
26 +48,49824
27 +50,27522
28 +52,20595
29 +54,33593
30 +56,16595
31 +58,14611
32 +60,28439
33 +62,8496
34 +64,8823
35 +66,15579
36 +68,6200
37 +70,8813
38 +72,8453
39 +74,4316
40 +76,3580
41 +78,6790
42 +80,3281
43 +82,2362
44 +84,4668
45 +86,1597
46 +88,1637
47 +90,3337
48 +92,1083
49 +94,971
50 +96,1641
51 +98,851
52 +100,878
53 +102,1059
54 +104,494
55 +106,404
56 +108,711
57 +110,454
58 +112,330
59 +114,487
60 +116,191
61 +118,181
62 +120,433
63 +122,131
64 +124,145
65 +126,204
66 +128,76
67 +130,78
68 +132,132
69 +134,50
70 +136,40
71 +138,93
72 +140,57
73 +142,30
74 +144,51
75 +146,22
76 +148,34
77 +150,37
78 +152,20
79 +154,13
80 +156,23
81 +158,10
82 +160,11
83 +162,8
84 +164,5
85 +166,1
86 +168,8
87 +170,6
88 +172,1
89 +174,3
90 +176,5
91 +178,4
92 +180,4
93 +182,1
94 +184,1
95 +196,1
96 +198,1
97 +210,2
98 +220,1
added data/gap_mod6_start_1e8.csv +7 −0
@@ -0,0 +1,7 @@
1 +p_mod6,gap_mod6,count
2 +1,0,1264047
3 +1,4,1616470
4 +2,1,1
5 +3,2,1
6 +5,0,1264465
7 +5,2,1616470
added data/gapstats_1e8.json +210 −0
@@ -0,0 +1,210 @@
1 +{
2 + "limit": 100000000,
3 + "scan_seconds": 0.3,
4 + "checkpoints": {
5 + "1000000": {
6 + "jumping_champion": 6,
7 + "N2": 8169,
8 + "N4": 8143,
9 + "N6": 13549,
10 + "N2_gt_N4": true,
11 + "mult6_local_max_up_to": 66,
12 + "pearson_consecutive_gaps": -0.04314,
13 + "HL_ratio_obs_over_pred": {
14 + "2": 1.1638,
15 + "4": 0.9079,
16 + "6": 2.127,
17 + "8": 1.0986,
18 + "10": 1.6979,
19 + "12": 0.8438,
20 + "14": 1.4478,
21 + "16": 1.167,
22 + "18": 1.2474,
23 + "20": 1.3534,
24 + "22": 1.4405,
25 + "24": 0.9367,
26 + "26": 1.0754,
27 + "28": 1.3251,
28 + "30": 1.1094,
29 + "32": 0.8116,
30 + "34": 0.9628,
31 + "36": 0.5327,
32 + "38": 0.7818,
33 + "40": 0.7755,
34 + "42": 0.7303,
35 + "44": 0.7661,
36 + "46": 0.6873,
37 + "48": 0.4844,
38 + "50": 0.4916,
39 + "52": 0.545,
40 + "54": 0.5899,
41 + "56": 0.3724,
42 + "58": 0.6091,
43 + "60": 0.6092,
44 + "62": 0.246,
45 + "64": 0.4307,
46 + "66": 0.4863,
47 + "68": 0.3178,
48 + "70": 0.4408,
49 + "72": 0.1881,
50 + "74": 0.4658,
51 + "76": 0.2717,
52 + "78": 0.3339,
53 + "80": 0.131,
54 + "82": 0.3591,
55 + "84": 0.2448,
56 + "86": 0.3905,
57 + "88": 0.1138,
58 + "90": 0.2385,
59 + "92": 0.1546,
60 + "96": 0.2053,
61 + "98": 0.206,
62 + "100": 0.2879,
63 + "112": 0.5587,
64 + "114": 0.4069
65 + }
66 + },
67 + "10000000": {
68 + "jumping_champion": 6,
69 + "N2": 58980,
70 + "N4": 58621,
71 + "N6": 99987,
72 + "N2_gt_N4": true,
73 + "mult6_local_max_up_to": 66,
74 + "pearson_consecutive_gaps": -0.03811,
75 + "HL_ratio_obs_over_pred": {
76 + "2": 1.1489,
77 + "4": 0.871,
78 + "6": 2.0393,
79 + "8": 1.0585,
80 + "10": 1.6133,
81 + "12": 0.8326,
82 + "14": 1.424,
83 + "16": 1.1671,
84 + "18": 1.2485,
85 + "20": 1.3614,
86 + "22": 1.3773,
87 + "24": 0.9891,
88 + "26": 1.141,
89 + "28": 1.4147,
90 + "30": 1.2232,
91 + "32": 0.8903,
92 + "34": 1.0759,
93 + "36": 0.6405,
94 + "38": 0.9417,
95 + "40": 0.8398,
96 + "42": 0.9292,
97 + "44": 0.8687,
98 + "46": 0.8307,
99 + "48": 0.736,
100 + "50": 0.7305,
101 + "52": 0.7709,
102 + "54": 0.7439,
103 + "56": 0.5409,
104 + "58": 0.8328,
105 + "60": 0.7985,
106 + "62": 0.5599,
107 + "64": 0.6578,
108 + "66": 0.6318,
109 + "68": 0.5484,
110 + "70": 0.6432,
111 + "72": 0.4055,
112 + "74": 0.4956,
113 + "76": 0.503,
114 + "78": 0.5206,
115 + "80": 0.3945,
116 + "82": 0.3845,
117 + "84": 0.5277,
118 + "86": 0.3299,
119 + "88": 0.4046,
120 + "90": 0.4117,
121 + "92": 0.274,
122 + "94": 0.3291,
123 + "96": 0.3057,
124 + "98": 0.2677,
125 + "100": 0.2272,
126 + "102": 0.237,
127 + "104": 0.3407,
128 + "106": 0.2218,
129 + "108": 0.1842,
130 + "110": 0.1903,
131 + "112": 0.236,
132 + "114": 0.1681,
133 + "116": 0.2483,
134 + "118": 0.1616,
135 + "120": 0.1492
136 + }
137 + },
138 + "100000000": {
139 + "jumping_champion": 6,
140 + "N2": 440312,
141 + "N4": 440257,
142 + "N6": 768752,
143 + "N2_gt_N4": true,
144 + "mult6_local_max_up_to": 66,
145 + "pearson_consecutive_gaps": -0.03161,
146 + "HL_ratio_obs_over_pred": {
147 + "2": 1.1233,
148 + "4": 0.8411,
149 + "6": 1.9794,
150 + "8": 1.0354,
151 + "10": 1.5517,
152 + "12": 0.8181,
153 + "14": 1.3853,
154 + "16": 1.1575,
155 + "18": 1.2411,
156 + "20": 1.3923,
157 + "22": 1.3623,
158 + "24": 1.0072,
159 + "26": 1.1746,
160 + "28": 1.4342,
161 + "30": 1.2777,
162 + "32": 0.9567,
163 + "34": 1.1223,
164 + "36": 0.6928,
165 + "38": 1.0298,
166 + "40": 0.8962,
167 + "42": 1.0292,
168 + "44": 0.9837,
169 + "46": 0.9288,
170 + "48": 0.8448,
171 + "50": 0.8412,
172 + "52": 0.9229,
173 + "54": 0.861,
174 + "56": 0.6814,
175 + "58": 0.9254,
176 + "60": 0.9739,
177 + "62": 0.6775,
178 + "64": 0.7893,
179 + "66": 0.7561,
180 + "68": 0.6979,
181 + "70": 0.7817,
182 + "72": 0.5203,
183 + "74": 0.6855,
184 + "76": 0.6376,
185 + "78": 0.6595,
186 + "80": 0.5209,
187 + "82": 0.5931,
188 + "84": 0.6405,
189 + "86": 0.5041,
190 + "88": 0.5793,
191 + "90": 0.595,
192 + "92": 0.4816,
193 + "94": 0.484,
194 + "96": 0.4483,
195 + "98": 0.4487,
196 + "100": 0.3112,
197 + "102": 0.4079,
198 + "104": 0.4355,
199 + "106": 0.3991,
200 + "108": 0.2648,
201 + "110": 0.406,
202 + "112": 0.36,
203 + "114": 0.3723,
204 + "116": 0.3333,
205 + "118": 0.3539,
206 + "120": 0.3075
207 + }
208 + }
209 + }
210 +}
\ No newline at end of file
added data/gapstats_1e8.npz +0 −0

Binary file not shown.

added data/maximal_gaps_1e8.csv +26 −0
@@ -0,0 +1,26 @@
1 +gap,start_prime,merit,csg_ratio
2 +1,2,1.442695,2.081369
3 +2,3,1.820478,1.657071
4 +4,7,2.055593,1.056366
5 +6,23,1.913574,0.610294
6 +8,89,1.782278,0.397065
7 +14,113,2.961466,0.626449
8 +18,523,2.875592,0.459390
9 +20,887,2.946443,0.434076
10 +22,1129,3.129851,0.445271
11 +34,1327,4.728345,0.657566
12 +36,9551,3.928244,0.428642
13 +44,15683,4.554709,0.471486
14 +52,19609,5.261164,0.532305
15 +72,31397,6.953520,0.671548
16 +86,155921,7.192377,0.601515
17 +96,360653,7.502537,0.586334
18 +112,370261,8.735012,0.681254
19 +114,492113,8.697998,0.663642
20 +118,1349533,8.359741,0.592248
21 +132,1357201,9.347823,0.661983
22 +148,2010733,10.197044,0.702566
23 +154,4652353,10.030689,0.653342
24 +180,17051707,10.809668,0.649161
25 +210,20831323,12.461452,0.739466
26 +220,47326693,12.448660,0.704405
added data/model_c4.json +50 −0
@@ -0,0 +1,50 @@
1 +{
2 + "model": "P(g1,g2) ~ W_HL(0,g1,g1+g2) * exp(-(g1+g2)/lambda)",
3 + "predictions": {
4 + "1e8": {
5 + "lambda": 18.4207,
6 + "rho": -0.008638,
7 + "rho_times_lambda": -0.15911
8 + },
9 + "1e9": {
10 + "lambda": 20.7233,
11 + "rho": -0.007659,
12 + "rho_times_lambda": -0.15872
13 + },
14 + "1e10": {
15 + "lambda": 23.0259,
16 + "rho": -0.006868,
17 + "rho_times_lambda": -0.15814
18 + },
19 + "4e10": {
20 + "lambda": 24.4121,
21 + "rho": -0.006461,
22 + "rho_times_lambda": -0.15774
23 + },
24 + "1e12": {
25 + "lambda": 27.631,
26 + "rho": -0.005673,
27 + "rho_times_lambda": -0.15676
28 + },
29 + "1e15": {
30 + "lambda": 34.5388,
31 + "rho": -0.004487,
32 + "rho_times_lambda": -0.15498
33 + },
34 + "1e20": {
35 + "lambda": 46.0517,
36 + "rho": -0.003335,
37 + "rho_times_lambda": -0.15358
38 + },
39 + "1e30": {
40 + "lambda": 69.0776,
41 + "rho": -0.002218,
42 + "rho_times_lambda": -0.1532
43 + },
44 + "1e40": {
45 + "lambda": 92.1034,
46 + "rho": -0.001671,
47 + "rho_times_lambda": -0.15388
48 + }
49 + }
50 +}
\ No newline at end of file
added data/summary_1e8.json +30 −0
@@ -0,0 +1,30 @@
1 +{
2 + "limit": 100000000,
3 + "n_gaps": 5761454,
4 + "scan_seconds": 0.6,
5 + "largest_gap": {
6 + "gap": 220,
7 + "after_prime": 47326693
8 + },
9 + "n_maximal_gaps": 25,
10 + "jumping_champion": 6,
11 + "gap_fraction_mod6": {
12 + "0": 0.4388669943385819,
13 + "1": 1.7356729742179665e-07,
14 + "2": 0.28056650283070905,
15 + "3": 0.0,
16 + "4": 0.28056632926341163,
17 + "5": 0.0
18 + },
19 + "best_merit": {
20 + "merit": 12.46145,
21 + "gap": 210,
22 + "after_prime": 20831323
23 + },
24 + "best_csg": {
25 + "csg": 0.73947,
26 + "gap": 210,
27 + "after_prime": 20831323
28 + },
29 + "method": "segmented sieve of Eratosthenes, numpy, deterministic (no seed)"
30 +}
\ No newline at end of file
added journal.md +203 −0
@@ -0,0 +1,203 @@
1 +<!--
2 +journal.md — Prime Mystery Engine: research journal
3 +Author: Simon-Pierre Boucher — contact@spboucher.ai
4 +-->
5 +
6 +# Research Journal
7 +
8 +## 2026-08-06 — Phase 0: Setup
9 +
10 +- Created structure: `/src`, `/data`, `/certs`, `journal.md`, `conjectures.md`, `records.md`.
11 +- Implemented `src/core.py`: odd-only numpy sieve, segmented sieve, deterministic Miller–Rabin
12 + (12 bases, valid for n < 3.317×10^24, Sorenson–Webster), BPSW (strong Lucas, Selfridge params).
13 +- Validation `src/test_core.py`: **37/37 PASS** in 0.3 s —
14 + π(10^2..10^8) exact (incl. π(10^8)=5761455 via segmented sieve), segment-vs-full sieve equality,
15 + full maximal-gap table below 10^6 reproduced, strong-pseudoprime & Carmichael traps,
16 + MR≡sieve on [0,40000), BPSW≡det-MR on 300 random 62-bit integers (seed 42).
17 +- Gate passed → allowed to proceed.
18 +
19 +## 2026-08-06 — Phase 1: Axis selection
20 +
21 +**Chosen axis: #10 — Prime deserts (large gaps).** Justification (5 lines):
22 +1. Fully verifiable: a gap is certified by two primality certificates + compositeness of the interior (sieve or covering) — zero ambiguity.
23 +2. Measurable success criteria exist at every scale: merit g/ln p, CSG ratio g/ln²p, reproduction of the known maximal-gap table below our compute bound.
24 +3. Rich conjecture surface: Cramér/Granville corrections, residues of gap endpoints, jumping champions — testable to 10^9+ locally.
25 +4. State of the art is precisely documented (records.md) so novelty checks are cheap and honest.
26 +5. The PROVER role has a concrete deliverable even without records: covering-system desert constructions with independent re-verification scripts.
27 +
28 +## 2026-08-06 — Phase 2: State of the art
29 +
30 +- Web-checked current records (see `records.md`): largest maximal gap 1854 (2026), merit record 41.94 (2017),
31 + CSG record 0.9206 (Nyman 1999), exhaustive bound ≈ 2×10^19.
32 +- Realistic targets: (i) independent verification of the maximal-gap table to ~4×10^9, (ii) massively tested
33 + statistical conjectures on gaps, (iii) certified covering-system desert + verification pipeline.
34 +
35 +## Cycle 1 — 2026-08-06 — COMPLETE (stopping criterion (a): a conjecture refuted + criterion (b): a construction certified)
36 +
37 +### EXPLORER (Phase 3)
38 +- `src/explore_gaps.py 1e8` (0.6 s): all 25 maximal gaps below 10^8 reproduced = published table;
39 + jumping champion 6; gaps ≡ 0 mod 6 are 43.9% of all gaps; best merit 12.46 (gap 210 after 20831323).
40 +- `src/analyze_gaps.py 1e8` with checkpoints 10^6/10^7/10^8 (bug found & fixed: segment boundaries
41 + must align on checkpoints — symptom: twin counts wrong at sub-segment checkpoints; after fix,
42 + N(2,x) = 8169 / 58980 / 440312 = literature values exactly).
43 +- Key quantified observations: (O1) N2−N4 margins +26/+359/+55 — tiny and non-monotone;
44 + (O2) multiples of 6 strict local maxima up to g = 66 at all three scales; (O3) ρ(x)·ln x ≈ −0.6 stable;
45 + (O4) naive HL gap model exp(−g/ln t) shows structured deviations (deficit at g = 36, 72, 100, 108) —
46 + parked for a future cycle (needs inclusion–exclusion model before conjecturing).
47 +
48 +### CONJECTURER (Phase 4)
49 +- C1 champion=6; C2 race N2>N4 for x ≥ 10^6; C3 G6(x) ≥ 66; C4 ρ·ln x ∈ [−0.65,−0.55] → c ≈ −0.6;
50 + C5 max CSG below 4×10^9 = 0.7395. Two trivial/parity statements discarded at birth. See conjectures.md.
51 +
52 +### ADVERSARY (Phase 5)
53 +- `src/adversary_race.py 4e9` (15 s, deterministic): range = 40× discovery.
54 +- **C2 REFUTED** — full-resolution scan (every gap event): first tie after 10^6 at end-prime 80966861,
55 + first strict overtake at 80966933; the race then changes leader forever in range (137.6M events with
56 + D ≤ 0; D(4×10^9) = −2270). Checkpoint-only verification had been misleading — lesson recorded.
57 +- C1, C3, C5 survive (C3 strengthened: G6 jumps to 216 at 10^9). C4 survives with revision (drift of
58 + ρ·ln x toward −0.565; "≈ −0.6 limit" retracted).
59 +- Cramér/Maier test applied: C2's refutation is exactly random-model behavior (fair random walk);
60 + C1/C3/C4 are structural (absent from a Cramér model).
61 +- All 32 maximal gaps below 4×10^9 = published table (independent verification, `data/adversary_4e9.json`).
62 +
63 +### PROVER (Phase 6)
64 +- Attempt 1 (`src/prover_desert.py`): pure covering system, primes ≤ 59, greedy residues + CRT →
65 + certified gap 90 at a 22-digit N, merit 1.78. Correct but weak.
66 +- Attempt 2 (`src/prover_desert2.py`): **hybrid covering** — greedy covers 243/259 positions for every
67 + shift t (PROVEN); 16 holes certified composite for the chosen t by explicit trial factors (15) and one
68 + strong MR witness (1); endpoints deterministic-MR prime (< 3.317×10^24 bound). Result:
69 + **certified gap 260 after N = 1116336781708038449369693 (25 digits), merit 4.6955** —
70 + ×2.9 vs attempt 1, ×4.3 vs the classic primorial baseline (61) with the same primes. NOT a record
71 + (records.md: merit record 41.94) — the deliverable is the certified pipeline.
72 +- Proven statement (elementary): for every t ≥ 0, the 243 covered positions of x + t·59# + i are
73 + composite — an explicit infinite family of long composite runs.
74 +- Independent verifications: `certs/verify_desert.py` PASSED; `certs/verify_c2_refutation.py` PASSED.
75 +
76 +### Decisions
77 +- Cycle closed. Next cycle: C4 at 10^10 (does ρ·ln x stabilize?) — see report_cycle1.md.
78 +
79 +## Cycle 2 — 2026-08-06 — COMPLETE (focus: C4 at 10^10/4×10^10, dual-machine)
80 +
81 +### Setup / budget
82 +- User directive: run C4 at 10^10 on this computer AND on M3U96a. Strategy: identical deterministic
83 + scan on both — laptop to 10^10 (40 s), M3U96a (M3 Ultra) to 4×10^10 (128 s, single-core numpy).
84 + Deploy via `cluster-gateway.sh stage` to `~/cluster-projects/conjoncture/`.
85 +- Sanity gate: `src/cycle2_c4.py 1e8` reproduces cycle-1 values exactly (ρ·ln x = −0.59604 / −0.61428 /
86 + −0.58221 at 10^6/10^7/10^8; D = 26/359/55). PASS.
87 +
88 +### EXPLORER/ADVERSARY (Phases 3+5 merged — C4 escalation)
89 +- **Cross-machine validation:** every overlapping checkpoint (10^6…10^10) bit-for-bit identical
90 + between laptop and M3U96a. Independent-hardware reproducibility ✓.
91 +- ρ₁·ln x continues drifting up: −0.58221 (10^8) → −0.56116 (10^10) → −0.55569 (4×10^10).
92 + Cycle-1 window [−0.65, −0.55] still holds at 4×10^10 — barely.
93 +- New: lag-2 correlation ρ₂·ln x drifts DOWN: −0.235 (10^8) → −0.249 (4×10^10).
94 +- Free continuations: champion = 6 up to 4×10^10 (C1 ✓); G6 ≥ 66 everywhere but NON-monotone
95 + (312 at 10^10 → 282 at 2×10^10 → 276 at 4×10^10 — tail noise, C3 statement unaffected);
96 + C2 race still swinging (D = +2681 at 10^10, +7160 at 2×10^10, −804 at 4×10^10);
97 + all 37 maximal gaps below 4×10^10 = published table incl. 384@20678048297, 394@22367084959,
98 + 456@25056082087 (web-checked OEIS A002386 / t5k.org after a wrong hand-written checklist
99 + briefly suggested a mismatch — the SCAN was right, the from-memory list was wrong; lesson:
100 + never hand-write "known" tables from memory, always fetch).
101 +- Max CSG below 4×10^10 = 0.79535 (gap 456 after 25056082087) — new in-range CSG high, known.
102 +
103 +### CONJECTURER (Phase 4)
104 +- `src/fit_c4.py`: ρ₁·ln x = c + d/ln x fits 9 checkpoints (10^8…4×10^10) with max residual 0.0025:
105 + **c = −0.486 ± 0.010 (LOO), d = −1.72** → C4′ stated (see conjectures.md), with the sharp
106 + falsifiable prediction ρ·ln x(10^12) = −0.549 ± 0.005, and predicted exit of the cycle-1 window
107 + near x ≈ 5×10^11. C6 (lag-2, c₂ ≈ −0.29) stated with lower confidence.
108 +
109 +### PROVER (Phase 6)
110 +- `src/model_c4.py`: first-order HL triple-correlation model P(g1,g2) ∝ S({0,g1,g1+g2})·e^{−(g1+g2)/λ}.
111 + Result: correct sign, correct c + d/λ drift form, magnitude ρ·λ ≈ −0.154 vs observed −0.49:
112 + **quantitatively REJECTED** (factor ~3.2). Honest conclusion: the anticorrelation is dominated by
113 + interior-compositeness constraints (inclusion–exclusion), not by the bare triple singular series.
114 + Building the full inclusion–exclusion model is the open PROVER problem for a later cycle.
115 +
116 +### Decisions
117 +- Cycle closed (criterion (a)-adjacent: C4 superseded by refined C4′ with a falsifiable constant).
118 +- Next precise action: test C4′ at 10^12 — predicted ρ·ln x = −0.549 ± 0.005. Needs ~25× more compute
119 + than 4×10^10 (~1 h on M3U96a single-core, or minutes core-proportionally across the cluster).
120 +
121 +## Cycle 3 — 2026-08-06 — COMPLETE (C4′ tested at 10^12, distributed on the cluster)
122 +
123 +### Setup / budget
124 +- User directive: distribute; exclude M3U96b and M2U64. Design: 1000 chunks of 10^9 (boundaries =
125 + multiples of 10^9 so checkpoints fall on chunk edges), mergeable streaming sums, 5000-wide junction
126 + buffer per chunk (max gap < 10^12 is 540, so the buffer always contains the 2 preceding gaps).
127 +- Small-scale gate first: 10 chunks × 10^9 merged locally reproduce the cycle-2 single-machine values
128 + at 10^10 EXACTLY (ρ, D, G6, champion). PASS.
129 +- Fleet reality vs plan: 18 candidate nodes → m2m16 unreachable; numpy installed on 5 more nodes
130 + (Apple CLT Python 3.9); 5 nodes had NO usable python3 (missing Xcode CLT: m4ma, m4mb, m4mc,
131 + M2M32c, m2m8b) → dropped. Final fleet: 12 nodes / 162 cores, core-proportional quotas
132 + (174/99/99/99/86/86/74/74/62/49/49/49 chunks).
133 +
134 +### Incident (instructive failure)
135 +- First launch: 9/12 nodes crashed instantly — `np.ndarray | None` annotation in core.py is
136 + Python ≥ 3.10 syntax; the 9 nodes run Apple's Python 3.9. Fix: `from __future__ import annotations`
137 + (annotations become lazy). Local test suite re-run (PASS), core.py re-pushed, import-checked on all
138 + nine, relaunched. Lesson: pin the LOWEST interpreter version in the fleet as the compatibility
139 + target, and import-check on every node class before a fleet launch.
140 +- zsh gotchas hit twice in orchestration (`$NODES` not word-split → `${=NODES}`; `$n:c…` eaten as a
141 + history-style modifier in scp targets → quote remote paths).
142 +
143 +### ADVERSARY (Phase 5 — the run)
144 +- 1000/1000 chunks collected; merge verified exact tiling of [2, 10^12).
145 +- **Cross-validation:** merged distributed results equal the cycle-2 single-machine scan exactly at
146 + 10^10, 2×10^10, 4×10^10 (|Δρ·ln x| = 0 at 5 decimals; D identical). Pipeline validated end-to-end.
147 +- **C4′ VERDICT: CONFIRMED.** ρ·ln x(10^12) = −0.54756 ∈ predicted −0.549 ± 0.005.
148 + Old C4 window [−0.65, −0.55] exited between 4×10^11 and 10^12 as predicted (≈ 5×10^11).
149 +- Checkpoints: −0.55381 (10^11), −0.55203 (2×10^11), −0.55027 (4×10^11), −0.54756 (10^12).
150 +- Free continuations: champion = 6 to 10^12 (C1); G6 → 480 at 10^12 (C3, growing); C2 race STILL
151 + swinging at 10^12 (D = −238 out of ~1.5×10^10 gap events — astonishingly tight);
152 + all 48 maximal gaps below 10^12 = published table (A005250/A002386; tail spot-checked online,
153 + ends 540 @ 738832927927); max CSG below 10^12 unchanged in significance (known values).
154 +- Compute: 17,719 core-seconds ≈ 4.9 core-hours; wall ≈ 5 min over two launches.
155 +
156 +### CONJECTURER (Phase 4 — refit)
157 +- Combined 13-point fit (10^8…10^12): c = −0.48449 (LOO [−0.48793, −0.48196]), d = −1.7571,
158 + max residual 0.00245. New predictions: ρ·ln x(10^13) = −0.5432 ± 0.004, ρ·ln x(10^14) = −0.5390 ± 0.004.
159 +- C6 firmed up: c₂ = −0.2806, d₂ = +0.7815 (opposite drift sign vs lag-1 confirmed);
160 + prediction ρ₂·ln x(10^13) = −0.2545 ± 0.004.
161 +
162 +### Decisions
163 +- Cycle closed (criterion (b)-adjacent: prediction confirmed, law strengthened).
164 +- **Fleet policy change (user directive, recorded):** future distributed runs use ONLY
165 + M3U96a + M4M64a + M4M64b, saturated, adding a Metal/GPU path (MLX or Metal kernels) for the sieve.
166 + Benchmark GPU vs numpy before committing a full run.
167 +- Next precise action: cycle 4 — C4′ at 10^13 on the 3-node fleet with a Metal-accelerated sieve
168 + (10× cycle 3's work: ~50 core-hours CPU-only, so the GPU path is the enabler).
169 +
170 +## Cycle 4 — 2026-08-06 — COMPLETE (C4′ and C6 tested at 10^13; CPU+Metal GPU on 3 nodes)
171 +
172 +### Setup / gates
173 +- Fleet per user directive: M3U96a (28c) + M4M64a (16c) + M4M64b (16c), saturated, plus Metal GPU.
174 +- MLX installed where missing (PEP 668 → `--user --break-system-packages`).
175 +- **GPU sieve** (`src/gpu_sieve.py`, `mx.fast.metal_kernel`): grid (blocks × base primes), byte marking
176 + from max(p², first multiple ≥ block), benign write races. Gates passed: byte-identical prime sets vs
177 + CPU sieve on 4 windows up to 10^13; full worker partial identical field-for-field on a 10^9 chunk.
178 + Bench: GPU 1.2–1.3 s vs CPU-core 2.7–3.6 s per 10^9 near 10^12 → GPU ≈ 2–3 CPU cores per node.
179 +- 10,000 chunks of 10^9; measured-rate quotas 4000/3000/3000 with GPU sub-quotas 370/400/400
180 + (27/15/15 CPU workers, 1 core reserved per node for GPU extraction). Local 10-chunk merge gate
181 + reproduced cycle-2 values exactly before launch.
182 +
183 +### The run
184 +- Wall ≈ 50 min; 147,738 core-seconds (~41 core-hours) including ~4% duplicated tail work from a
185 + late rebalance (M4 nodes finished their quotas and absorbed 500 tail chunks of M3U96a's queue;
186 + duplicates are byte-identical, tiling asserted at merge). GPU processed 1170 chunks (11.7%).
187 +- Merge: 10,000 unique chunks tile [2, 10^13) exactly; **all 10 overlapping checkpoints identical to
188 + cycles 2/3** (|Δρ·ln x| = 0, D exact).
189 +
190 +### VERDICTS
191 +- **C4′ CONFIRMED (2nd out-of-sample hit): ρ·ln x(10^13) = −0.54264 vs predicted −0.5432 ± 0.004.**
192 +- **C6 CONFIRMED (1st hit): ρ₂·ln x(10^13) = −0.25247 vs predicted −0.2545 ± 0.004.**
193 +- 16-point refits: c = −0.48354 (LOO ±0.002), d = −1.7772; c₂ = −0.27762, d₂ = +0.7187.
194 + Next: ρ·ln x(10^14) = −0.5387 ± 0.004; ρ₂·ln x(10^14) = −0.2553 ± 0.004.
195 +- Continuations: champion 6 to 10^13; G6 = 576 at 10^13; C2 race STILL swinging (D: −98,967 at 2×10^12,
196 + +8,870 at 10^13); all 53 maximal gaps below 10^13 = published table (new in range: 582, 588, 602,
197 + 652, 674 — tail web-checked t5k/Nicely); max CSG below 10^13 = 0.79754 (gap 652 after 2614941710599).
198 +
199 +### Decisions
200 +- Cycle closed. Next precise action: cycle 5 — 10^14 (~10× work: ~400 CPU-core-hours equivalent;
201 + with the 3-node fleet ≈ 8–9 h wall, GPU included — schedule as an overnight run) to test
202 + ρ·ln x(10^14) = −0.5387 ± 0.004. Alternative next action if compute is deferred: the PROVER
203 + inclusion–exclusion model for the constant c (main open problem).
added paper/main.pdf +0 −0

Binary file not shown.

added paper/main.tex +797 −0
@@ -0,0 +1,797 @@
1 +% =============================================================================
2 +% main.tex — Prime Mystery Engine: methods and results to date (detailed)
3 +% Author: Simon-Pierre Boucher — contact@spboucher.ai
4 +% Build: pdflatex main.tex (twice) or latexmk -pdf main.tex
5 +% =============================================================================
6 +\documentclass[11pt]{article}
7 +
8 +\usepackage[margin=1.05in]{geometry}
9 +\usepackage{amsmath,amssymb,amsthm}
10 +\usepackage{booktabs}
11 +\usepackage{microtype}
12 +\usepackage{algorithm}
13 +\usepackage{algpseudocode}
14 +\usepackage[hidelinks]{hyperref}
15 +\usepackage{xcolor}
16 +
17 +\newtheorem{theorem}{Theorem}
18 +\newtheorem{lemma}{Lemma}
19 +\newtheorem{proposition}{Proposition}
20 +\newtheorem{conjecture}{Conjecture}
21 +\newtheorem{observation}{Observation}
22 +\theoremstyle{definition}
23 +\newtheorem{result}{Result}
24 +\newtheorem{remark}{Remark}
25 +
26 +\newcommand{\lnx}{\ln x}
27 +\newcommand{\statuslbl}[1]{\textsc{\small #1}}
28 +\newcommand{\Sing}{\mathfrak{S}}
29 +\DeclareMathOperator{\li}{li}
30 +
31 +\title{A Verification-First Computational Laboratory for Prime Gap Statistics:\\
32 +Certified Deserts, a Refuted Race Conjecture, and an Empirical Scaling Law\\
33 +for Consecutive-Gap Correlations}
34 +\author{Simon-Pierre Boucher\\
35 +\texttt{contact@spboucher.ai}}
36 +\date{August 6, 2026 --- working draft v3}
37 +
38 +\begin{document}
39 +\maketitle
40 +
41 +\begin{abstract}
42 +We describe an autonomous, verification-first research pipeline for the statistics of gaps
43 +between consecutive primes, and the results it has produced to date. The pipeline enforces a
44 +strict epistemic protocol (every claim is \emph{proven}, \emph{certified}, \emph{verified up
45 +to an explicit bound}, or \emph{explicitly conjectural}) and cycles through four adversarial
46 +roles: exploration, conjecture, attack, and proof. Main outcomes so far: (i) the refutation,
47 +with minimal counterexample, of the conjecture that gaps of size 2 outnumber gaps of size 4
48 +from $10^6$ onward --- the two counts in fact exchange the lead indefinitely throughout the
49 +tested range, with the difference at $10^{12}$ being $-238$ out of $\approx 3.76\times
50 +10^{10}$ gap events; (ii) a fully certified prime desert of length 260 at a 25-digit
51 +integer, constructed by a hybrid covering system with independently re-verifiable
52 +certificates; (iii) an empirical scaling law for the Pearson correlation of consecutive
53 +gaps, $\rho(x)\lnx = c + d/\lnx + o(1/\lnx)$ with $c = -0.4835 \pm 0.002$ and $d = -1.777$,
54 +whose out-of-sample predictions at $x = 10^{12}$ ($-0.549 \pm 0.005$, measured $-0.54756$)
55 +and $x = 10^{13}$ ($-0.5432 \pm 0.004$, measured $-0.54264$) were both subsequently
56 +confirmed; and (iv) the quantitative rejection of the
57 +first-order Hardy--Littlewood triple-correlation model for this constant, which predicts
58 +the correct sign and drift form but only about $30\%$ of the observed magnitude. All scans
59 +are deterministic; their accumulators are exact integers below $2^{53}$, which yields
60 +bit-identical results across heterogeneous machines and exact chunked/merged decomposition.
61 +All published gap tables below $10^{13}$ (53 maximal gaps) were reproduced exactly.
62 +Computations ran on a small cluster of Apple-silicon machines, latterly using a custom
63 +Metal GPU sieve kernel (validated byte-identical against the CPU implementation) alongside
64 +saturated CPU cores; the correlation scan now extends to $10^{13}$, where the law's predictions for both the
65 +lag-1 and lag-2 statistics were confirmed out of sample.
66 +\end{abstract}
67 +
68 +\tableofcontents
69 +
70 +\section{Introduction}
71 +
72 +Let $p_1 < p_2 < \cdots$ denote the primes and
73 +\begin{equation}
74 +g_n \;=\; p_{n+1} - p_n
75 +\end{equation}
76 +the sequence of prime gaps. This project treats the empirical study of $(g_n)$ as an
77 +engineering discipline: no numerical claim is admitted without a certificate, a
78 +re-verification script, or an explicit statement of the verified range. The work reported
79 +here was carried out by an autonomous research loop (the ``Prime Mystery Engine'')
80 +operating under the fixed protocol of Section~\ref{sec:protocol} on the hardware of
81 +Section~\ref{sec:compute}.
82 +
83 +Three kinds of outputs are reported. First, \emph{negative results}: a natural conjecture
84 +about the race between twin gaps ($g=2$) and quadruple gaps ($g=4$) is refuted with a
85 +minimal counterexample (Section~\ref{sec:race}). Second, \emph{constructions}: a certified
86 +prime desert obtained from a hybrid covering system (Section~\ref{sec:desert}). Third,
87 +\emph{quantitative empirical laws}: a two-parameter scaling law for the correlation of
88 +consecutive gaps that has survived two rounds of genuine out-of-sample prediction
89 +(Section~\ref{sec:c4}), together with the failure of the natural first-order model to
90 +explain its constant (Section~\ref{sec:model}).
91 +
92 +Throughout, every statement carries one of the labels \statuslbl{proven},
93 +\statuslbl{certified}, \statuslbl{verified up to $X$}, \statuslbl{conjectured}, or
94 +\statuslbl{refuted}. Nothing below is claimed to be new to the literature unless
95 +explicitly discussed; where we suspect a phenomenon is known, we say so.
96 +
97 +\section{Notation and background}
98 +\label{sec:background}
99 +
100 +$\pi(x)$ is the prime-counting function; the prime number theorem gives
101 +$\pi(x) \sim \li(x) = \int_2^x dt/\ln t$, so the \emph{local scale} of gaps near $x$ is
102 +\begin{equation}
103 +\lambda \;=\; \lnx, \qquad \mathbb{E}[\,g \mid p \approx x\,] \sim \lambda .
104 +\end{equation}
105 +For an even $g$ we write
106 +\begin{equation}
107 +N(g, x) \;=\; \#\{\, n : g_n = g,\; p_{n+1} \le x \,\},
108 +\end{equation}
109 +the number of gaps of size exactly $g$ ending below $x$ (the ``ending-prime'' attribution
110 +is the convention used by all our accumulators; see Section~\ref{sec:merge}).
111 +
112 +\paragraph{Hardy--Littlewood singular series.}
113 +For a finite set $\mathcal{H} = \{h_1, \dots, h_k\}$ of integers let
114 +$\nu_p(\mathcal{H})$ be the number of distinct residues of $\mathcal{H}$ modulo $p$. The
115 +Hardy--Littlewood $k$-tuple conjecture predicts
116 +\begin{equation}
117 +\#\{\, n \le x : n + h_1, \dots, n + h_k \ \text{all prime} \,\}
118 +\;\sim\; \Sing(\mathcal{H}) \int_2^x \frac{dt}{(\ln t)^k},
119 +\qquad
120 +\Sing(\mathcal{H}) \;=\; \prod_{p} \frac{1 - \nu_p(\mathcal{H})/p}{(1 - 1/p)^k}.
121 +\label{eq:HL}
122 +\end{equation}
123 +For a pair $\mathcal{H} = \{0, g\}$ this specializes to
124 +\begin{equation}
125 +\Sing(\{0,g\}) \;=\; C(g) \;=\; 2 C_2 \prod_{\substack{p \mid g \\ p > 2}} \frac{p-1}{p-2},
126 +\qquad
127 +C_2 = \prod_{p > 2}\Bigl(1 - \frac{1}{(p-1)^2}\Bigr) = 0.6601618158\ldots
128 +\label{eq:singular-pair}
129 +\end{equation}
130 +Note $C(2) = C(4) = 2C_2$: the twin and quadruple constants coincide, which is what makes
131 +the race of Section~\ref{sec:race} a second-order question.
132 +
133 +\paragraph{Merit and the Cram\'er--Shanks--Granville ratio.}
134 +For a gap $g$ after $p$,
135 +\begin{equation}
136 +\text{merit}(g, p) = \frac{g}{\ln p},
137 +\qquad
138 +\mathrm{CSG}(g, p) = \frac{g}{\ln^2 p}.
139 +\end{equation}
140 +Cram\'er's heuristic suggests $\limsup \mathrm{CSG} = 1$; Granville's correction argues
141 +$\limsup \ge 2e^{-\gamma} \approx 1.1229$. The largest CSG value ever observed is
142 +$0.9206\ldots$ (Nyman's gap 1132 after 1693182318746371).
143 +
144 +\paragraph{Cram\'er random model.}
145 +In the Cram\'er model each integer $n$ is independently ``prime'' with probability
146 +$1/\ln n$; consecutive gaps near $x$ are then asymptotically i.i.d.\
147 +$\mathrm{Exponential}(\lambda)$ up to discretization,
148 +so \emph{any} statistic that vanishes for independent gaps (such as the correlations of
149 +Section~\ref{sec:c4}) measures genuinely arithmetic structure. We use this as a
150 +significance filter (the ``Cram\'er/Maier test'').
151 +
152 +\section{Methodological protocol}
153 +\label{sec:protocol}
154 +
155 +Work proceeds in cycles. Each cycle passes through four roles in order:
156 +\begin{enumerate}
157 + \item \textbf{Explorer} --- generates data at small scale, checks it against known
158 + values, then scales up; produces quantified observations only.
159 + \item \textbf{Conjecturer} --- converts observations into precise statements with
160 + explicit quantifiers and bounds; discards at birth any statement with a modular
161 + obstruction (admissibility is checked mod $2,3,5,7$ systematically) or that is
162 + trivially known.
163 + \item \textbf{Adversary} --- attacks each conjecture on a range $10\times$--$100\times$
164 + the discovery range, hunting for the \emph{minimal} counterexample rather than
165 + merely re-verifying; applies the Cram\'er/Maier test.
166 + \item \textbf{Prover} --- for survivors, attempts a proof, a reduction to a standard
167 + conjecture, or an explicit model; for constructions, produces a complete
168 + certificate and an independent verification script.
169 +\end{enumerate}
170 +Non-negotiable constraints: small scale before large scale; determinism; every announced
171 +object re-verifiable by a standalone script; every ``known value'' fetched from a source
172 +(OEIS, t5k.org, the Nicely tables) rather than transcribed from memory --- a rule adopted
173 +after an instructive incident (Section~\ref{sec:lessons}).
174 +
175 +\section{Computational infrastructure}
176 +\label{sec:compute}
177 +
178 +\subsection{Hardware}
179 +
180 +All computation ran on Apple-silicon machines from a private cluster: a Mac Studio
181 +M3~Ultra (28 CPU cores, 96\,GB RAM, 80-core GPU; node \texttt{M3U96a}), two Mac Studios
182 +M4~Max (16 CPU cores, 64\,GB each; nodes \texttt{M4M64a/b}), and, in one earlier
183 +campaign, nine additional small nodes (M1--M4 class, 8--16 cores) for a 162-core fleet.
184 +Orchestration ran from a laptop over SSH; file distribution used an rsync gateway
185 +through \texttt{M3U96a}.
186 +
187 +\subsection{Primality primitives}
188 +\label{sec:primitives}
189 +
190 +\paragraph{Deterministic Miller--Rabin.}
191 +Write $n - 1 = 2^s d$ with $d$ odd. A base $a$ with $a \not\equiv 0 \pmod n$ is a
192 +\emph{strong witness} for the compositeness of $n$ if
193 +\begin{equation}
194 +a^{d} \not\equiv 1 \pmod n
195 +\quad\text{and}\quad
196 +a^{2^r d} \not\equiv -1 \pmod n \ \ \text{for all } 0 \le r < s.
197 +\label{eq:mr}
198 +\end{equation}
199 +If no $a$ in the 12-element base set $\{2, 3, 5, 7, 11, 13, 17, 19, 23, 29, 31, 37\}$
200 +is a strong witness, then $n$ is prime for every
201 +\begin{equation}
202 +n \;<\; 3.317\times 10^{24} \;=\; 3{,}317{,}044{,}064{,}679{,}887{,}385{,}961{,}981
203 +\end{equation}
204 +(Sorenson--Webster \cite{SW}). All primality assertions in our certificates stay below
205 +this bound, so they are unconditional.
206 +
207 +\paragraph{Baillie--PSW.}
208 +Beyond the bound (used only as a cross-check, never in a certificate) we apply BPSW: a
209 +strong base-2 test plus a strong Lucas test with Selfridge parameters --- the first
210 +$D \in \{5, -7, 9, -11, \dots\}$ with Jacobi symbol $\bigl(\tfrac{D}{n}\bigr) = -1$,
211 +$P = 1$, $Q = (1-D)/4$. With $n + 1 = 2^s d$, $n$ is a strong Lucas probable prime if
212 +\begin{equation}
213 +U_d \equiv 0 \ \ \text{or} \ \ V_{2^r d} \equiv 0 \pmod n \ \ \text{for some } 0 \le r < s,
214 +\end{equation}
215 +where the Lucas sequences are computed by the binary ladder
216 +\begin{equation}
217 +U_{2k} = U_k V_k, \qquad
218 +V_{2k} = V_k^2 - 2Q^k, \qquad
219 +U_{k+1} = \tfrac{1}{2}(P U_k + V_k), \qquad
220 +V_{k+1} = \tfrac{1}{2}(D U_k + P V_k) \pmod n .
221 +\end{equation}
222 +No BPSW pseudoprime is known; the test is deterministic below $2^{64}$.
223 +
224 +\paragraph{Validation gate.}
225 +No computation is permitted until a 37-check gate passes ($0.3$\,s): $\pi(10^k)$ for
226 +$k \le 8$ (e.g.\ $\pi(10^8) = 5{,}761{,}455$ via the segmented path), equality of
227 +segmented and monolithic sieves on randomized windows, the complete table of maximal
228 +gaps below $10^6$, strong-pseudoprime traps ($2047, 3277, 4033, \dots$), Carmichael
229 +traps ($561, 1105, 41041, 321197185, \dots$), and cross-agreement of the three tests on
230 +$[0, 4\times 10^4)$ and on 300 random 62-bit integers (fixed seed 42).
231 +
232 +\subsection{Segmented sieve and cost model}
233 +\label{sec:sieve}
234 +
235 +To enumerate primes in $[\ell, h)$ we precompute base primes up to $\sqrt{h}$ and, per
236 +segment of length $S = 5\times 10^7$, mark for every base prime $p$ the arithmetic
237 +progression starting at $\max(p^2, \lceil \ell/p \rceil p)$. The marking work per
238 +segment is
239 +\begin{equation}
240 +W(S, h) \;=\; \sum_{p \le \sqrt{h}} \Bigl( \frac{S}{p} + O(1) \Bigr)
241 +\;=\; S \bigl( \ln\ln\sqrt{h} + M \bigr) + O\bigl(\pi(\sqrt h)\bigr),
242 +\label{eq:mertens}
243 +\end{equation}
244 +by Mertens' second theorem ($M = 0.2614\ldots$), i.e.\ $O(\ln\ln h)$ byte-writes per
245 +integer --- the procedure is memory-bandwidth-bound, which is why both the NumPy
246 +implementation and the GPU kernel of Section~\ref{sec:gpu} are effective. Measured
247 +throughput: $\approx 3\times 10^8$ integers per second per CPU core near $10^{10}$,
248 +falling mildly with $h$ through the $\pi(\sqrt h)$ term ($\pi(\sqrt{10^{13}}) =
249 +216{,}816$ base primes).
250 +
251 +\subsection{Mergeable streaming statistics; exact distribution}
252 +\label{sec:merge}
253 +
254 +All statistics of a scan are accumulated in objects that form a commutative monoid
255 +under merging, so a scan of $[2, X)$ decomposes into independent chunks:
256 +
257 +\begin{itemize}
258 +\item \textbf{Histogram} $H[g] = N(g, \cdot)$ and the race counter
259 + $D = H[2] - H[4]$ (integer vectors; merge $=$ addition).
260 +\item \textbf{Record lists}: chunk-local running maxima $(g, p)$; the global merge
261 + keeps increasing records in range order (any global record is a chunk-local
262 + record, so no information is lost).
263 +\item \textbf{Pearson sums}: for lag $j \in \{1, 2\}$,
264 +\begin{equation}
265 +T^{(j)} \;=\; \Bigl( n,\ \textstyle\sum a,\ \sum b,\ \sum a^2,\ \sum b^2,\ \sum ab \Bigr),
266 +\qquad (a, b) = (g_{n-j},\, g_n),
267 +\label{eq:sums}
268 +\end{equation}
269 + summed over all counted pairs; merge $=$ componentwise addition. The Pearson
270 + coefficient is reconstructed at any checkpoint as
271 +\begin{equation}
272 +\rho \;=\;
273 +\frac{\frac{1}{n}\sum ab - \frac{1}{n}\sum a \cdot \frac{1}{n}\sum b}
274 +{\sqrt{\bigl(\frac{1}{n}\sum a^2 - (\frac{1}{n}\sum a)^2\bigr)
275 + \bigl(\frac{1}{n}\sum b^2 - (\frac{1}{n}\sum b)^2\bigr)}} .
276 +\label{eq:pearson}
277 +\end{equation}
278 +\end{itemize}
279 +
280 +\begin{proposition}[attribution and tiling]
281 +Attribute each gap $g_n$ (and each pair $(g_{n-j}, g_n)$) to the chunk containing the
282 +\emph{ending} prime $p_{n+1}$. Then chunks $[\ell_i, \ell_{i+1})$ tiling $[2, X)$ induce a
283 +partition of all gap events, and the ordered cumulative merge reproduces the exact prefix
284 +statistics at every chunk boundary.
285 +\end{proposition}
286 +
287 +\begin{lemma}[junction buffer]
288 +\label{lem:buffer}
289 +Let $G^* $ be the maximal prime gap below $X$. If each worker sieves from $\ell - B$
290 +with $B \ge 3G^* + 1$, it observes the two gaps preceding its first counted gap, so the
291 +lag-1 and lag-2 pair sums \eqref{eq:sums} are exact across chunk boundaries.
292 +\end{lemma}
293 +\begin{proof}[Proof sketch]
294 +The first counted gap ends at some $p_{n+1} \ge \ell$, hence starts at
295 +$p_n \ge \ell - G^*$; its two predecessors start at $\ge \ell - 3G^*$. Any window of
296 +length $G^*$ below $X$ contains a prime by definition of $G^*$, so sieving from
297 +$\ell - B$ recovers a full prime stream from below $\ell - 3G^*$.
298 +\end{proof}
299 +We use $B = 5000$; the maximal gap below $10^{13}$ is well under $800$
300 +($G^* = 540$ below $10^{12}$), so the margin is ample.
301 +
302 +\begin{proposition}[exactness of the accumulators]
303 +\label{prop:exact}
304 +All components of $H$, $D$, and $T^{(j)}$ are integers, and every partial sum arising in
305 +their accumulation is bounded by $\sum g_n \le X \le 10^{13} < 2^{53}$, resp.\
306 +$\sum g_n^2 \approx \lambda^2 \pi(X) \lesssim 10^{15}$ for the ranges used here. Hence
307 +the IEEE-754 double-precision accumulation is \emph{exact} (every intermediate value is
308 +an integer representable in the 53-bit significand), the results are independent of
309 +summation order and of the chunk decomposition, and identical bit-for-bit across
310 +machines; the only rounding occurs in the final evaluation of \eqref{eq:pearson}.
311 +\end{proposition}
312 +
313 +This proposition is not academic: it is what makes the cross-machine and
314 +chunked-vs-monolithic validations come out \emph{exactly} equal (laptop vs.\ M3~Ultra
315 +up to $4\times 10^{10}$; 10-chunk merge vs.\ single scan at $10^{10}$; 1000-chunk
316 +distributed merge vs.\ prior scans at $10^{10}, 2\times 10^{10}, 4\times 10^{10}$ ---
317 +race counters exactly equal, correlations equal to all published decimals).
318 +
319 +\subsection{A Metal GPU sieve kernel}
320 +\label{sec:gpu}
321 +
322 +For the $10^{13}$ campaign the marking phase moved to the GPU via a custom Metal kernel
323 +(MLX \texttt{fast.metal\_kernel}). The segment is a byte array initialized to 1; the
324 +dispatch grid is $(\text{blocks}) \times (\text{base primes})$ with block width
325 +$10^6$; the thread for block $b$ and prime $p$ executes
326 +\begin{equation}
327 +\text{for } m = \max\bigl(p^2, \lceil \mathrm{blo}/p\rceil\, p\bigr),\,
328 +m {+} p,\, m {+} 2p, \dots < \mathrm{bhi}: \quad \texttt{out}[m - \ell] \leftarrow 0,
329 +\end{equation}
330 +in 64-bit integer arithmetic. Write races are benign (all threads store the same value).
331 +Prime extraction (\texttt{nonzero} + differences) remains on the CPU; one CPU core per
332 +node is reserved for it. Acceptance criteria (all met): byte-identical prime sets to the
333 +CPU sieve on windows spanning $[2, 10^6)$ to $[10^{13} - 5\times 10^7,\, 10^{13})$, and
334 +a worker partial identical field-for-field to the CPU worker's on a full $10^9$ chunk.
335 +Measured single-chunk times near $10^{12}$ ($10^9$ integers): GPU $1.2$--$1.3$\,s versus
336 +$2.7$--$3.6$\,s per CPU core, on all three nodes; the GPU path is thus worth
337 +$\approx 2$--$3$ CPU cores per node; in the $10^{13}$ campaign it processed 1170 of the
338 +10{,}000 chunks ($11.7\%$).
339 +
340 +\begin{algorithm}[t]
341 +\caption{Distributed correlation scan of $[2, X)$ (one worker chunk)}
342 +\begin{algorithmic}[1]
343 +\Require chunk $[\ell, h)$ with $\ell, h$ multiples of $10^9$ (or $\ell = 2$); buffer $B = 5000$
344 +\State sieve primes in $[\ell - B, h)$ segment by segment (NumPy bytes or Metal kernel)
345 +\State form gaps between consecutive primes; \emph{count} those whose end prime is $\ge \ell$
346 +\State accumulate $H$, $D$, records, and the lag-1/lag-2 sums $T^{(1)}, T^{(2)}$ of
347 + \eqref{eq:sums}, using the buffered predecessor gaps at the boundary
348 + (Lemma~\ref{lem:buffer})
349 +\State emit all accumulators as JSON \Comment{ordered cumulative merge afterwards}
350 +\end{algorithmic}
351 +\end{algorithm}
352 +
353 +\subsection{Campaigns}
354 +
355 +\begin{center}
356 +\begin{tabular}{@{}llllll@{}}
357 +\toprule
358 +Campaign & Range & Fleet & Chunks & Wall time & Cost \\
359 +\midrule
360 +Cycle 1 & $4\times 10^9$ & 1 laptop core & --- & 15\,s & --- \\
361 +Cycle 2 & $10^{10}$ / $4\times 10^{10}$ & laptop / M3 Ultra (1 core) & --- & 40 / 128\,s & --- \\
362 +Cycle 3 & $10^{12}$ & 12 nodes, 162 CPU cores & 1000 & $\approx 5$\,min & 17{,}719 core-s \\
363 +Cycle 4 & $10^{13}$ & 3 nodes, 58 CPU cores $+$ 3 GPUs & 10{,}000 & $\approx 50$\,min & 147{,}738 core-s \\
364 +\bottomrule
365 +\end{tabular}
366 +\end{center}
367 +Chunk quotas are proportional to measured per-node rates (e.g.\ cycle 4:
368 +4000/3000/3000 with per-node GPU shares 370/400/400).
369 +
370 +\section{Result I: the 2-vs-4 race conjecture is false}
371 +\label{sec:race}
372 +
373 +Define
374 +\begin{equation}
375 +D(x) \;=\; N(2, x) - N(4, x).
376 +\end{equation}
377 +By \eqref{eq:singular-pair}, $C(2) = C(4)$, so under Hardy--Littlewood
378 +$N(2,x) \sim N(4,x)$ and $\operatorname{sign} D$ is a second-order phenomenon,
379 +analogous to a Chebyshev bias. Checkpoints at $10^6, 10^7, 10^8$ gave
380 +$D = +26, +359, +55$, suggesting:
381 +
382 +\begin{conjecture}[C2; \statuslbl{refuted}]
383 +$N(2, x) > N(4, x)$ for all $x \ge 10^6$.
384 +\end{conjecture}
385 +
386 +\begin{result}[\statuslbl{refuted with minimal counterexample}]
387 +\label{res:race}
388 +A full-resolution scan (every gap event inspected via cumulative sums) shows
389 +\[
390 +D = 0 \ \text{first at end-prime } 80{,}966{,}861, \qquad
391 +D = -1 \ \text{first at end-prime } 80{,}966{,}933 .
392 +\]
393 +The race then changes leader indefinitely throughout the tested range. Up to
394 +$4\times 10^9$ there are $137{,}574{,}763$ gap events with $D \le 0$; sampled values:
395 +\begin{center}
396 +\begin{tabular}{@{}lcccccccc@{}}
397 +\toprule
398 +$x$ & $10^9$ & $4{\cdot}10^9$ & $10^{10}$ & $2{\cdot}10^{10}$ & $4{\cdot}10^{10}$ &
399 +$10^{11}$ & $4{\cdot}10^{11}$ & $10^{12}$ \\
400 +\midrule
401 +$D(x)$ & $-173$ & $-2270$ & $+2681$ & $+7160$ & $-804$ & $+2888$ & $+22767$ & $-238$ \\
402 +\bottomrule
403 +\end{tabular}
404 +\end{center}
405 +\end{result}
406 +
407 +\paragraph{Null model.} Restricting to steps where $g \in \{2, 4\}$, $D$ performs a
408 +$\pm 1$ walk with $\approx$ balanced steps; with $N_{\pm}(x) = N(2,x) + N(4,x)$ steps,
409 +the diffusive scale is $\sqrt{N_\pm(x)}$ --- about $6\times 10^4$ at $x = 10^{12}$
410 +(where $N_\pm \approx 3.7\times 10^9$). The observed $|D|$ stays well within this
411 +envelope (max sampled $|D| \approx 2.3\times 10^4$), so the data are consistent with a
412 +recurrent, unbiased race; the original conjecture was precisely the anti-random-model
413 +claim, and the Cram\'er/Maier test correctly flagged it. The refutation is
414 +re-verifiable by a standalone, stdlib-only script
415 +(\texttt{certs/verify\_c2\_refutation.py}, $\approx 30$\,s).
416 +
417 +\paragraph{Methodological point.} The conjecture \emph{passed} at every decade
418 +checkpoint while failing $\sim 1.4\times 10^8$ times in between: race-type statements
419 +must be attacked at full resolution. Open follow-up: whether the \emph{logarithmic
420 +density} of the $\{D > 0\}$ region exists and is biased (the analogue of the
421 +Rubinstein--Sarnak analysis for $\pi(x; 4, 1)$ vs $\pi(x; 4, 3)$).
422 +
423 +\section{Result II: a certified prime desert from a hybrid covering system}
424 +\label{sec:desert}
425 +
426 +\begin{result}[\statuslbl{certified}]
427 +\label{res:desert}
428 +Let $N = 1{,}116{,}336{,}781{,}708{,}038{,}449{,}369{,}693$ (25 digits). Then $N$ and
429 +$N + 260$ are consecutive primes: the gap after $N$ is exactly $260$, with
430 +$\mathrm{merit} = 260/\ln N = 4.6955$.
431 +\end{result}
432 +
433 +\paragraph{Construction.}
434 +Let $Q = \{2, 3, \dots, 59\}$ be the 17 primes up to $m = 59$ and $P = 59\# =
435 +\prod_{p \in Q} p = 1{,}922{,}760{,}350{,}154{,}212{,}639{,}070$. Choose one residue
436 +class $a_p \bmod p$ per $p \in Q$, greedily maximizing coverage of positions
437 +$\{1, \dots, L\}$, $L = 259$, under the constraint that no class contains $0$ or
438 +$L + 1$:
439 +\begin{equation}
440 +\mathcal{C} \;=\; \bigl\{ i \in \{1,\dots,L\} : \exists\, p \in Q,\ i \equiv a_p
441 +\!\!\pmod p \bigr\}, \qquad |\mathcal{C}| = 243 .
442 +\end{equation}
443 +By the Chinese Remainder Theorem pick $x$ with
444 +\begin{equation}
445 +x \;\equiv\; -a_p \pmod{p} \quad \text{for all } p \in Q, \qquad 0 \le x < P .
446 +\label{eq:crt}
447 +\end{equation}
448 +
449 +\begin{lemma}[\statuslbl{proven}]
450 +For every integer $t \ge 0$ and every $i \in \mathcal{C}$, the number $x + tP + i$ is
451 +composite.
452 +\end{lemma}
453 +\begin{proof}
454 +If $i \equiv a_p \pmod p$ then $x + tP + i \equiv -a_p + 0 + a_p \equiv 0 \pmod p$ by
455 +\eqref{eq:crt}, and $x + tP + i > P > p$, so $p$ is a proper divisor.
456 +\end{proof}
457 +
458 +This yields an explicit infinite family of runs of $\ge 243$ composites in arithmetic
459 +progression of modulus $P$. For the $16$ uncovered positions (``holes'') and the two
460 +endpoints, search over the shift $t$: at $t = 580$,
461 +\begin{itemize}
462 +\item $N = x + 580 P$ and $N + 260$ both pass deterministic Miller--Rabin
463 + (Section~\ref{sec:primitives}; both $< 3.317\times 10^{24}$, so unconditionally
464 + prime);
465 +\item each hole $N + i$ is composite, certified by an explicit prime factor found by
466 + trial division (15 holes) or by an explicit strong witness $a$ satisfying
467 + \eqref{eq:mr} (1 hole) --- a strong witness \emph{is} a compositeness proof.
468 +\end{itemize}
469 +Hence every $N + i$, $1 \le i \le 259$, carries an explicit compositeness certificate
470 +and both endpoints an unconditional primality certificate: the gap is exactly 260.
471 +
472 +\paragraph{Endpoint-search heuristic.}
473 +$N$ and $N + L + 1$ are constructed coprime to all $p \le 59$, so by the standard
474 +sieve heuristic each is prime with probability
475 +$\approx e^{\gamma} \ln(59) / \ln N \approx 0.148$ per shift, both endpoints with
476 +probability $\approx 0.022$; including hole compositeness
477 +($\approx 0.856^{16} \approx 0.083$) the expected number of shifts is a few hundred
478 +($t = 580$ observed), against
479 +$t_{\max} = \lfloor (3.317\times 10^{24} - x - L - 1)/P \rfloor \approx 1725$ available
480 +below the deterministic bound.
481 +
482 +\paragraph{Calibration.}
483 +Same prime set $Q$: the classic primorial construction ($a_p \equiv 0$) yields a gap of
484 +only $61$ (next prime after 59); pure covering with no holes yields $90$; the hybrid
485 +yields $260$. This is \emph{not} a record of any kind (Section~\ref{sec:sota}); the
486 +deliverable is the certified, independently re-verifiable pipeline
487 +(\texttt{certs/desert\_certificate.json} $+$ \texttt{certs/verify\_desert.py}, the
488 +verifier carrying its own independent Miller--Rabin implementation).
489 +
490 +\section{Result III: a scaling law for consecutive-gap correlations}
491 +\label{sec:c4}
492 +
493 +Let $\rho(x)$ (resp.\ $\rho_2(x)$) be the Pearson correlation \eqref{eq:pearson} of the
494 +pairs $(g_n, g_{n+1})$ (resp.\ $(g_n, g_{n+2})$) over all gaps ending below $x$. In the
495 +Cram\'er model $\rho \equiv \rho_2 \equiv 0$; the observed anticorrelation is
496 +structural. All 16 checkpoints:
497 +
498 +\begin{center}
499 +\begin{tabular}{@{}lcc@{}}
500 +\toprule
501 +$x$ & $\rho(x)\lnx$ & $\rho_2(x)\lnx$ \\
502 +\midrule
503 +$10^{8}$ & $-0.58221$ & $-0.23535$ \\
504 +$2\times 10^{8}$ & $-0.57397$ & $-0.23980$ \\
505 +$4\times 10^{8}$ & $-0.57089$ & $-0.24457$ \\
506 +$10^{9}$ & $-0.57055$ & $-0.24252$ \\
507 +$2\times 10^{9}$ & $-0.56652$ & $-0.24340$ \\
508 +$4\times 10^{9}$ & $-0.56513$ & $-0.24524$ \\
509 +$10^{10}$ & $-0.56116$ & $-0.24783$ \\
510 +$2\times 10^{10}$ & $-0.55954$ & $-0.24814$ \\
511 +$4\times 10^{10}$ & $-0.55569$ & $-0.24929$ \\
512 +$10^{11}$ & $-0.55381$ & $-0.24949$ \\
513 +$2\times 10^{11}$ & $-0.55203$ & $-0.25035$ \\
514 +$4\times 10^{11}$ & $-0.55027$ & $-0.25076$ \\
515 +$10^{12}$ & $-0.54756$ & $-0.25141$ \\
516 +$2\times 10^{12}$ & $-0.54601$ & $-0.25162$ \\
517 +$4\times 10^{12}$ & $-0.54448$ & $-0.25205$ \\
518 +$10^{13}$ & $-0.54264$ & $-0.25247$ \\
519 +\bottomrule
520 +\end{tabular}
521 +\end{center}
522 +
523 +\begin{conjecture}[C4$'$; \statuslbl{conjectured}, one prediction confirmed]
524 +\label{conj:c4p}
525 +With $\lambda = \lnx$,
526 +\begin{equation}
527 +\rho(x)\,\lambda \;=\; c + \frac{d}{\lambda} + o\!\Bigl(\frac{1}{\lambda}\Bigr),
528 +\qquad c = -0.4835 \pm 0.002, \quad d = -1.777 .
529 +\label{eq:c4law}
530 +\end{equation}
531 +\end{conjecture}
532 +
533 +The $1/\lambda$ form is the natural first correction: the mean gap is $\lambda$, all
534 +discrete/arithmetic effects (the parity of gaps, the mod-6 structure, the singular
535 +series) enter through moments whose relative size is $O(1/\lambda)$, and the model of
536 +Section~\ref{sec:model} produces exactly this drift shape.
537 +
538 +\paragraph{Fit.}
539 +Least squares on the 16 points $(\lambda_i, v_i)$, $v_i = \rho(x_i)\lambda_i$:
540 +\begin{equation}
541 +(\hat c, \hat d) \;=\; \arg\min_{c, d} \sum_{i=1}^{13}
542 +\Bigl( v_i - c - \frac{d}{\lambda_i} \Bigr)^{\!2}
543 +\;=\; (-0.48354,\ -1.7772),
544 +\end{equation}
545 +with maximum residual $0.00255$ and leave-one-out range
546 +$\hat c \in [-0.48570, -0.48165]$ (16 refits, each dropping one point). The quoted
547 +uncertainty on $c$ is this leave-one-out spread, not a formal confidence interval.
548 +
549 +\paragraph{Out-of-sample confirmation.}
550 +The law was first fitted on the 9 checkpoints up to $4\times 10^{10}$
551 +($\hat c = -0.48618$, $\hat d = -1.7226$) \emph{before} the $10^{12}$ campaign, giving
552 +the prediction
553 +\[
554 +\rho\lnx(10^{12}) \;=\; \hat c + \hat d / \ln 10^{12} \;=\; -0.549 \pm 0.005 .
555 +\]
556 +The subsequent measurement returned $-0.54756$. The previously conjectured static
557 +window $\rho\lnx \in [-0.65, -0.55]$ was exited between $4\times 10^{11}$ and
558 +$10^{12}$, as the law predicted (exit at $\lambda = \hat d/(-0.55 - \hat c)
559 +\Rightarrow x \approx 5\times 10^{11}$).
560 +
561 +\paragraph{Second confirmation at $10^{13}$.}
562 +The 13-point refit after the $10^{12}$ campaign predicted
563 +$\rho\lnx(10^{13}) = -0.5432 \pm 0.004$ and, for the lag-2 statistic,
564 +$\rho_2\lnx(10^{13}) = -0.2545 \pm 0.004$. The subsequent $10^{13}$ campaign
565 +(Section~\ref{sec:compute}) measured
566 +\[
567 +\rho\lnx(10^{13}) = -0.54264, \qquad \rho_2\lnx(10^{13}) = -0.25247,
568 +\]
569 +confirming both. Current out-of-sample predictions from the 16-point fit:
570 +\begin{equation}
571 +\rho\lnx(10^{14}) = -0.5387 \pm 0.004, \qquad
572 +\rho\lnx(10^{15}) = -0.5350 \pm 0.004, \qquad
573 +\rho_2\lnx(10^{14}) = -0.2553 \pm 0.004 .
574 +\end{equation}
575 +
576 +\paragraph{Lag 2.}
577 +The same fit form on $\rho_2$ gives
578 +\begin{equation}
579 +c_2 = -0.2806, \qquad d_2 = +0.782,
580 +\end{equation}
581 +with the drift in the \emph{opposite} direction (lag-1 relaxes upward toward its
582 +constant, lag-2 downward), and the prediction $\rho_2\lnx(10^{13}) = -0.2545 \pm
583 +0.004$ (\statuslbl{conjectured}; $\rho_2 < 0$ \statuslbl{verified up to $10^{12}$}).
584 +
585 +\paragraph{Robustness checks.}
586 +(i) \emph{Out-of-sample record}: the law's only two genuine forecasts to date both hit ---
587 +\begin{center}
588 +\begin{tabular}{@{}lccc@{}}
589 +\toprule
590 +target & fit data used & predicted & measured \\
591 +\midrule
592 +$\rho\lnx(10^{12})$ & 9 pts $\le 4\times 10^{10}$ & $-0.5485 \pm 0.005$ & $-0.54756$ \\
593 +$\rho\lnx(10^{13})$ & 13 pts $\le 10^{12}$ & $-0.5432 \pm 0.004$ & $-0.54264$ \\
594 +$\rho_2\lnx(10^{13})$ & 13 pts $\le 10^{12}$ & $-0.2545 \pm 0.004$ & $-0.25247$ \\
595 +\bottomrule
596 +\end{tabular}
597 +\end{center}
598 +(ii) \emph{Split-sample stability}: fitting \eqref{eq:c4law} on the 8 low checkpoints
599 +($10^8$--$2\times 10^{10}$) gives $c = -0.4890$, on the 8 high checkpoints
600 +($4\times 10^{10}$--$10^{13}$) gives $c = -0.4830$ --- consistent within the quoted
601 +uncertainty. (iii) \emph{Functional-form sensitivity}: adding a $1/\lambda^2$ term
602 +(three-parameter fit) leaves the in-range residuals unchanged (max $0.0025$) but moves
603 +the extrapolated constant to $c = -0.4749$: with only $\lambda \in [18.4, 29.9]$
604 +observed, the \emph{asymptotic} constant is determined only at the $\pm 0.01$ level,
605 +even though in-range and near-range predictions are tightly pinned. We therefore quote
606 +$c = -0.4835 \pm 0.002$ (statistical, fixed form) with a systematic form-uncertainty
607 +of order $0.01$. (iv) \emph{Dependence caveat}: the 16 checkpoints are nested prefixes
608 +of one dataset, hence strongly positively correlated; ordinary least squares is used
609 +only as a summarizing device, and the confirmations that matter are the genuinely
610 +out-of-sample ones in (i).
611 +
612 +We have not located these constants in the literature. The qualitative anticorrelation
613 +of consecutive gaps is known empirically, and $c, c_2$ may be derivable from existing
614 +conditional results (e.g.\ pair-correlation heuristics); we make no novelty claim
615 +pending a proper literature pass.
616 +
617 +\section{Result IV: the first-order Hardy--Littlewood model fails quantitatively}
618 +\label{sec:model}
619 +
620 +The natural first-order model for the joint law of two consecutive gaps at scale
621 +$\lambda$ treats the triple $\{p,\, p + g_1,\, p + g_1 + g_2\}$ as an HL pattern and
622 +ignores the requirement that the \emph{interior} of each gap be prime-free:
623 +\begin{equation}
624 +\Pr(g_1, g_2) \;\propto\; \Sing\bigl(\{0,\, g_1,\, g_1 + g_2\}\bigr)\,
625 +e^{-(g_1 + g_2)/\lambda},
626 +\qquad g_1, g_2 \in 2\mathbb{N}.
627 +\label{eq:triple-model}
628 +\end{equation}
629 +Using \eqref{eq:HL} with $k = 3$ and writing $\nu_p = \nu_p(\{0, g_1, g_1{+}g_2\})$:
630 +$\nu_p = 1$ iff $p \mid g_1$ and $p \mid g_2$; $\nu_p = 2$ iff exactly one of
631 +$p \mid g_1$, $p \mid g_2$, $p \mid g_1{+}g_2$ holds; else $\nu_p = 3$. For
632 +$p > g_1 + g_2$ always $\nu_p = 3$, so the factor
633 +$(1 - 3/p)/(1 - 1/p)^3 = 1 + O(1/p^2)$ is common to all pairs and cancels in the
634 +normalization; the computable part is the finite product over $p \le g_1 + g_2$.
635 +If $\nu_3 = 3$ the pattern is inadmissible and the weight vanishes --- e.g.\
636 +$(g_1, g_2) \equiv (1, 1) \pmod 3$, so the pair of consecutive gaps $(4, 4)$ is
637 +\emph{forbidden}: this arithmetic coupling is the model's only source of correlation
638 +(with $\Sing \equiv 1$ it factorizes and $\rho = 0$).
639 +
640 +\begin{result}[\statuslbl{model rejected}]
641 +Numerical evaluation of \eqref{eq:triple-model} (grid $g_i \le 14\lambda$, exact finite
642 +singular products) gives
643 +\begin{center}
644 +\begin{tabular}{@{}lccccccc@{}}
645 +\toprule
646 +$x$ & $10^8$ & $10^{10}$ & $10^{12}$ & $10^{15}$ & $10^{20}$ & $10^{30}$ & $10^{40}$\\
647 +\midrule
648 +$\rho^{\mathrm{model}}\lambda$ & $-0.1591$ & $-0.1581$ & $-0.1568$ & $-0.1550$ &
649 +$-0.1536$ & $-0.1532$ & $-0.1539$ \\
650 +\bottomrule
651 +\end{tabular}
652 +\end{center}
653 +The model reproduces the sign and the $c + d/\lambda$ drift form, but its plateau
654 +($\approx -0.154$) is only $\approx 32\%$ of the observed $c = -0.4845$. The dominant
655 +share of the anticorrelation must therefore come from the interior-compositeness
656 +(inclusion--exclusion) constraints that \eqref{eq:triple-model} ignores. Constructing
657 +the corrected model --- in the spirit of the Odlyzko--Rubinstein--Wolf treatment of
658 +jumping champions --- is the main open problem left by this cycle.
659 +\end{result}
660 +
661 +\section{Secondary observations}
662 +\label{sec:secondary}
663 +
664 +\paragraph{Jumping champion.}
665 +$\operatorname{argmax}_g N(g, x) = 6$ at every checkpoint from $10^6$ to $10^{12}$
666 +(\statuslbl{verified}), consistent with the Odlyzko--Rubinstein--Wolf conjecture that 6
667 +is the champion from $\approx 947$ to $\approx 1.7\times 10^{35}$.
668 +
669 +\paragraph{Local dominance of multiples of 6.}
670 +By \eqref{eq:singular-pair}, $3 \mid g$ multiplies $C(g)$ by
671 +$(3-1)/(3-2) = 2$, so multiples of 6 should be strict local maxima of
672 +$g \mapsto N(g, x)$:
673 +\begin{equation}
674 +G_6(x) \;=\; \max\bigl\{ G : \forall\, 6 \mid g \le G,\
675 +N(g,x) > \max(N(g{-}2,x),\, N(g{+}2,x)) \bigr\}
676 +\end{equation}
677 +grows through $66\ (10^8)$, $216\ (10^9)$, $312\ (10^{10})$, $480\ (10^{12})$
678 +(\statuslbl{verified}; the growth is sample-size-limited, as HL predicts).
679 +
680 +\paragraph{CSG maxima.}
681 +$\max_{p \le 4\times 10^9} \mathrm{CSG} = 0.73947$ (gap 210 after 20{,}831{,}323);
682 +$\max_{p \le 10^{13}} \mathrm{CSG} = 0.79754$ (gap 652 after 2{,}614{,}941{,}710{,}599) ---
683 +all agree with the published record tables.
684 +
685 +\paragraph{Reproduction of published tables.}
686 +Every campaign reproduces the known tables in its range (\statuslbl{verified}): all 53
687 +maximal gaps below $10^{13}$ (OEIS A005250/A002386; 53 below $10^{13}$, ending with gap 674 after
688 +7{,}177{,}162{,}611{,}713), twin-pair counts $N(2, 10^{6,7,8}) = 8169,\ 58980,\ 440312$,
689 +and $\pi(10^k)$, $k \le 8$.
690 +
691 +\section{Threats to validity}
692 +\label{sec:threats}
693 +
694 +We list what could invalidate each headline claim, and why we believe it does not.
695 +\begin{enumerate}
696 +\item \emph{Software error in the scan.} Mitigations: the 37-check validation gate;
697 + exact reproduction of independent published tables in every campaign (53 maximal
698 + gaps, twin counts, $\pi(10^k)$, CSG maxima); bit-identical cross-machine and
699 + chunked-vs-monolithic agreement (Proposition~\ref{prop:exact}); two independent
700 + sieve implementations (NumPy CPU and Metal GPU) forced to byte-identical outputs.
701 + A bug would have to affect all of these coherently.
702 +\item \emph{Floating-point artefacts in $\rho$.} Excluded by
703 + Proposition~\ref{prop:exact}: all accumulators are exact integers; rounding
704 + enters only in the final division, at relative level $\sim 10^{-16}$.
705 +\item \emph{Selection effects / overfitting in C4$'$.} The law has two free parameters
706 + and sixteen highly correlated points; the defense is not the fit quality but the
707 + two pre-registered out-of-sample confirmations (Section~\ref{sec:c4},
708 + robustness item (i)) and the split-sample stability. The asymptotic constant
709 + carries an explicit form-systematic of $\pm 0.01$.
710 +\item \emph{Certificate validity for the desert.} Endpoints rely on the
711 + Sorenson--Webster bound for the 12-base Miller--Rabin test; both endpoints are
712 + $\approx 10^{24} < 3.317\times 10^{24}$. Interior certificates are elementary
713 + (explicit divisors) or strong witnesses \eqref{eq:mr}, checked by an independent
714 + implementation. No probabilistic step remains.
715 +\item \emph{Literature novelty.} We claim novelty for nothing; phenomena analogous to
716 + C4$'$/C6 may exist in the pair-correlation literature. All ``known value''
717 + comparisons were fetched from OEIS/t5k/Nicely at run time.
718 +\end{enumerate}
719 +
720 +\section{Instructive failures}
721 +\label{sec:lessons}
722 +
723 +Recorded because they shaped the protocol:
724 +\begin{enumerate}
725 + \item \emph{Checkpoint-only verification is a trap}: C2 held at every decade
726 + checkpoint while failing $\sim 1.4\times 10^8$ times in between
727 + (Result~\ref{res:race}).
728 + \item \emph{Never transcribe known values from memory}: a hand-written ``known
729 + table'' briefly contradicted a correct scan; fetching OEIS resolved it in the
730 + scan's favour.
731 + \item \emph{Fleet heterogeneity}: a Python-3.10-only type annotation crashed 9 of 12
732 + nodes running Python 3.9 (fixed by lazy annotation evaluation); five further
733 + nodes had no usable interpreter at all. Target the lowest interpreter in the
734 + fleet and import-check per node class before launching.
735 + \item \emph{Naive models fail beyond their remit}: the single-gap model
736 + $N(g,x) \approx C(g) \int_2^x e^{-g/\ln t} (\ln t)^{-2} dt$ shows structured
737 + residuals (deficits at $g = 36, 72, 100, 108$ relative to neighbours) --- do
738 + not conjecture on top of it without the inclusion--exclusion correction.
739 +\end{enumerate}
740 +
741 +\section{State of the art (for honesty of any claim)}
742 +\label{sec:sota}
743 +
744 +Published records against which any of our outputs must be measured: largest known
745 +maximal gap $1854$ after $p = 101412319996363309069$ (2026), with exhaustive
746 +maximal-gap verification to $\approx 2\times 10^{19}$; largest known gap between
747 +probable primes $16{,}045{,}848$ (2024); record merit $41.938\ldots$ (Gapcoin, 2017);
748 +record CSG ratio $0.9206\ldots$ (Nyman, 1999). Ford--Green--Konyagin--Maynard--Tao
749 +\cite{FGKMT} give the best unconditional lower bound
750 +\begin{equation}
751 +\max_{p_{n+1} \le x} g_n \;\gg\;
752 +\frac{\ln x \,\ln\ln x \,\ln\ln\ln\ln x}{\ln\ln\ln x},
753 +\end{equation}
754 +while Baker--Harman--Pintz give $g_n \ll p_n^{0.525}$ unconditionally. None of our
755 +constructions approaches the computational records; none is claimed to.
756 +
757 +\section{Reproducibility}
758 +
759 +Everything is re-runnable end to end: \texttt{src/test\_core.py} (validation gate, 37
760 +checks), \texttt{src/cycle2\_c4.py} (single-machine scans),
761 +\texttt{src/worker\_c4.py} $+$ \texttt{src/merge\_c4.py} (distributed scans with
762 +built-in cross-validation against prior campaigns), \texttt{src/gpu\_sieve.py} (Metal
763 +kernel with its correctness gate), \texttt{src/model\_c4.py} (the HL triple model),
764 +\texttt{certs/verify\_desert.py} and \texttt{certs/verify\_c2\_refutation.py}
765 +(standalone verifiers, stdlib-only where feasible). All scans are deterministic and
766 +seed-free; all raw checkpoint data are kept as JSON under \texttt{/data}; the research
767 +journal (\texttt{journal.md}) records every cycle, decision, and failure.
768 +
769 +\paragraph{Data and code availability.}
770 +The complete repository (source, certificates, raw checkpoint JSONs, journal, and this
771 +manuscript) is available at \url{https://github.com/spboucher-ai/prime-mystery-engine}.
772 +Re-verification of the two headline certificates requires only a stock Python~3:
773 +\texttt{python3 certs/verify\_desert.py} and
774 +\texttt{python3 certs/verify\_c2\_refutation.py}.
775 +
776 +\begin{thebibliography}{9}
777 +\bibitem{FGKMT} K.~Ford, B.~Green, S.~Konyagin, J.~Maynard, T.~Tao,
778 +\emph{Long gaps between primes}, J.~Amer.~Math.~Soc.~31 (2018), 65--105.
779 +\bibitem{ORW} A.~Odlyzko, M.~Rubinstein, M.~Wolf,
780 +\emph{Jumping champions}, Experiment.~Math.~8 (1999), 107--118.
781 +\bibitem{SW} J.~Sorenson, J.~Webster,
782 +\emph{Strong pseudoprimes to twelve prime bases}, Math.~Comp.~86 (2017), 985--1003.
783 +\bibitem{BPSW} R.~Baillie, S.~S.~Wagstaff,
784 +\emph{Lucas pseudoprimes}, Math.~Comp.~35 (1980), 1391--1417.
785 +\bibitem{RS} M.~Rubinstein, P.~Sarnak, \emph{Chebyshev's bias},
786 +Experiment.~Math.~3 (1994), 173--197.
787 +\bibitem{Granville} A.~Granville, \emph{Harald Cram\'er and the distribution of prime
788 +numbers}, Scand.~Actuar.~J.~1 (1995), 12--28.
789 +\bibitem{A005250} OEIS Foundation, \emph{Sequences A005250, A002386 (record gaps
790 +between primes)}, \url{https://oeis.org}.
791 +\bibitem{t5k} \emph{Table of known maximal gaps between primes},
792 +\url{https://t5k.org/notes/GapsTable.html}.
793 +\bibitem{Nicely} T.~R.~Nicely, \emph{First occurrence prime gaps},
794 +\url{https://faculty.lynchburg.edu/~nicely/gaps/gaplist.html}.
795 +\end{thebibliography}
796 +
797 +\end{document}
added records.md +37 −0
@@ -0,0 +1,37 @@
1 +<!--
2 +records.md — Prime Mystery Engine: state of the art (Phase 2)
3 +Author: Simon-Pierre Boucher — contact@spboucher.ai
4 +Last updated: 2026-08-06
5 +-->
6 +
7 +# State of the Art — Axis #10: Prime Deserts (large prime gaps)
8 +
9 +## What is known (checked against the literature, 2026-08-06)
10 +
11 +| Quantity | Known record | Source |
12 +|---|---|---|
13 +| Largest **maximal** prime gap | 1854, after p = 101412319996363309069 (85th maximal gap, Robert Smith / Brian Kehrig code, 2026) | primerecords.dk, t5k.org GapsTable |
14 +| Exhaustive verification bound for maximal gaps | ≈ 2×10^19 (beyond 2^64) | t5k.org/notes/GapsTable.html |
15 +| Largest known gap (any size, PRP endpoints) | 16,045,848 near a 385,713-digit number (A. Höglund, 2024) | primerecords.dk/primegaps/megagap3.htm |
16 +| Highest known **merit** g/ln(p) | 41.93878… (gap 8350 after an 87-digit prime, Gapcoin network, 2017) | primegap-list-project |
17 +| Highest Cramér–Shanks–Granville ratio g/ln²(p) | 0.9206… (Nyman's gap 1132 after 1693182318746371, 1999) | MathWorld Prime Gaps |
18 +| Theory (lower bound) | Ford–Green–Konyagin–Maynard–Tao 2016: max gap ≥ c·(ln x · lnln x · lnlnlnln x)/lnlnln x | published |
19 +| Theory (upper bound) | Baker–Harman–Pintz: gap = O(x^0.525); RH gives O(√x ln x) | published |
20 +| Heuristic | Cramér: limsup g/ln²p = 1; Granville correction: ≥ 2e^{-γ} ≈ 1.1229 | published |
21 +| Jumping champion | 6 is the most common gap from ≈ 947 up to ≈ 1.7×10^35 (conjectured, verified in ranges); then 30 | published (Odlyzko–Rubinstein–Wolf) |
22 +
23 +## What would count as NEW here
24 +
25 +1. **A new maximal gap** — requires exhaustive sieving beyond 2×10^19: **out of reach** locally; not the goal.
26 +2. **A gap with merit > 41.94** — requires massive targeted search (Gapcoin-scale): out of reach in one cycle; a merit > 20 gap found by our own sieve pipeline would still only be "personal best" territory (thousands of merit-20+ gaps are catalogued). Any gap we find must be checked against the primegap-list-project tables before any claim.
27 +3. **Structural conjectures about gap statistics** (residues of maximal-gap primes, corrections to Cramér in arithmetic progressions, gap-pair correlations) — realistic: findings are likely "(a) probably known", but a precise, massively tested statement with clean data is a legitimate output.
28 +4. **Certified desert constructions** (covering systems): any construction is reproducible and certifiable, but primorial/covering constructions have low merit (~2–4) — value is the **certificate pipeline**, not records.
29 +
30 +## Exact thresholds to beat (for honesty in any claim)
31 +
32 +- Maximal gap: must exceed 1854 AND be proven maximal (impossible without exhaustive sieve > 2×10^19) — or simply reproduce/verify the known table below our compute bound (verification, not record).
33 +- Merit record: > 41.93878.
34 +- CSG ratio record: > 0.92064.
35 +- First-occurrence gaps: the Nicely / primegap-list-project tables catalogue first known occurrences for every even gap; before claiming a "first occurrence", check those tables.
36 +
37 +Sources: [t5k.org GapsTable](https://t5k.org/notes/GapsTable.html), [primerecords.dk top-20](http://primerecords.dk/primegaps/gaps20.htm), [primerecords.dk megagap](http://primerecords.dk/primegaps/megagap3.htm), [primegap-list-project merit record](https://primegap-list-project.github.io/record/2017/12/31/new-prime-gap-of-maximum-merit/), [MathWorld Prime Gaps](https://mathworld.wolfram.com/PrimeGaps.html), [Wikipedia Prime gap](https://en.wikipedia.org/wiki/Prime_gap), [Nicely first-occurrence tables](https://faculty.lynchburg.edu/~nicely/gaps/gaplist.html)
added report_cycle1.md +65 −0
@@ -0,0 +1,65 @@
1 +<!--
2 +report_cycle1.md — Prime Mystery Engine: Cycle 1 synthesis (Phase 7)
3 +Author: Simon-Pierre Boucher — contact@spboucher.ai
4 +Date: 2026-08-06
5 +-->
6 +
7 +# Cycle 1 Report — Prime Deserts (axis #10)
8 +
9 +## Results
10 +
11 +1. **REFUTATION (main result).** Conjecture C2 — "N(2,x) > N(4,x) for all x ≥ 10^6" — is **false**.
12 + Minimal counterexample: the count of gap-4 pairs ties the count of twin pairs at end-prime
13 + **80966861** and strictly overtakes it at **80966933**. Beyond that, the race changes leader
14 + endlessly up to 4×10^9 (last lead-change at 3999999979, i.e. still swinging at the scan boundary;
15 + difference −2270 at 4×10^9). Epistemic status: REFUTED, re-verifiable in ~30 s with a
16 + stdlib-only script. Consistent with the random-walk heuristic (equal singular series) — the
17 + right follow-up question is the logarithmic density of each leader (Chebyshev-bias analogue).
18 +
19 +2. **CERTIFIED CONSTRUCTION.** A prime desert of length **260** at 25 digits:
20 + N = 1116336781708038449369693 and N+260 are consecutive primes (merit 4.6955).
21 + Method: hybrid covering system mod primes ≤ 59 (243 positions composite for *every* CRT shift —
22 + proven; 16 holes certified by explicit factors/MR witnesses; endpoints by deterministic
23 + Miller–Rabin, below the 3.317×10^24 validity bound). ×4.3 over the classic primorial baseline.
24 + **Not a record** (records.md) — the certified, independently re-verifiable pipeline is the point.
25 +
26 +3. **INDEPENDENT VERIFICATIONS of known tables** (validation of the whole engine):
27 + all 32 maximal gaps below 4×10^9, twin counts at 10^6/10^7/10^8, π(10^k) k ≤ 8,
28 + max CSG ratio 0.73947 below 4×10^9 — all equal to published values.
29 +
30 +## Conjecture status table
31 +
32 +| # | Statement (short) | Status |
33 +|---|---|---|
34 +| C1 | Jumping champion = 6 | SURVIVOR — VERIFIED UP TO 4×10^9 (known, ORW) |
35 +| C2 | N(2,x) > N(4,x) for x ≥ 10^6 | **REFUTED** — min. counterexample 80966861/80966933 |
36 +| C3 | Multiples of 6 local maxima; G6(x) ≥ 66, → ∞ | SURVIVOR, strengthened (G6 = 216 at 10^9..4×10^9) |
37 +| C4 | ρ(x) < 0, ρ·ln x ∈ [−0.65, −0.55] | SURVIVOR with revision (drift → −0.565; limit clause retracted) |
38 +| C5 | max CSG below 4×10^9 = 0.7395 | VERIFIED (matches published table; not new) |
39 +
40 +## Instructive failures
41 +
42 +- **Checkpoint-only verification is a trap**: C2 "held" at every decade checkpoint while failing
43 + ~137 million times in between. Full-resolution (every-event) adversary scans are now mandatory
44 + for race-type conjectures.
45 +- A first analysis run had segment boundaries misaligned with checkpoints (caught because twin
46 + counts disagreed with literature values — the validate-against-known-values rule paid off).
47 +- The naive HL gap model exp(−g/ln t) is inadequate beyond g ≈ 40 (structured residuals at
48 + g = 36, 72, 100, 108): do not conjecture on it before building the inclusion–exclusion model.
49 +
50 +## Files and re-verification
51 +
52 +- `src/core.py` + `src/test_core.py` — primitives; run `python3 src/test_core.py` (37/37 PASS, 0.3 s).
53 +- `src/explore_gaps.py`, `src/analyze_gaps.py`, `src/adversary_race.py` — deterministic scans;
54 + outputs in `/data` (`adversary_4e9.json`, `gapstats_1e8.*`, `maximal_gaps_1e8.csv`, …).
55 +- `certs/desert_certificate.json` + `certs/verify_desert.py` — desert; run `python3 certs/verify_desert.py`.
56 +- `certs/verify_c2_refutation.py` — refutation; run `python3 certs/verify_c2_refutation.py`.
57 +- `src/prover_desert.py` (attempt 1), `src/prover_desert2.py` (attempt 2, current certificate).
58 +
59 +## Next directions (ranked)
60 +
61 +1. **C4 at 10^10** (cluster job): does ρ·ln x stabilize? Best quantitative lead — structural
62 + (absent from Cramér model), cheap to test, precise falsifiable target c ∈ [−0.60, −0.50].
63 +2. C2 follow-up: logarithmic density of the gap-2 lead (Chebyshev-bias analogue) — needs the
64 + full crossing record, one scan with lead-time accounting.
65 +3. Proper inclusion–exclusion HL model for N(g,x), then revisit the g = 36/72/100/108 deficits.
added report_cycle2.md +67 −0
@@ -0,0 +1,67 @@
1 +<!--
2 +report_cycle2.md — Prime Mystery Engine: Cycle 2 synthesis (Phase 7)
3 +Author: Simon-Pierre Boucher — contact@spboucher.ai
4 +Date: 2026-08-06
5 +-->
6 +
7 +# Cycle 2 Report — C4 escalation to 4×10^10 (dual-machine)
8 +
9 +## Results
10 +
11 +1. **C4 refined into C4′ (main result).** The consecutive-gap anticorrelation obeys, empirically,
12 + **ρ(x)·ln x = c + d/ln x** with **c = −0.486 ± 0.010, d = −1.72** (9 checkpoints 10^8…4×10^10,
13 + max residual 0.0025, leave-one-out spread on c: [−0.494, −0.482]). Falsifiable prediction for
14 + cycle 3: **ρ·ln x = −0.549 ± 0.005 at x = 10^12**; the cycle-1 window [−0.65, −0.55] should be
15 + exited near x ≈ 5×10^11. Epistemic status: CONJECTURED (consistent up to 4×10^10).
16 +
17 +2. **First-order HL model rejected (PROVER).** The triple-correlation model
18 + P(g₁,g₂) ∝ S({0,g₁,g₁+g₂})·e^{−(g₁+g₂)/λ} reproduces the sign and the c + d/λ drift form but
19 + gives ρ·λ ≈ −0.154, a factor ~3.2 too small. The anticorrelation is therefore dominated by
20 + interior-compositeness (inclusion–exclusion) effects, not the bare singular series. Open problem.
21 +
22 +3. **Dual-machine reproducibility.** Laptop (10^10) and M3U96a (4×10^10) produced bit-for-bit
23 + identical values at every shared checkpoint — deterministic pipeline validated on independent
24 + hardware.
25 +
26 +4. **New conjecture C6.** Lag-2 gaps also anticorrelate: ρ₂·ln x → c₂ ≈ −0.29 ± 0.03, drifting in the
27 + opposite direction to lag-1. Low-confidence fit; to be firmed up.
28 +
29 +5. **Free continuations.** Champion = 6 (C1) and G6 ≥ 66 (C3) hold to 4×10^10 (G6 is non-monotone:
30 + 312 → 282 → 276 — tail noise). C2 race keeps changing leader (D = −804 at 4×10^10). All 37
31 + maximal gaps below 4×10^10 match OEIS A002386 (web-checked); max CSG = 0.79535 (456 after
32 + 25056082087, known).
33 +
34 +## Instructive failure
35 +
36 +A hand-written "known table" check briefly claimed our maximal gaps disagreed with the literature;
37 +fetching OEIS A002386 showed the scan was right and the from-memory list wrong. Rule reinforced:
38 +**never transcribe "known" values from memory — always fetch the source.**
39 +
40 +## Conjecture table (cumulative)
41 +
42 +| # | Statement (short) | Status |
43 +|---|---|---|
44 +| C1 | Jumping champion = 6 | VERIFIED UP TO 4×10^10 |
45 +| C2 | N(2,x) > N(4,x) for x ≥ 10^6 | REFUTED (cycle 1); leader still swapping at 4×10^10 |
46 +| C3 | G6(x) ≥ 66 | VERIFIED UP TO 4×10^10 |
47 +| C4 | ρ·ln x ∈ [−0.65, −0.55] | VERIFIED UP TO 4×10^10, predicted to FAIL ≈ 5×10^11 — superseded |
48 +| C4′ | ρ·ln x = c + d/ln x, c = −0.486(10), d ≈ −1.72 | CONJECTURED; falsifiable at 10^12 |
49 +| C5 | Maximal-gap table reproduction | VERIFIED UP TO 4×10^10 (37/37 = A002386) |
50 +| C6 | ρ₂·ln x → c₂ ≈ −0.29 | CONJECTURED (new, low confidence) |
51 +
52 +## Files and re-verification
53 +
54 +- `src/cycle2_c4.py` — deterministic scan; sanity: `python3 src/cycle2_c4.py 1e8 sanity`
55 + must reproduce cycle-1 checkpoints exactly.
56 +- `data/cycle2_c4_1e10_laptop.json`, `data/cycle2_c4_4e10_M3U96a.json` — raw results (identical on
57 + shared checkpoints; diff them to re-verify cross-machine agreement).
58 +- `src/fit_c4.py` → `data/fit_c4.json` — the C4′ fit; `src/model_c4.py` → `data/model_c4.json` — the
59 + rejected first-order model.
60 +
61 +## Next precise action (one)
62 +
63 +Run the C4′ test at 10^12 distributed core-proportionally across the MacLustr cluster
64 +(mergeable streaming sums make the scan embarrassingly parallel; ~284 cores ⇒ minutes),
65 +to accept/refute **ρ·ln x(10^12) = −0.549 ± 0.005**.
66 +
67 +Sources: [OEIS A002386](https://oeis.org/A002386), [t5k.org GapsTable](https://t5k.org/notes/GapsTable.html)
added report_cycle3.md +66 −0
@@ -0,0 +1,66 @@
1 +<!--
2 +report_cycle3.md — Prime Mystery Engine: Cycle 3 synthesis (Phase 7)
3 +Author: Simon-Pierre Boucher — contact@spboucher.ai
4 +Date: 2026-08-06
5 +-->
6 +
7 +# Cycle 3 Report — C4′ tested at 10^12 (distributed, 12 nodes / 162 cores)
8 +
9 +## Headline
10 +
11 +**C4′'s prediction was CONFIRMED.** Cycle 2 predicted ρ(x)·ln x = −0.549 ± 0.005 at x = 10^12;
12 +the distributed scan measured **−0.54756**. The superseded cycle-1 window [−0.65, −0.55] was exited
13 +between 4×10^11 and 10^12, matching the predicted exit ≈ 5×10^11. This is the engine's first
14 +confirmed quantitative prediction.
15 +
16 +## Results
17 +
18 +1. **Refined law.** 13-point fit (10^8…10^12): ρ·ln x = c + d/ln x with
19 + **c = −0.4845 ± 0.003 (LOO), d = −1.757**, max residual 0.0025.
20 + New falsifiable predictions: **ρ·ln x(10^13) = −0.5432 ± 0.004**, ρ·ln x(10^14) = −0.5390 ± 0.004.
21 +2. **C6 firmed up.** Lag-2: c₂ = −0.2806, d₂ = +0.782 (drift sign opposite to lag-1);
22 + prediction ρ₂·ln x(10^13) = −0.2545 ± 0.004.
23 +3. **Pipeline validation.** The 1000-chunk distributed merge reproduces the cycle-2 single-machine
24 + checkpoints exactly (10^10, 2×10^10, 4×10^10) and all 48 maximal gaps below 10^12 equal the
25 + published table (tail spot-checked: 540 after 738832927927).
26 +4. **Free continuations.** Champion = 6 up to 10^12 (C1). G6 grows to 480 (C3). The C2 twin-vs-quad
27 + race is still changing leader at 10^12: D = −238 out of ~1.5×10^10 gap events.
28 +5. **Cost.** 17,719 core-seconds (≈ 4.9 core-hours), ≈ 5 minutes wall across the fleet.
29 +
30 +## Instructive failures (logged for future cycles)
31 +
32 +- 9/12 nodes crashed on launch: a Python-3.10-only annotation (`np.ndarray | None`) met Apple's
33 + Python 3.9. Fixed with `from __future__ import annotations`; rule: target the fleet's LOWEST
34 + interpreter and import-check per node class before launching.
35 +- 5 nodes had no usable python3 at all (missing Xcode CLT), 1 was unreachable — the theoretical
36 + 228-core fleet was really 162 cores.
37 +- zsh word-splitting and modifier gotchas in orchestration (documented in journal).
38 +
39 +## Conjecture table (cumulative)
40 +
41 +| # | Statement (short) | Status |
42 +|---|---|---|
43 +| C1 | Jumping champion = 6 | VERIFIED UP TO 10^12 |
44 +| C2 | N(2,x) > N(4,x), x ≥ 10^6 | REFUTED (cycle 1); race still swinging at 10^12 (D = −238) |
45 +| C3 | G6(x) ≥ 66 | VERIFIED UP TO 10^12 (G6 = 480 and growing) |
46 +| C4 | ρ·ln x ∈ [−0.65, −0.55] | REFUTED-as-predicted (exited between 4×10^11 and 10^12) |
47 +| C4′ | ρ·ln x = c + d/ln x, c = −0.4845(30) | **PREDICTION CONFIRMED at 10^12**; next test 10^13 |
48 +| C5 | Maximal-gap table reproduction | VERIFIED UP TO 10^12 (48/48 = published table) |
49 +| C6 | ρ₂·ln x → c₂ ≈ −0.28 | CONJECTURED, firmed (13-point fit); prediction at 10^13 stated |
50 +
51 +## Files and re-verification
52 +
53 +- `src/worker_c4.py`, `src/merge_c4.py` — distributed chunk worker + merger. Re-verify the pipeline:
54 + run any subset of chunks locally and merge; the built-in cross-validation block compares against
55 + `data/cycle2_c4_4e10_M3U96a.json`.
56 +- `data/cycle3_c4_1e12.json` — merged result (checkpoints, maximal gaps, cross-val report).
57 +- Fits reproducible from the two checkpoint JSONs (script in journal / `src/fit_c4.py` pattern).
58 +
59 +## Next precise action (one)
60 +
61 +Cycle 4: test ρ·ln x(10^13) = −0.5432 ± 0.004 on the **3-node fleet only (M3U96a, M4M64a, M4M64b —
62 +user directive), saturated, with a Metal/GPU sieve path** (benchmark MLX/Metal kernel vs numpy first;
63 +10^13 is ~10× cycle 3's work, ~50 CPU core-hours, so GPU acceleration is the enabler).
64 +
65 +Sources: [OEIS A005250](https://oeis.org/A005250), [OEIS A002386](https://oeis.org/A002386),
66 +[t5k.org GapsTable](https://t5k.org/notes/GapsTable.html)
added report_cycle4.md +58 −0
@@ -0,0 +1,58 @@
1 +<!--
2 +report_cycle4.md — Prime Mystery Engine: Cycle 4 synthesis (Phase 7)
3 +Author: Simon-Pierre Boucher — contact@spboucher.ai
4 +Date: 2026-08-06
5 +-->
6 +
7 +# Cycle 4 Report — C4′/C6 at 10^13, CPU + Metal GPU on 3 nodes
8 +
9 +## Headline
10 +
11 +**Both out-of-sample predictions confirmed at 10^13.**
12 +ρ·ln x(10^13) = **−0.54264** (predicted −0.5432 ± 0.004 — C4′'s second consecutive hit);
13 +ρ₂·ln x(10^13) = **−0.25247** (predicted −0.2545 ± 0.004 — C6's first hit).
14 +The scaling law ρ(x)·ln x = c + d/ln x now stands twice-tested; 16-point refit:
15 +**c = −0.48354 (LOO ± 0.002), d = −1.7772** (max residual 0.0026).
16 +
17 +## The compute
18 +
19 +- Fleet (user directive): M3U96a + M4M64a + M4M64b only — 58 CPU cores saturated + **3 GPUs via a
20 + custom Metal kernel** (MLX `fast.metal_kernel`), which passed byte-identity gates against the CPU
21 + sieve before use and processed **1170 of 10,000 chunks (11.7%)**.
22 +- 10,000 chunks of 10^9, measured-rate quotas, ~50 min wall, 147,738 core-seconds (~41 core-hours,
23 + incl. ~4% duplicated tail work from a late rebalance; duplicates byte-identical, tiling asserted).
24 +- **Validation**: 10,000 unique chunks tile [2, 10^13) exactly; all 10 overlapping checkpoints
25 + bit-identical to cycles 2/3; all 53 maximal gaps below 10^13 equal the published table
26 + (new in range: 582, 588, 602, 652, 674 — web-checked against t5k/Nicely).
27 +
28 +## Conjecture table (cumulative)
29 +
30 +| # | Statement | Status |
31 +|---|---|---|
32 +| C1 | Champion = 6 | VERIFIED UP TO 10^13 |
33 +| C2 | N(2,x) > N(4,x) | REFUTED (cycle 1); race still swinging at 10^13 (D = +8870) |
34 +| C3 | G6 ≥ 66 | VERIFIED UP TO 10^13 (G6 = 576) |
35 +| C4′ | ρ·ln x = c + d/ln x | **CONFIRMED at 10^12 AND 10^13**; c = −0.48354 ± 0.002 |
36 +| C5 | Maximal-gap table | VERIFIED UP TO 10^13 (53/53); max CSG 0.79754 (652 after 2614941710599) |
37 +| C6 | ρ₂·ln x → c₂ | **CONFIRMED at 10^13**; c₂ = −0.27762, d₂ = +0.719 |
38 +
39 +## Predictions on record (falsifiable)
40 +
41 +- ρ·ln x(10^14) = **−0.5387 ± 0.004**; ρ·ln x(10^15) = −0.5350 ± 0.004.
42 +- ρ₂·ln x(10^14) = **−0.2553 ± 0.004**.
43 +
44 +## Files and re-verification
45 +
46 +- `src/gpu_sieve.py` — Metal kernel + its correctness/benchmark gate (`python3 src/gpu_sieve.py`).
47 +- `src/gpu_runner.py`, refactored `src/worker_c4.py` (CPU/GPU modes, identical outputs).
48 +- `data/cycle4_c4_1e13.json` — merged result with built-in cross-validation records.
49 +- Paper updated: `paper/main.tex` (12 pp., compiles clean).
50 +
51 +## Next precise action (one)
52 +
53 +Cycle 5: 10^14 (~10× the work; ≈ 8–9 h wall on the 3-node fleet — overnight run) to test
54 +ρ·ln x(10^14) = −0.5387 ± 0.004. If compute is deferred: attack the main open problem instead —
55 +the inclusion–exclusion correction to the HL triple model that should explain c ≈ −0.484.
56 +
57 +Sources: [t5k.org GapsTable](https://t5k.org/notes/GapsTable.html),
58 +[Nicely maximal gaps](https://faculty.lynchburg.edu/~nicely/gaps/gaps.html)
added src/adversary_race.py +145 −0
@@ -0,0 +1,145 @@
1 +#!/usr/bin/env python3
2 +# =============================================================================
3 +# adversary_race.py — Cycle 1 ADVERSARY: attack C1..C5 up to 4e9
4 +# Author: Simon-Pierre Boucher — contact@spboucher.ai
5 +# =============================================================================
6 +# Single deterministic segmented scan of [2, LIMIT). Attacks:
7 +# C2: FULL-RESOLUTION hunt for crossings of D(x) = N(2,x) - N(4,x):
8 +# every gap event is inspected (vectorized cumsum), every x where
9 +# D <= 0 after x >= START_TRACK is recorded (minimal counterexample).
10 +# Also records ALL sign-change epochs from the very start (context).
11 +# C1/C3/C4: histogram + correlation snapshots at decade checkpoints.
12 +# C5: complete maximal-gap record list with CSG ratios.
13 +# Output: data/adversary_<L>.json (+ console digest).
14 +# Usage: python3 adversary_race.py 4e9
15 +# =============================================================================
16 +
17 +import json
18 +import sys
19 +import time
20 +from math import log
21 +from pathlib import Path
22 +
23 +import numpy as np
24 +
25 +from core import primes_upto, sieve_segment
26 +
27 +DATA = Path(__file__).resolve().parent.parent / "data"
28 +DATA.mkdir(exist_ok=True)
29 +
30 +
31 +def pearson(S):
32 + n = S["n"]
33 + cov = S["xy"] / n - (S["x"] / n) * (S["y"] / n)
34 + vx = S["xx"] / n - (S["x"] / n) ** 2
35 + vy = S["yy"] / n - (S["y"] / n) ** 2
36 + return cov / (vx * vy) ** 0.5
37 +
38 +
39 +def main():
40 + limit = int(float(sys.argv[1])) if len(sys.argv) > 1 else 4 * 10**9
41 + seg_size = 50_000_000
42 + cps = [c for c in (10**6, 10**7, 10**8, 10**9, 2 * 10**9, 4 * 10**9) if c <= limit]
43 + bounds = sorted({limit, *cps, *range(2, limit, seg_size)} - {2})
44 + base = primes_upto(int(limit ** 0.5) + 1)
45 +
46 + hist = np.zeros(5000, dtype=np.int64)
47 + S = dict(n=0, x=0.0, y=0.0, xx=0.0, yy=0.0, xy=0.0)
48 + D = 0 # running N2 - N4
49 + last_nonpos_prime = None # last prime p (end of gap) where D <= 0
50 + n_nonpos_events = 0
51 + zero_epochs = [] # (prime, D) samples where D==0 (first 200)
52 + maximal = []
53 + best = 0
54 + checkpoints_out = {}
55 + prev = None
56 + prev_gap = None
57 + t0 = time.time()
58 + lo = 2
59 + cps_left = list(cps)
60 + for hi in bounds:
61 + primes = sieve_segment(lo, hi, base)
62 + if len(primes):
63 + if prev is not None:
64 + primes = np.concatenate(([prev], primes))
65 + if len(primes) >= 2:
66 + gaps = np.diff(primes)
67 + ends = primes[1:]
68 + hist += np.bincount(gaps, minlength=len(hist))[: len(hist)]
69 + # --- C2 full-resolution race ---
70 + delta = (gaps == 2).astype(np.int64) - (gaps == 4).astype(np.int64)
71 + run = D + np.cumsum(delta)
72 + nonpos = np.nonzero(run <= 0)[0]
73 + if len(nonpos):
74 + n_nonpos_events += len(nonpos)
75 + last_nonpos_prime = int(ends[nonpos[-1]])
76 + zeros = nonpos[run[nonpos] == 0]
77 + for i in zeros[: max(0, 200 - len(zero_epochs))]:
78 + zero_epochs.append((int(ends[i]), 0))
79 + D = int(run[-1])
80 + # --- correlations ---
81 + gl = np.concatenate(([prev_gap], gaps)) if prev_gap is not None else gaps
82 + a, b = gl[:-1].astype(np.float64), gl[1:].astype(np.float64)
83 + S["n"] += len(a)
84 + S["x"] += a.sum(); S["y"] += b.sum()
85 + S["xx"] += (a * a).sum(); S["yy"] += (b * b).sum()
86 + S["xy"] += (a * b).sum()
87 + prev_gap = int(gaps[-1])
88 + # --- C5 records ---
89 + if int(gaps.max()) > best:
90 + starts = primes[:-1]
91 + for i in np.nonzero(gaps > best)[0]:
92 + g, p = int(gaps[i]), int(starts[i])
93 + if g > best:
94 + best = g
95 + maximal.append((g, p))
96 + prev = primes[-1]
97 + while cps_left and hi >= cps_left[0]:
98 + c = cps_left.pop(0)
99 + nz = np.nonzero(hist)[0]
100 + champ = int(nz[np.argmax(hist[nz])])
101 + g6 = 0
102 + for g in range(6, int(nz.max()) - 2, 6):
103 + if hist[g] > hist[g - 2] and hist[g] > hist[g + 2]:
104 + g6 = g
105 + else:
106 + break
107 + rho = pearson(S)
108 + checkpoints_out[str(c)] = {
109 + "champion": champ,
110 + "N2": int(hist[2]), "N4": int(hist[4]),
111 + "D=N2-N4": int(hist[2] - hist[4]),
112 + "G6": g6,
113 + "rho": round(rho, 6),
114 + "rho_times_lnx": round(rho * log(c), 4),
115 + }
116 + print(f" cp {c:.0e}: champ={champ} D={hist[2]-hist[4]} G6={g6} "
117 + f"rho*lnx={rho*log(c):.4f} last_nonpos={last_nonpos_prime} "
118 + f"[{time.time()-t0:.0f}s]", flush=True)
119 + lo = hi
120 +
121 + out = {
122 + "limit": limit,
123 + "scan_seconds": round(time.time() - t0, 1),
124 + "C2_race": {
125 + "final_D": D,
126 + "n_events_D_nonpositive": int(n_nonpos_events),
127 + "last_prime_with_D_nonpositive": last_nonpos_prime,
128 + "first_zero_epochs_sample": zero_epochs[:50],
129 + "definition": "D(p) = N2 - N4 counting gaps whose END prime <= p",
130 + },
131 + "checkpoints": checkpoints_out,
132 + "maximal_gaps": [
133 + {"gap": g, "after": p, "merit": round(g / log(p), 4),
134 + "csg": round(g / log(p) ** 2, 5)} for g, p in maximal
135 + ],
136 + "max_csg": max(round(g / log(p) ** 2, 5) for g, p in maximal if p > 100),
137 + }
138 + with open(DATA / f"adversary_{limit:.0e}.json".replace("+0", "").replace("+", ""), "w") as f:
139 + json.dump(out, f, indent=2)
140 + print(json.dumps(out["C2_race"], indent=2))
141 + print("max CSG:", out["max_csg"], "| largest gap:", maximal[-1])
142 +
143 +
144 +if __name__ == "__main__":
145 + main()
added src/analyze_gaps.py +166 −0
@@ -0,0 +1,166 @@
1 +#!/usr/bin/env python3
2 +# =============================================================================
3 +# analyze_gaps.py — Cycle 1 EXPLORER/ADVERSARY: gap statistics vs Hardy-Littlewood
4 +# Author: Simon-Pierre Boucher — contact@spboucher.ai
5 +# =============================================================================
6 +# One segmented scan up to LIMIT with checkpoints. Produces (in /data):
7 +# gapstats_<L>.npz : per-checkpoint gap histograms, consecutive-gap pair
8 +# matrix (g_n, g_{n+1}) for gaps < 400, correlation sums
9 +# gapstats_<L>.json: HL-normalized ratios, champion tables, correlations
10 +# HL model used (documented):
11 +# N(g, x) ~ C(g) * INT_2^x exp(-g/ln t) / ln^2 t dt,
12 +# C(g) = 2*C2 * prod_{p | g, p odd} (p-1)/(p-2), C2 = 0.6601618158468696.
13 +# Deterministic; no randomness.
14 +# Usage: python3 analyze_gaps.py 1e8 [checkpoints default 1e6,1e7,1e8,1e9,4e9 <= limit]
15 +# =============================================================================
16 +
17 +import json
18 +import sys
19 +import time
20 +from math import exp, log
21 +from pathlib import Path
22 +
23 +import numpy as np
24 +
25 +from core import primes_upto, sieve_segment
26 +
27 +DATA = Path(__file__).resolve().parent.parent / "data"
28 +DATA.mkdir(exist_ok=True)
29 +
30 +C2 = 0.6601618158468696
31 +GMAX = 400 # pair matrix covers gaps < GMAX
32 +
33 +
34 +def singular(g: int) -> float:
35 + """C(g) = 2*C2 * prod_{p|g, p>2} (p-1)/(p-2)."""
36 + c = 2 * C2
37 + d, m = 3, g
38 + while d * d <= m:
39 + if m % d == 0:
40 + if d > 2:
41 + c *= (d - 1) / (d - 2)
42 + while m % d == 0:
43 + m //= d
44 + d += 2 if d > 2 else 1
45 + if m > 2:
46 + c *= (m - 1) / (m - 2)
47 + return c
48 +
49 +
50 +def hl_integral(g: int, x: int) -> float:
51 + """INT_2^x exp(-g/ln t)/ln^2 t dt, trapezoid on a log grid."""
52 + ts = np.exp(np.linspace(log(3), log(x), 4000))
53 + ys = np.exp(-g / np.log(ts)) / np.log(ts) ** 2
54 + return float(np.trapezoid(ys, ts))
55 +
56 +
57 +def scan(limit: int, checkpoints, segment_size: int = 50_000_000):
58 + base = primes_upto(int(limit ** 0.5) + 1)
59 + hist = np.zeros(4000, dtype=np.int64)
60 + pair = np.zeros((GMAX, GMAX), dtype=np.int64)
61 + # streaming sums for Pearson correlation of (g_n, g_{n+1})
62 + S = dict(n=0, x=0.0, y=0.0, xx=0.0, yy=0.0, xy=0.0)
63 + cp_hists, cp_corr = {}, {}
64 + cps = sorted(checkpoints)
65 + # segment boundaries aligned on checkpoints so each checkpoint histogram
66 + # covers exactly [2, checkpoint)
67 + bounds = sorted({limit, *cps, *range(2, limit, segment_size)} - {2})
68 + prev = None # last prime seen
69 + prev_gap = None # gap ending at prev
70 + t0 = time.time()
71 + lo = 2
72 + for hi in bounds:
73 + primes = sieve_segment(lo, hi, base)
74 + if len(primes) == 0:
75 + continue
76 + if prev is not None:
77 + primes = np.concatenate(([prev], primes))
78 + if len(primes) >= 2:
79 + gaps = np.diff(primes)
80 + hist += np.bincount(gaps, minlength=len(hist))[: len(hist)]
81 + # consecutive pairs, including the junction pair across segments
82 + if prev_gap is not None:
83 + gl = np.concatenate(([prev_gap], gaps))
84 + else:
85 + gl = gaps
86 + a, b = gl[:-1].astype(np.float64), gl[1:].astype(np.float64)
87 + am, bm = gl[:-1], gl[1:]
88 + mask = (am < GMAX) & (bm < GMAX)
89 + np.add.at(pair, (am[mask], bm[mask]), 1)
90 + S["n"] += len(a)
91 + S["x"] += a.sum(); S["y"] += b.sum()
92 + S["xx"] += (a * a).sum(); S["yy"] += (b * b).sum()
93 + S["xy"] += (a * b).sum()
94 + prev_gap = int(gaps[-1])
95 + prev = primes[-1]
96 + while cps and hi >= cps[0]:
97 + c = cps.pop(0)
98 + cp_hists[c] = hist.copy()
99 + cp_corr[c] = dict(S)
100 + print(f" checkpoint {c:.1e} reached at {time.time()-t0:.1f}s")
101 + lo = hi
102 + return hist, pair, cp_hists, cp_corr, time.time() - t0
103 +
104 +
105 +def pearson(S):
106 + n = S["n"]
107 + cov = S["xy"] / n - (S["x"] / n) * (S["y"] / n)
108 + vx = S["xx"] / n - (S["x"] / n) ** 2
109 + vy = S["yy"] / n - (S["y"] / n) ** 2
110 + return cov / (vx * vy) ** 0.5
111 +
112 +
113 +def main():
114 + limit = int(float(sys.argv[1])) if len(sys.argv) > 1 else 10**8
115 + cps = [int(c) for c in (10**6, 10**7, 10**8, 10**9, 4 * 10**9) if c <= limit]
116 + tag = f"{limit:.0e}".replace("+0", "").replace("+", "")
117 + hist, pair, cp_hists, cp_corr, dt = scan(limit, cps)
118 +
119 + np.savez_compressed(
120 + DATA / f"gapstats_{tag}.npz",
121 + hist=hist, pair=pair,
122 + checkpoints=np.array(sorted(cp_hists)),
123 + **{f"hist_{c}": h for c, h in cp_hists.items()},
124 + )
125 +
126 + out = {"limit": limit, "scan_seconds": round(dt, 1), "checkpoints": {}}
127 + for c in sorted(cp_hists):
128 + h = cp_hists[c]
129 + nz = np.nonzero(h)[0]
130 + even = nz[nz % 2 == 0]
131 + # HL comparison for even gaps up to 120
132 + ratios = {}
133 + for g in [int(v) for v in even if 2 <= v <= 120]:
134 + pred = singular(g) * hl_integral(g, c)
135 + ratios[g] = round(int(h[g]) / pred, 4) if pred > 0 else None
136 + champ = int(nz[np.argmax(h[nz])])
137 + # largest G such that every multiple of 6 up to G is a strict local max
138 + # of the gap histogram (h[g] > h[g-2] and h[g] > h[g+2])
139 + mult6_localmax_upto = 0
140 + for g in range(6, int(even.max()) - 2, 6):
141 + if h[g] > h[g - 2] and h[g] > h[g + 2]:
142 + mult6_localmax_upto = g
143 + else:
144 + break
145 + out["checkpoints"][str(c)] = {
146 + "jumping_champion": champ,
147 + "N2": int(h[2]), "N4": int(h[4]), "N6": int(h[6]),
148 + "N2_gt_N4": bool(h[2] > h[4]),
149 + "mult6_local_max_up_to": int(mult6_localmax_upto),
150 + "pearson_consecutive_gaps": round(pearson(cp_corr[c]), 5),
151 + "HL_ratio_obs_over_pred": ratios,
152 + }
153 + with open(DATA / f"gapstats_{tag}.json", "w") as f:
154 + json.dump(out, f, indent=2)
155 + # console digest
156 + for c in sorted(cp_hists):
157 + o = out["checkpoints"][str(c)]
158 + r = o["HL_ratio_obs_over_pred"]
159 + rv = [v for v in r.values() if v]
160 + print(f"x={c:.0e}: champ={o['jumping_champion']} N2={o['N2']} N4={o['N4']} "
161 + f"N2>N4={o['N2_gt_N4']} 6|g-localmax-upto={o['mult6_local_max_up_to']} "
162 + f"rho={o['pearson_consecutive_gaps']} HLratio[min,max]=[{min(rv):.3f},{max(rv):.3f}]")
163 +
164 +
165 +if __name__ == "__main__":
166 + main()
added src/core.py +207 −0
@@ -0,0 +1,207 @@
1 +#!/usr/bin/env python3
2 +# =============================================================================
3 +# core.py — Prime Mystery Engine: core primitives
4 +# Author: Simon-Pierre Boucher — contact@spboucher.ai
5 +# =============================================================================
6 +# Building blocks (Phase 0 of the protocol):
7 +# - primes_upto(n) : numpy Sieve of Eratosthenes (odd-only)
8 +# - segmented_primes(lo,hi) : segmented sieve, constant memory per segment
9 +# - is_prime(n) : deterministic Miller-Rabin for n < 3.317e24
10 +# - bpsw(n) : Baillie-PSW (MR base 2 + strong Lucas)
11 +# - prime_count(n) : pi(n) via sieve (validation helper)
12 +# All routines are pure-Python/numpy, no hidden prime tables beyond the
13 +# 12 deterministic MR bases (which are witnesses, not a prime list).
14 +# =============================================================================
15 +
16 +from __future__ import annotations # keep annotations lazy (Python 3.9 nodes)
17 +
18 +import numpy as np
19 +
20 +# Deterministic Miller-Rabin witness set: correct for all n < 3.317e24
21 +# (Sorenson & Webster 2015). Documented per the Golden Rule.
22 +_MR_BASES = (2, 3, 5, 7, 11, 13, 17, 19, 23, 29, 31, 37)
23 +_MR_LIMIT = 3_317_044_064_679_887_385_961_981
24 +
25 +
26 +def primes_upto(n: int) -> np.ndarray:
27 + """All primes <= n, via odd-only Sieve of Eratosthenes (numpy)."""
28 + if n < 2:
29 + return np.array([], dtype=np.int64)
30 + if n < 3:
31 + return np.array([2], dtype=np.int64)
32 + # index i represents the odd number 2i+1; index 0 (=1) is not prime
33 + size = (n + 1) // 2
34 + sieve = np.ones(size, dtype=bool)
35 + sieve[0] = False
36 + for i in range(1, (int(n ** 0.5) + 1) // 2 + 1):
37 + if sieve[i]:
38 + p = 2 * i + 1
39 + start = (p * p) // 2
40 + sieve[start::p] = False
41 + odds = 2 * np.nonzero(sieve)[0].astype(np.int64) + 1
42 + return np.concatenate(([np.int64(2)], odds))
43 +
44 +
45 +def sieve_segment(lo: int, hi: int, base_primes: np.ndarray | None = None) -> np.ndarray:
46 + """Primes in [lo, hi) via segmented sieve. base_primes must cover sqrt(hi)."""
47 + if hi <= 2:
48 + return np.array([], dtype=np.int64)
49 + lo = max(lo, 2)
50 + if base_primes is None:
51 + base_primes = primes_upto(int(hi ** 0.5) + 1)
52 + seg = np.ones(hi - lo, dtype=bool)
53 + for p in base_primes:
54 + p = int(p)
55 + if p * p >= hi:
56 + break
57 + start = max(p * p, ((lo + p - 1) // p) * p)
58 + seg[start - lo::p] = False
59 + if lo <= 1:
60 + seg[: 2 - lo] = False
61 + return lo + np.nonzero(seg)[0].astype(np.int64)
62 +
63 +
64 +def segmented_primes(lo: int, hi: int, segment_size: int = 10_000_000):
65 + """Yield numpy arrays of primes covering [lo, hi) segment by segment."""
66 + base = primes_upto(int(hi ** 0.5) + 1)
67 + for start in range(lo, hi, segment_size):
68 + end = min(start + segment_size, hi)
69 + yield sieve_segment(start, end, base)
70 +
71 +
72 +def prime_count(n: int) -> int:
73 + """pi(n), computed by segmented sieve (validation helper)."""
74 + total = 0
75 + for chunk in segmented_primes(2, n + 1):
76 + total += len(chunk)
77 + return total
78 +
79 +
80 +def _mr_witness(n: int, a: int, d: int, r: int) -> bool:
81 + """True if a is a Miller-Rabin witness for compositeness of n."""
82 + x = pow(a, d, n)
83 + if x == 1 or x == n - 1:
84 + return False
85 + for _ in range(r - 1):
86 + x = x * x % n
87 + if x == n - 1:
88 + return False
89 + return True
90 +
91 +
92 +def is_prime(n: int) -> bool:
93 + """Deterministic primality for n < 3.317e24 (Miller-Rabin, 12 bases).
94 + For larger n, falls back to BPSW (documented: no known counterexample,
95 + verified exhaustively below 2^64)."""
96 + if n < 2:
97 + return False
98 + for p in (2, 3, 5, 7, 11, 13, 17, 19, 23, 29, 31, 37):
99 + if n % p == 0:
100 + return n == p
101 + d, r = n - 1, 0
102 + while d % 2 == 0:
103 + d //= 2
104 + r += 1
105 + if n < _MR_LIMIT:
106 + return not any(_mr_witness(n, a, d, r) for a in _MR_BASES if a % n)
107 + return bpsw(n)
108 +
109 +
110 +def _jacobi(a: int, n: int) -> int:
111 + """Jacobi symbol (a/n), n odd positive."""
112 + a %= n
113 + result = 1
114 + while a:
115 + while a % 2 == 0:
116 + a //= 2
117 + if n % 8 in (3, 5):
118 + result = -result
119 + a, n = n, a
120 + if a % 4 == 3 and n % 4 == 3:
121 + result = -result
122 + a %= n
123 + return result if n == 1 else 0
124 +
125 +
126 +def _strong_lucas(n: int) -> bool:
127 + """Strong Lucas probable-prime test (Selfridge parameters). n odd, >2,
128 + not a perfect square, no small factors."""
129 + # Selfridge: first D in 5,-7,9,-11,... with Jacobi(D/n) = -1
130 + D = 5
131 + while True:
132 + j = _jacobi(D % n, n)
133 + if j == -1:
134 + break
135 + if j == 0 and abs(D) != n:
136 + return False
137 + D = -D - 2 if D > 0 else -D + 2
138 + if abs(D) > 1_000_000: # would indicate a square slipped through
139 + raise ArithmeticError("no suitable D found; n likely a square")
140 + Q = (1 - D) // 4
141 + # factor n+1 = d * 2^s
142 + d, s = n + 1, 0
143 + while d % 2 == 0:
144 + d //= 2
145 + s += 1
146 + # Lucas sequences U_d, V_d by binary ladder
147 + U, V, k = 0, 2, 0
148 + Qk = 1
149 + inv2 = (n + 1) // 2 # inverse of 2 mod n (n odd)
150 + for bit in bin(d)[2:]:
151 + # double: U_{2k}=U_k V_k ; V_{2k}=V_k^2 - 2 Q^k
152 + U, V = U * V % n, (V * V - 2 * Qk) % n
153 + Qk = Qk * Qk % n
154 + if bit == "1":
155 + # increment: U_{k+1}=(P U + V)/2 ; V_{k+1}=(D U + P V)/2 ; P=1
156 + U, V = (U + V) * inv2 % n, (D * U + V) * inv2 % n
157 + Qk = Qk * Q % n
158 + if U == 0 or V == 0:
159 + return True
160 + for _ in range(s - 1):
161 + V = (V * V - 2 * Qk) % n
162 + if V == 0:
163 + return True
164 + Qk = Qk * Qk % n
165 + return False
166 +
167 +
168 +def _is_square(n: int) -> bool:
169 + r = math_isqrt(n)
170 + return r * r == n
171 +
172 +
173 +from math import isqrt as math_isqrt # noqa: E402
174 +
175 +
176 +def bpsw(n: int) -> bool:
177 + """Baillie-PSW: strong MR base 2 + strong Lucas (Selfridge).
178 + No known counterexample; deterministic below 2^64."""
179 + if n < 2:
180 + return False
181 + for p in (2, 3, 5, 7, 11, 13, 17, 19, 23, 29, 31, 37):
182 + if n % p == 0:
183 + return n == p
184 + d, r = n - 1, 0
185 + while d % 2 == 0:
186 + d //= 2
187 + r += 1
188 + if _mr_witness(n, 2, d, r):
189 + return False
190 + if _is_square(n):
191 + return False
192 + return _strong_lucas(n)
193 +
194 +
195 +def gaps_in_range(lo: int, hi: int, segment_size: int = 50_000_000):
196 + """Iterate consecutive prime gaps over [lo, hi). Yields (p, gap) where
197 + gap = next_prime(p) - p, for all consecutive pairs with p in [lo, hi)."""
198 + prev = None
199 + for chunk in segmented_primes(max(lo, 2), hi, segment_size):
200 + if len(chunk) == 0:
201 + continue
202 + if prev is not None:
203 + yield int(prev), int(chunk[0] - prev)
204 + diffs = np.diff(chunk)
205 + for p, g in zip(chunk[:-1], diffs):
206 + yield int(p), int(g)
207 + prev = chunk[-1]
added src/cycle2_c4.py +138 −0
@@ -0,0 +1,138 @@
1 +#!/usr/bin/env python3
2 +# =============================================================================
3 +# cycle2_c4.py — Cycle 2 ADVERSARY: C4 (scaled gap anticorrelation) to 10^10+
4 +# Author: Simon-Pierre Boucher — contact@spboucher.ai
5 +# =============================================================================
6 +# Attacks C4: rho(x) < 0 and rho(x)*ln x -> c in [-0.60, -0.50] ?
7 +# Single deterministic segmented scan of [2, LIMIT). Collects:
8 +# - lag-1 Pearson rho(x) and lag-2 Pearson rho2(x) of consecutive gaps,
9 +# at checkpoints {1,2,4}x10^k (streaming sums; junction gaps handled)
10 +# - C1/C3 continuation: jumping champion, G6(x) at checkpoints
11 +# - C2 race D = N2 - N4 at checkpoints (lead-change census continuation)
12 +# - C5 continuation: maximal gaps (validated against published table)
13 +# Output: data/cycle2_c4_<L>.json — bit-for-bit reproducible on any machine.
14 +# Usage: python3 cycle2_c4.py 1e10 [hosttag]
15 +# =============================================================================
16 +
17 +import json
18 +import platform
19 +import sys
20 +import time
21 +from math import log
22 +from pathlib import Path
23 +
24 +import numpy as np
25 +
26 +from core import primes_upto, sieve_segment
27 +
28 +DATA = Path(__file__).resolve().parent.parent / "data"
29 +DATA.mkdir(exist_ok=True)
30 +
31 +
32 +def pearson(S):
33 + n = S["n"]
34 + if n < 2:
35 + return 0.0
36 + cov = S["xy"] / n - (S["x"] / n) * (S["y"] / n)
37 + vx = S["xx"] / n - (S["x"] / n) ** 2
38 + vy = S["yy"] / n - (S["y"] / n) ** 2
39 + return cov / (vx * vy) ** 0.5
40 +
41 +
42 +def acc(S, a, b):
43 + S["n"] += len(a)
44 + S["x"] += a.sum(); S["y"] += b.sum()
45 + S["xx"] += (a * a).sum(); S["yy"] += (b * b).sum()
46 + S["xy"] += (a * b).sum()
47 +
48 +
49 +def main():
50 + limit = int(float(sys.argv[1])) if len(sys.argv) > 1 else 10**10
51 + tag = sys.argv[2] if len(sys.argv) > 2 else platform.node()
52 + seg_size = 50_000_000
53 + cps = sorted(c for k in range(6, 12) for c in (10**k, 2 * 10**k, 4 * 10**k)
54 + if c <= limit)
55 + bounds = sorted({limit, *cps, *range(2, limit, seg_size)} - {2})
56 + base = primes_upto(int(limit ** 0.5) + 1)
57 +
58 + hist = np.zeros(6000, dtype=np.int64)
59 + S1 = dict(n=0, x=0.0, y=0.0, xx=0.0, yy=0.0, xy=0.0) # lag-1
60 + S2 = dict(n=0, x=0.0, y=0.0, xx=0.0, yy=0.0, xy=0.0) # lag-2
61 + D = 0
62 + maximal = []
63 + best = 0
64 + out_cps = {}
65 + prev = None
66 + tail = [] # last up-to-2 gaps before current segment
67 + t0 = time.time()
68 + lo = 2
69 + cps_left = list(cps)
70 + for hi in bounds:
71 + primes = sieve_segment(lo, hi, base)
72 + if len(primes):
73 + if prev is not None:
74 + primes = np.concatenate(([prev], primes))
75 + if len(primes) >= 2:
76 + gaps = np.diff(primes)
77 + hist += np.bincount(gaps, minlength=len(hist))[: len(hist)]
78 + D += int((gaps == 2).sum() - (gaps == 4).sum())
79 + gl = np.concatenate((np.array(tail, dtype=gaps.dtype), gaps))
80 + k = len(tail)
81 + if k >= 1:
82 + acc(S1, gl[k - 1:-1].astype(np.float64),
83 + gl[k:].astype(np.float64))
84 + else: # very first segment: all internal lag-1 pairs
85 + acc(S1, gl[:-1].astype(np.float64), gl[1:].astype(np.float64))
86 + if k >= 2:
87 + acc(S2, gl[k - 2:-2].astype(np.float64),
88 + gl[k:].astype(np.float64))
89 + elif len(gl) > 2: # first segments: internal lag-2 pairs
90 + acc(S2, gl[:-2].astype(np.float64), gl[2:].astype(np.float64))
91 + if int(gaps.max()) > best:
92 + starts = primes[:-1]
93 + for i in np.nonzero(gaps > best)[0]:
94 + g, p = int(gaps[i]), int(starts[i])
95 + if g > best:
96 + best = g
97 + maximal.append((g, p))
98 + tail = [int(v) for v in gaps[-2:]]
99 + prev = primes[-1]
100 + while cps_left and hi >= cps_left[0]:
101 + c = cps_left.pop(0)
102 + nz = np.nonzero(hist)[0]
103 + champ = int(nz[np.argmax(hist[nz])])
104 + g6 = 0
105 + for g in range(6, int(nz.max()) - 2, 6):
106 + if hist[g] > hist[g - 2] and hist[g] > hist[g + 2]:
107 + g6 = g
108 + else:
109 + break
110 + r1, r2 = pearson(S1), pearson(S2)
111 + out_cps[str(c)] = {
112 + "rho1": round(r1, 7), "rho1_lnx": round(r1 * log(c), 5),
113 + "rho2": round(r2, 7), "rho2_lnx": round(r2 * log(c), 5),
114 + "champion": champ, "G6": g6, "D_N2_minus_N4": D,
115 + "n_gaps": int(S1["n"] + 1),
116 + }
117 + print(f" cp {c:.0e}: rho1*lnx={r1*log(c):+.5f} rho2*lnx={r2*log(c):+.5f} "
118 + f"champ={champ} G6={g6} D={D} [{time.time()-t0:.0f}s]", flush=True)
119 + lo = hi
120 +
121 + out = {
122 + "limit": limit,
123 + "host": tag,
124 + "scan_seconds": round(time.time() - t0, 1),
125 + "checkpoints": out_cps,
126 + "maximal_gaps": [{"gap": g, "after": p, "csg": round(g / log(p) ** 2, 5)}
127 + for g, p in maximal],
128 + "method": "deterministic segmented sieve; streaming Pearson sums; "
129 + "gaps attributed to END prime; no randomness",
130 + }
131 + fname = DATA / f"cycle2_c4_{limit:.0e}_{tag}.json".replace("+", "").replace("e0", "e")
132 + with open(fname, "w") as f:
133 + json.dump(out, f, indent=2)
134 + print("written:", fname)
135 +
136 +
137 +if __name__ == "__main__":
138 + main()
added src/explore_gaps.py +120 −0
@@ -0,0 +1,120 @@
1 +#!/usr/bin/env python3
2 +# =============================================================================
3 +# explore_gaps.py — Cycle 1 EXPLORER: prime-gap statistics up to a bound
4 +# Author: Simon-Pierre Boucher — contact@spboucher.ai
5 +# =============================================================================
6 +# Fully vectorized (numpy) segmented scan of [2, LIMIT). Produces:
7 +# data/maximal_gaps_<L>.csv : record (maximal) gaps with merit and CSG ratio
8 +# data/gap_histogram_<L>.csv : count of each even gap g
9 +# data/gap_mod6_start_<L>.csv : histogram of (start prime mod 6, gap mod 6)
10 +# data/summary_<L>.json : headline numbers (deterministic, no seed used)
11 +# Method: segmented Sieve of Eratosthenes (core.py), np.diff per segment,
12 +# np.bincount accumulation. No per-prime Python loop.
13 +# Usage: python3 explore_gaps.py 1e8
14 +# =============================================================================
15 +
16 +import json
17 +import sys
18 +import time
19 +from math import log
20 +from pathlib import Path
21 +
22 +import numpy as np
23 +
24 +from core import primes_upto, sieve_segment
25 +
26 +DATA = Path(__file__).resolve().parent.parent / "data"
27 +DATA.mkdir(exist_ok=True)
28 +
29 +
30 +def scan(limit: int, segment_size: int = 50_000_000):
31 + base = primes_upto(int(limit ** 0.5) + 1)
32 + hist = np.zeros(4000, dtype=np.int64) # gap -> count
33 + mod6 = np.zeros((6, 6), dtype=np.int64) # (p mod 6, gap mod 6) -> count
34 + maximal = [] # (gap, start prime) records
35 + best = 0
36 + prev = None # last prime of previous segment
37 + n_gaps = 0
38 + t0 = time.time()
39 + for lo in range(2, limit, segment_size):
40 + hi = min(lo + segment_size, limit)
41 + primes = sieve_segment(lo, hi, base)
42 + if len(primes) == 0:
43 + continue
44 + if prev is not None:
45 + primes = np.concatenate(([prev], primes))
46 + if len(primes) >= 2:
47 + gaps = np.diff(primes)
48 + starts = primes[:-1]
49 + n_gaps += len(gaps)
50 + hist += np.bincount(gaps, minlength=len(hist))[: len(hist)]
51 + np.add.at(mod6, (starts % 6, gaps % 6), 1)
52 + # record gaps within this segment
53 + seg_best = int(gaps.max())
54 + if seg_best > best:
55 + idx = np.nonzero(gaps > best)[0]
56 + for i in idx:
57 + g, p = int(gaps[i]), int(starts[i])
58 + if g > best:
59 + best = g
60 + maximal.append((g, p))
61 + prev = primes[-1]
62 + dt = time.time() - t0
63 + return hist, mod6, maximal, n_gaps, dt
64 +
65 +
66 +def main():
67 + limit = int(float(sys.argv[1])) if len(sys.argv) > 1 else 10**8
68 + tag = f"{limit:.0e}".replace("+0", "").replace("+", "")
69 + hist, mod6, maximal, n_gaps, dt = scan(limit)
70 +
71 + # maximal gaps with merit and Cramér-Shanks-Granville ratio
72 + with open(DATA / f"maximal_gaps_{tag}.csv", "w") as f:
73 + f.write("gap,start_prime,merit,csg_ratio\n")
74 + for g, p in maximal:
75 + f.write(f"{g},{p},{g/log(p):.6f},{g/log(p)**2:.6f}\n")
76 +
77 + nz = np.nonzero(hist)[0]
78 + with open(DATA / f"gap_histogram_{tag}.csv", "w") as f:
79 + f.write("gap,count\n")
80 + for g in nz:
81 + f.write(f"{g},{hist[g]}\n")
82 +
83 + with open(DATA / f"gap_mod6_start_{tag}.csv", "w") as f:
84 + f.write("p_mod6,gap_mod6,count\n")
85 + for a in range(6):
86 + for b in range(6):
87 + if mod6[a, b]:
88 + f.write(f"{a},{b},{mod6[a, b]}\n")
89 +
90 + # headline stats
91 + champion = int(nz[np.argmax(hist[nz])])
92 + g_arr = nz.astype(float)
93 + c_arr = hist[nz].astype(float)
94 + total = c_arr.sum()
95 + frac_mod6 = {r: float(c_arr[g_arr % 6 == r].sum() / total) for r in range(6)}
96 + gmax, pmax = maximal[-1]
97 + best_merit = max((g / log(p), g, p) for g, p in maximal if p > 100)
98 + best_csg = max((g / log(p) ** 2, g, p) for g, p in maximal if p > 100)
99 + summary = {
100 + "limit": limit,
101 + "n_gaps": int(n_gaps),
102 + "scan_seconds": round(dt, 1),
103 + "largest_gap": {"gap": gmax, "after_prime": pmax},
104 + "n_maximal_gaps": len(maximal),
105 + "jumping_champion": champion,
106 + "gap_fraction_mod6": frac_mod6,
107 + "best_merit": {"merit": round(best_merit[0], 5), "gap": best_merit[1], "after_prime": best_merit[2]},
108 + "best_csg": {"csg": round(best_csg[0], 5), "gap": best_csg[1], "after_prime": best_csg[2]},
109 + "method": "segmented sieve of Eratosthenes, numpy, deterministic (no seed)",
110 + }
111 + with open(DATA / f"summary_{tag}.json", "w") as f:
112 + json.dump(summary, f, indent=2)
113 + print(json.dumps(summary, indent=2))
114 + print(f"\nmaximal gaps ({len(maximal)}):")
115 + for g, p in maximal:
116 + print(f" gap {g:4d} after {p:>15d} merit {g/log(p):7.4f} csg {g/log(p)**2:.4f}")
117 +
118 +
119 +if __name__ == "__main__":
120 + main()
added src/fit_c4.py +66 −0
@@ -0,0 +1,66 @@
1 +#!/usr/bin/env python3
2 +# =============================================================================
3 +# fit_c4.py — Cycle 2 CONJECTURER: extrapolate rho1(x)*ln x from scan data
4 +# Author: Simon-Pierre Boucher — contact@spboucher.ai
5 +# =============================================================================
6 +# Fits rho1*lnx = c + d/lambda (lambda = ln x) by least squares on checkpoints
7 +# x >= 1e8 (below that, transient), reports the extrapolated constant c with
8 +# leave-one-out spread, and compares with the HL triple-correlation model
9 +# (data/model_c4.json), which predicts the same 1/lambda-drift FORM.
10 +# Usage: python3 fit_c4.py data/cycle2_c4_4e10_M3U96a.json
11 +# =============================================================================
12 +
13 +import json
14 +import sys
15 +from math import log
16 +from pathlib import Path
17 +
18 +import numpy as np
19 +
20 +DATA = Path(__file__).resolve().parent.parent / "data"
21 +
22 +
23 +def fit(lams, vals):
24 + A = np.vstack([np.ones_like(lams), 1.0 / lams]).T
25 + (c, d), res, *_ = np.linalg.lstsq(A, vals, rcond=None)
26 + return c, d
27 +
28 +
29 +def main():
30 + src = Path(sys.argv[1]) if len(sys.argv) > 1 else DATA / "cycle2_c4_4e10_M3U96a.json"
31 + d = json.load(open(src))
32 + pts = [(float(x), v["rho1_lnx"]) for x, v in d["checkpoints"].items()
33 + if float(x) >= 1e8]
34 + pts.sort()
35 + lams = np.array([log(x) for x, _ in pts])
36 + vals = np.array([v for _, v in pts])
37 + c, dd = fit(lams, vals)
38 + # leave-one-out spread on c
39 + cs = []
40 + for i in range(len(pts)):
41 + m = np.ones(len(pts), dtype=bool); m[i] = False
42 + cs.append(fit(lams[m], vals[m])[0])
43 + resid = vals - (c + dd / lams)
44 + print(f"points used (x >= 1e8): {len(pts)}")
45 + for (x, v), r in zip(pts, resid):
46 + print(f" x={x:.0e} rho1*lnx={v:+.5f} fit-residual={r:+.5f}")
47 + print(f"\nfit: rho1*lnx = c + d/lnx c = {c:+.5f} d = {dd:+.4f}")
48 + print(f"leave-one-out c range: [{min(cs):+.5f}, {max(cs):+.5f}]")
49 + print(f"max |residual| = {np.abs(resid).max():.5f}")
50 + model = DATA / "model_c4.json"
51 + if model.exists():
52 + m = json.load(open(model))
53 + big = m["predictions"].get("1e40", {})
54 + print(f"\nHL triple model for comparison: rho*lambda -> ~{big.get('rho_times_lambda')}"
55 + f" (same 1/lambda-drift form, ~3.6x smaller magnitude)")
56 + out = {"source": src.name, "n_points": len(pts),
57 + "c_extrapolated": round(float(c), 5), "d": round(float(dd), 4),
58 + "c_loo_range": [round(min(cs), 5), round(max(cs), 5)],
59 + "max_abs_residual": round(float(np.abs(resid).max()), 6)}
60 + with open(DATA / "fit_c4.json", "w") as f:
61 + json.dump(out, f, indent=2)
62 + print("written:", DATA / "fit_c4.json")
63 +
64 +
65 +if __name__ == "__main__":
66 + main()
added src/gpu_runner.py +28 −0
@@ -0,0 +1,28 @@
1 +#!/usr/bin/env python3
2 +# =============================================================================
3 +# gpu_runner.py — Cycle 4: persistent GPU worker looping over a chunk list
4 +# Author: Simon-Pierre Boucher — contact@spboucher.ai
5 +# =============================================================================
6 +# Processes chunks sequentially in ONE process so the compiled Metal kernel
7 +# and base-prime array are reused. Skips chunks whose output already exists
8 +# (idempotent restart). Usage: python3 gpu_runner.py <chunks_file>
9 +# =============================================================================
10 +
11 +import sys
12 +from pathlib import Path
13 +
14 +from gpu_sieve import sieve_segment_gpu
15 +from worker_c4 import run
16 +
17 +def main():
18 + chunks = Path(sys.argv[1]).read_text().split()
19 + for i in range(0, len(chunks), 3):
20 + lo, hi, out = chunks[i], chunks[i + 1], chunks[i + 2]
21 + if Path(out).exists():
22 + continue
23 + run(int(float(lo)), int(float(hi)), out,
24 + sieve_fn=sieve_segment_gpu, tag="/gpu")
25 +
26 +
27 +if __name__ == "__main__":
28 + main()
added src/gpu_sieve.py +100 −0
@@ -0,0 +1,100 @@
1 +#!/usr/bin/env python3
2 +# =============================================================================
3 +# gpu_sieve.py — Cycle 4: segmented sieve on Apple GPU via MLX Metal kernel
4 +# Author: Simon-Pierre Boucher — contact@spboucher.ai
5 +# =============================================================================
6 +# One thread per (block, prime): marks multiples of base_primes[y] inside
7 +# block x of the segment byte-array. Write races are benign (all store 0).
8 +# Marking starts at max(p*p, first multiple >= block start) — identical
9 +# semantics to core.sieve_segment (validated in __main__ before any use).
10 +# =============================================================================
11 +
12 +import numpy as np
13 +
14 +import mlx.core as mx
15 +
16 +_SRC = r"""
17 + uint b = thread_position_in_grid.x; // block index
18 + uint pi = thread_position_in_grid.y; // prime index
19 + int64_t lo = params[0];
20 + int64_t hi = params[1];
21 + int64_t B = params[2];
22 + int64_t p = primes[pi];
23 + if (p * p >= hi) return;
24 + int64_t blo = lo + (int64_t)b * B;
25 + int64_t bhi = blo + B < hi ? blo + B : hi;
26 + if (blo >= bhi) return;
27 + int64_t start = ((blo + p - 1) / p) * p;
28 + if (start < p * p) start = p * p;
29 + for (int64_t m = start; m < bhi; m += p) {
30 + out[m - lo] = 0;
31 + }
32 +"""
33 +
34 +_kernel = mx.fast.metal_kernel(
35 + name="segmented_sieve",
36 + input_names=["primes", "params"],
37 + output_names=["out"],
38 + source=_SRC,
39 +)
40 +
41 +_BLOCK = 1_000_000
42 +
43 +
44 +def sieve_segment_gpu(lo: int, hi: int, base_primes) -> np.ndarray:
45 + """Primes in [lo, hi) — GPU-marked composites, CPU extraction."""
46 + lo = max(lo, 2)
47 + n = hi - lo
48 + if n <= 0:
49 + return np.array([], dtype=np.int64)
50 + primes_mx = mx.array(np.asarray(base_primes, dtype=np.int64))
51 + params = mx.array(np.array([lo, hi, _BLOCK], dtype=np.int64))
52 + nblocks = (n + _BLOCK - 1) // _BLOCK
53 + (out,) = _kernel(
54 + inputs=[primes_mx, params],
55 + grid=(nblocks, len(base_primes), 1),
56 + threadgroup=(min(nblocks, 32), 8, 1),
57 + output_shapes=[(n,)],
58 + output_dtypes=[mx.uint8],
59 + init_value=1,
60 + )
61 + flags = np.array(out, copy=False)
62 + if lo <= 2:
63 + # positions of 0 and 1 if present (lo==2 means none)
64 + pass
65 + res = lo + np.nonzero(flags)[0].astype(np.int64)
66 + # base primes < sqrt(hi) inside the segment survive (marking starts at p^2)
67 + return res
68 +
69 +
70 +if __name__ == "__main__":
71 + import sys
72 + import time
73 + from core import primes_upto, sieve_segment
74 +
75 + print("device:", mx.default_device())
76 + # correctness gate: GPU == CPU on varied windows
77 + ok = True
78 + windows = [(2, 10**6), (999_000_000, 1_001_000_000),
79 + (10**12, 10**12 + 5 * 10**7), (10**13 - 5 * 10**7, 10**13)]
80 + for lo, hi in windows:
81 + base = primes_upto(int(hi ** 0.5) + 1)
82 + g = sieve_segment_gpu(lo, hi, base)
83 + c = sieve_segment(lo, hi, base)
84 + same = np.array_equal(g, c)
85 + ok &= same
86 + print(f" [{'PASS' if same else 'FAIL'}] [{lo:.2e},{hi:.2e}): "
87 + f"{len(g)} primes (gpu) vs {len(c)} (cpu)")
88 + if not ok:
89 + sys.exit(1)
90 + # benchmark: one 1e9 chunk near 1e12, GPU vs single-core CPU
91 + lo, hi = 10**12, 10**12 + 10**9
92 + base = primes_upto(int(hi ** 0.5) + 1)
93 + for name, fn in (("gpu", sieve_segment_gpu), ("cpu", sieve_segment)):
94 + t0 = time.time()
95 + tot = 0
96 + for s in range(lo, hi, 50_000_000):
97 + tot += len(fn(s, min(s + 50_000_000, hi), base))
98 + print(f" bench {name}: 1e9-chunk near 1e12 in {time.time()-t0:.1f}s "
99 + f"({tot} primes)")
100 + print("GPU SIEVE VALIDATED")
added src/merge_c4.py +117 −0
@@ -0,0 +1,117 @@
1 +#!/usr/bin/env python3
2 +# =============================================================================
3 +# merge_c4.py — Cycle 3: merge distributed worker partials into checkpoints
4 +# Author: Simon-Pierre Boucher — contact@spboucher.ai
5 +# =============================================================================
6 +# Reads all worker partial JSONs, verifies the chunks tile [2, LIMIT) exactly,
7 +# merges cumulatively in range order, and emits checkpoint statistics at every
8 +# requested boundary. Cross-validates overlapping checkpoints against the
9 +# cycle-2 single-machine scan when available (5-decimal tolerance: summation
10 +# order differs, float64 pairwise sums agree to ~1e-9).
11 +# Usage: python3 merge_c4.py <partials_dir> <limit> [out.json]
12 +# =============================================================================
13 +
14 +import json
15 +import sys
16 +from math import log
17 +from pathlib import Path
18 +
19 +import numpy as np
20 +
21 +DATA = Path(__file__).resolve().parent.parent / "data"
22 +CHECKPOINTS = [10**10, 2 * 10**10, 4 * 10**10,
23 + 10**11, 2 * 10**11, 4 * 10**11, 10**12,
24 + 2 * 10**12, 4 * 10**12, 10**13]
25 +
26 +
27 +def pearson(S):
28 + n = S["n"]
29 + cov = S["xy"] / n - (S["x"] / n) * (S["y"] / n)
30 + vx = S["xx"] / n - (S["x"] / n) ** 2
31 + vy = S["yy"] / n - (S["y"] / n) ** 2
32 + return cov / (vx * vy) ** 0.5
33 +
34 +
35 +def main():
36 + pdir = Path(sys.argv[1])
37 + limit = int(float(sys.argv[2]))
38 + out_path = Path(sys.argv[3]) if len(sys.argv) > 3 else DATA / f"cycle3_c4_{limit:.0e}.json".replace("+", "").replace("e0", "e")
39 + parts = [json.load(open(p)) for p in sorted(pdir.glob("part_*.json"))]
40 + parts.sort(key=lambda d: d["lo"])
41 + # tiling check
42 + assert parts[0]["lo"] == 2, "first chunk must start at 2"
43 + for a, b in zip(parts, parts[1:]):
44 + assert a["hi"] == b["lo"], f"gap/overlap between chunks at {a['hi']} vs {b['lo']}"
45 + assert parts[-1]["hi"] == limit, f"last chunk ends at {parts[-1]['hi']}, expected {limit}"
46 + print(f"{len(parts)} chunks tile [2, {limit:.0e}) exactly; "
47 + f"hosts: {sorted(set(p['host'] for p in parts))}")
48 +
49 + hist = np.zeros(6000, dtype=np.int64)
50 + S1 = dict(n=0, x=0.0, y=0.0, xx=0.0, yy=0.0, xy=0.0)
51 + S2 = dict(n=0, x=0.0, y=0.0, xx=0.0, yy=0.0, xy=0.0)
52 + D = 0
53 + records = []
54 + best = 0
55 + cps = {}
56 + want = [c for c in CHECKPOINTS if c <= limit]
57 + for part in parts:
58 + for g, c in part["hist"].items():
59 + hist[int(g)] += c
60 + for k in S1:
61 + S1[k] += part["S1"][k]
62 + S2[k] += part["S2"][k]
63 + D += part["D"]
64 + for g, p in part["records"]:
65 + if g > best:
66 + best = g
67 + records.append((g, p))
68 + if want and part["hi"] == want[0]:
69 + c = want.pop(0)
70 + nz = np.nonzero(hist)[0]
71 + champ = int(nz[np.argmax(hist[nz])])
72 + g6 = 0
73 + for g in range(6, int(nz.max()) - 2, 6):
74 + if hist[g] > hist[g - 2] and hist[g] > hist[g + 2]:
75 + g6 = g
76 + else:
77 + break
78 + r1, r2 = pearson(S1), pearson(S2)
79 + cps[str(c)] = {"rho1": round(r1, 7), "rho1_lnx": round(r1 * log(c), 5),
80 + "rho2": round(r2, 7), "rho2_lnx": round(r2 * log(c), 5),
81 + "champion": champ, "G6": g6, "D_N2_minus_N4": D}
82 + print(f"cp {c:.0e}: rho1*lnx={r1*log(c):+.5f} rho2*lnx={r2*log(c):+.5f} "
83 + f"champ={champ} G6={g6} D={D}")
84 + assert not want, f"missing chunk boundaries for checkpoints {want}"
85 +
86 + # cross-validation against previous cycles' scans
87 + xval = {}
88 + for ref_name in ("cycle2_c4_4e10_M3U96a.json", "cycle3_c4_1e12.json"):
89 + ref_file = DATA / ref_name
90 + if not ref_file.exists():
91 + continue
92 + ref = json.load(open(ref_file))["checkpoints"]
93 + for c in list(ref):
94 + if c not in cps:
95 + continue
96 + d = abs(cps[c]["rho1_lnx"] - ref[c]["rho1_lnx"])
97 + ok = (d < 2e-5 and cps[c]["D_N2_minus_N4"] == ref[c]["D_N2_minus_N4"]
98 + and cps[c]["champion"] == ref[c]["champion"])
99 + xval[c] = {"ref": ref_name, "rho1_lnx_diff": round(d, 7), "exact_D_match": ok}
100 + print(f"x-val {float(c):.0e} vs {ref_name}: |drho1_lnx|={d:.2e} "
101 + f"D {'==' if ok else '!='} -> {'OK' if ok else 'FAIL'}")
102 +
103 + total_s = sum(p["seconds"] for p in parts)
104 + out = {"limit": limit, "n_chunks": len(parts),
105 + "hosts": sorted(set(p["host"] for p in parts)),
106 + "total_core_seconds": round(total_s, 1),
107 + "checkpoints": cps,
108 + "maximal_gaps": [{"gap": g, "after": p, "csg": round(g / log(p) ** 2, 5)}
109 + for g, p in records],
110 + "cross_validation_vs_cycle2": xval}
111 + with open(out_path, "w") as f:
112 + json.dump(out, f, indent=2)
113 + print("written:", out_path, f"(total {total_s:.0f} core-seconds)")
114 +
115 +
116 +if __name__ == "__main__":
117 + main()
added src/model_c4.py +86 −0
@@ -0,0 +1,86 @@
1 +#!/usr/bin/env python3
2 +# =============================================================================
3 +# model_c4.py — Cycle 2 PROVER: HL triple-correlation model for the C4 constant
4 +# Author: Simon-Pierre Boucher — contact@spboucher.ai
5 +# =============================================================================
6 +# Heuristic model for the joint law of consecutive gaps (g1, g2) at scale
7 +# lambda = ln x:
8 +# P(g1, g2) PROPORTIONAL TO W(g1, g2) * exp(-(g1+g2)/lambda)
9 +# where W is the Hardy-Littlewood singular series of the triple {0, g1, g1+g2}
10 +# restricted to primes p <= diameter (the tail p > g1+g2 is constant across
11 +# pairs and cancels in the normalization):
12 +# W = prod_{p odd, p <= G} [ (1 - nu_p/p) / (1 - 1/p)^3 ] * [p=2 term const]
13 +# nu_p = #{0, g1, g1+g2 mod p} (in {1,2,3}; equals 3 unless p | g1,
14 +# p | g2 or p | g1+g2)
15 +# KNOWN LIMITATION (documented): no inclusion-exclusion for "no prime strictly
16 +# inside the gaps" — this is the crude first-order model. If W == 1 (Cramer),
17 +# rho == 0 identically; all predicted anticorrelation is singular-series
18 +# coupling (the p | g1+g2 factor prevents factorization W = u(g1) v(g2)).
19 +# Output: model rho(lambda) * lambda for lambdas of interest + large-lambda
20 +# trend, printed and saved to data/model_c4.json. Deterministic.
21 +# =============================================================================
22 +
23 +import json
24 +from math import exp, log
25 +from pathlib import Path
26 +
27 +import numpy as np
28 +
29 +DATA = Path(__file__).resolve().parent.parent / "data"
30 +
31 +
32 +def small_primes(n):
33 + s = np.ones(n + 1, dtype=bool); s[:2] = False
34 + for i in range(2, int(n ** 0.5) + 1):
35 + if s[i]:
36 + s[i * i:: i] = False
37 + return np.nonzero(s)[0]
38 +
39 +
40 +def model_rho(lam: float, gmax: int) -> float:
41 + """Pearson correlation of (g1,g2) under the HL-coupled exponential model."""
42 + gs = np.arange(2, gmax + 1, 2, dtype=np.int64)
43 + g1 = gs[:, None]
44 + g2 = gs[None, :]
45 + gsum = g1 + g2
46 + logW = np.zeros((len(gs), len(gs)), dtype=np.float64)
47 + for p in small_primes(2 * gmax):
48 + if p == 2:
49 + continue # constant factor for even gaps (nu=1), cancels
50 + # nu_p on the grid
51 + d1 = (g1 % p == 0)
52 + d2 = (g2 % p == 0)
53 + ds = (gsum % p == 0)
54 + nu = np.full(logW.shape, 3, dtype=np.int64)
55 + nu[d1 | d2 | ds] = 2
56 + nu[d1 & d2] = 1
57 + logW += np.log(1 - nu / p) - 3 * log(1 - 1 / p)
58 + P = np.exp(logW - gsum / lam)
59 + P /= P.sum()
60 + m1 = (P * g1).sum(); m2 = (P * g2).sum()
61 + v1 = (P * (g1 - m1) ** 2).sum(); v2 = (P * (g2 - m2) ** 2).sum()
62 + cov = (P * (g1 - m1) * (g2 - m2)).sum()
63 + return float(cov / (v1 * v2) ** 0.5)
64 +
65 +
66 +def main():
67 + out = {"model": "P(g1,g2) ~ W_HL(0,g1,g1+g2) * exp(-(g1+g2)/lambda)",
68 + "predictions": {}}
69 + print(f"{'lambda':>8} {'x=e^lam':>10} {'rho_model':>12} {'rho*lambda':>12}")
70 + for lam, label in [(log(1e8), "1e8"), (log(1e9), "1e9"), (log(1e10), "1e10"),
71 + (log(4e10), "4e10"), (log(1e12), "1e12"),
72 + (log(1e15), "1e15"), (log(1e20), "1e20"),
73 + (log(1e30), "1e30"), (log(1e40), "1e40")]:
74 + gmax = int(14 * lam) # truncation: weight e^-14 ~ 8e-7
75 + r = model_rho(lam, gmax)
76 + out["predictions"][label] = {"lambda": round(lam, 4),
77 + "rho": round(r, 6),
78 + "rho_times_lambda": round(r * lam, 5)}
79 + print(f"{lam:8.3f} {label:>10} {r:12.6f} {r*lam:12.5f}")
80 + with open(DATA / "model_c4.json", "w") as f:
81 + json.dump(out, f, indent=2)
82 + print("written:", DATA / "model_c4.json")
83 +
84 +
85 +if __name__ == "__main__":
86 + main()
added src/prover_desert.py +167 −0
@@ -0,0 +1,167 @@
1 +#!/usr/bin/env python3
2 +# =============================================================================
3 +# prover_desert.py — Cycle 1 PROVER: certified prime desert via covering system
4 +# Author: Simon-Pierre Boucher — contact@spboucher.ai
5 +# =============================================================================
6 +# Construction:
7 +# 1. Prime set Q = primes <= m. Greedily pick a residue class a_p mod p for
8 +# each p in Q so that every position i in 1..L satisfies i = a_p (mod p)
9 +# for some p, while NO class contains position 0 or L+1.
10 +# 2. CRT: x = -a_p (mod p) for all p. Then for EVERY t >= 0,
11 +# N = x + t*P (P = prod Q) has N+i = 0 (mod q_i) with N+i > q_i,
12 +# hence N+1 .. N+L are ALL composite. <-- PROVEN by covering
13 +# 3. Search the smallest t with N and N+L+1 both prime (deterministic
14 +# Miller-Rabin, valid since N + L + 1 < 3.317e24). Then the gap after N
15 +# is EXACTLY L+1. <-- CERTIFIED
16 +# Baseline to beat: the classic primorial construction with the same Q gives
17 +# only L = (next prime after m) - 1 - 1 interior composites (gap ~ next prime).
18 +# Output: certs/desert_certificate.json (+ console report)
19 +# Deterministic: greedy tie-break is lexicographic; t-search is sequential.
20 +# =============================================================================
21 +
22 +import json
23 +import time
24 +from math import log, prod
25 +from pathlib import Path
26 +
27 +from core import is_prime, primes_upto
28 +
29 +CERTS = Path(__file__).resolve().parent.parent / "certs"
30 +CERTS.mkdir(exist_ok=True)
31 +
32 +MR_LIMIT = 3_317_044_064_679_887_385_961_981
33 +
34 +
35 +def greedy_cover(qs, L, order):
36 + """Try to cover positions 1..L with one class per prime (order given).
37 + Classes containing position 0 or L+1 are forbidden. Returns {p: a_p} or None."""
38 + uncovered = set(range(1, L + 1))
39 + classes = {}
40 + for p in order:
41 + best_class, best_hits = None, -1
42 + for a in range(p):
43 + if a == 0 % p or a == (L + 1) % p:
44 + continue
45 + hits = sum(1 for i in uncovered if i % p == a)
46 + if hits > best_hits:
47 + best_class, best_hits = a, hits
48 + if best_class is None:
49 + return None
50 + classes[p] = best_class
51 + uncovered = {i for i in uncovered if i % p != best_class}
52 + return classes if not uncovered else None
53 +
54 +
55 +def max_cover(qs):
56 + """Largest L (with its classes) coverable by greedy, over a few orderings."""
57 + best = (0, None)
58 + orders = [sorted(qs), sorted(qs, reverse=True)]
59 + for order in orders:
60 + L = 1
61 + last_good = None
62 + # increase L until greedy fails twice in a row (L must keep parity odd)
63 + fails = 0
64 + while fails < 6:
65 + cand = greedy_cover(qs, L, order)
66 + if cand:
67 + last_good = (L, cand)
68 + fails = 0
69 + else:
70 + fails += 1
71 + L += 2 # keep L odd: p=2 then covers odds, sparing 0 and L+1
72 + if last_good and last_good[0] > best[0]:
73 + best = last_good
74 + return best
75 +
76 +
77 +def crt(classes):
78 + """x with x = -a_p (mod p) for all p; 0 <= x < prod(p)."""
79 + x, mod = 0, 1
80 + for p, a in sorted(classes.items()):
81 + r = (-a) % p
82 + # solve x + mod*k = r (mod p)
83 + k = ((r - x) * pow(mod, -1, p)) % p
84 + x += mod * k
85 + mod *= p
86 + return x, mod
87 +
88 +
89 +def main():
90 + t0 = time.time()
91 + m = 59
92 + qs = [int(p) for p in primes_upto(m)]
93 + L, classes = max_cover(qs)
94 + assert classes is not None, "no covering found"
95 + print(f"prime set: primes <= {m} ({len(qs)} primes); covered interval length L = {L}")
96 + baseline = int(primes_upto(200)[len(qs)]) - 1 # next prime after m, minus 1
97 + print(f"baseline (classic primorial construction): {baseline} composites -> improvement x{L/baseline:.2f}")
98 +
99 + # verify covering exhaustively before CRT (belt and braces)
100 + for i in range(1, L + 1):
101 + assert any(i % p == a for p, a in classes.items()), f"position {i} uncovered"
102 + for p, a in classes.items():
103 + assert 0 % p != a and (L + 1) % p != a, f"class of {p} hits an endpoint"
104 +
105 + x, P = crt(classes)
106 + print(f"P = {m}# = {P}")
107 + # search smallest t with both endpoints prime; keep N + L + 1 under MR limit
108 + t_max = (MR_LIMIT - x - L - 1) // P
109 + found = None
110 + for t in range(1, t_max + 1):
111 + N = x + t * P
112 + if is_prime(N) and is_prime(N + L + 1):
113 + found = (t, N)
114 + break
115 + assert found, f"no prime endpoints for t <= {t_max}"
116 + t, N = found
117 + gap = L + 1
118 + merit = gap / log(N)
119 + print(f"t = {t}; N = {N} ({len(str(N))} digits)")
120 + print(f"CERTIFIED GAP: next_prime({N}) - {N} = {gap}, merit = {merit:.4f}, "
121 + f"csg = {gap/log(N)**2:.5f}")
122 +
123 + # interior divisor table: smallest covering prime per position
124 + divisors = {}
125 + for i in range(1, L + 1):
126 + q = min(p for p, a in classes.items() if i % p == a)
127 + assert (N + i) % q == 0 and N + i > q
128 + divisors[i] = q
129 +
130 + cert = {
131 + "title": "Certified prime desert via covering system",
132 + "author": "Simon-Pierre Boucher — contact@spboucher.ai",
133 + "date": "2026-08-06",
134 + "claim": f"The gap between consecutive primes N and N+{gap} is exactly {gap}.",
135 + "N": str(N),
136 + "gap": gap,
137 + "merit": round(merit, 5),
138 + "csg_ratio": round(gap / log(N) ** 2, 6),
139 + "construction": {
140 + "prime_set_max": m,
141 + "primorial_P": str(P),
142 + "crt_residue_x": str(x),
143 + "shift_t": t,
144 + "classes_a_p": {str(p): a for p, a in sorted(classes.items())},
145 + "note": "for EVERY t>=0, x+t*P+1 .. x+t*P+L are composite (covering); "
146 + "this specific t makes both endpoints prime",
147 + },
148 + "interior_divisors": {str(i): q for i, q in sorted(divisors.items())},
149 + "endpoint_primality": {
150 + "method": "deterministic Miller-Rabin, bases 2..37 (12 bases)",
151 + "validity_bound": str(MR_LIMIT),
152 + "both_endpoints_below_bound": bool(N + gap < MR_LIMIT),
153 + },
154 + "baseline_same_primes": baseline,
155 + "epistemic_status": "PROVEN (interior compositeness, by covering) + "
156 + "CERTIFIED (endpoint primality, deterministic MR). "
157 + "NOT a record (records.md: merit record is 41.94); "
158 + "value = fully certified pipeline.",
159 + }
160 + path = CERTS / "desert_certificate.json"
161 + with open(path, "w") as f:
162 + json.dump(cert, f, indent=2)
163 + print(f"certificate written: {path} ({time.time()-t0:.1f}s)")
164 +
165 +
166 +if __name__ == "__main__":
167 + main()
added src/prover_desert2.py +179 −0
@@ -0,0 +1,179 @@
1 +#!/usr/bin/env python3
2 +# =============================================================================
3 +# prover_desert2.py — Cycle 1 PROVER (attempt 2): hybrid covering desert
4 +# Author: Simon-Pierre Boucher — contact@spboucher.ai
5 +# =============================================================================
6 +# Upgrade over prover_desert.py: the covering system (primes <= 59) covers MOST
7 +# of positions 1..L; the few uncovered positions are certified composite for
8 +# the specific shift t by an explicit small factor (trial division) or an
9 +# explicit Miller-Rabin witness (a valid compositeness certificate).
10 +# Endpoints N, N+L+1 certified prime by deterministic MR (N+L+1 < 3.317e24).
11 +# Deterministic search; no randomness.
12 +# Output: certs/desert_certificate.json (overwrites attempt 1 - superseded)
13 +# =============================================================================
14 +
15 +import json
16 +import time
17 +from math import log
18 +from pathlib import Path
19 +
20 +from core import is_prime, primes_upto
21 +
22 +CERTS = Path(__file__).resolve().parent.parent / "certs"
23 +MR_LIMIT = 3_317_044_064_679_887_385_961_981
24 +TRIAL_PRIMES = [int(p) for p in primes_upto(1_000_000)]
25 +
26 +
27 +def greedy_cover_partial(qs, L):
28 + """One class per prime, greedy max coverage of 1..L; classes may not touch
29 + positions 0 or L+1. Returns (classes, uncovered_list)."""
30 + uncovered = set(range(1, L + 1))
31 + classes = {}
32 + for p in sorted(qs):
33 + best_a, best_hits = None, -1
34 + for a in range(p):
35 + if a == 0 or a == (L + 1) % p:
36 + continue
37 + hits = sum(1 for i in uncovered if i % p == a)
38 + if hits > best_hits:
39 + best_a, best_hits = a, hits
40 + classes[p] = best_a
41 + uncovered = {i for i in uncovered if i % p != best_a}
42 + return classes, sorted(uncovered)
43 +
44 +
45 +def crt(classes):
46 + x, mod = 0, 1
47 + for p, a in sorted(classes.items()):
48 + r = (-a) % p
49 + k = ((r - x) * pow(mod, -1, p)) % p
50 + x += mod * k
51 + mod *= p
52 + return x, mod
53 +
54 +
55 +def mr_witness_for(n):
56 + """Return a base a proving n composite (strong witness), or None."""
57 + d, r = n - 1, 0
58 + while d % 2 == 0:
59 + d //= 2
60 + r += 1
61 + for a in (2, 3, 5, 7, 11, 13, 17, 19, 23, 29, 31, 37):
62 + x = pow(a, d, n)
63 + if x == 1 or x == n - 1:
64 + continue
65 + for _ in range(r - 1):
66 + x = x * x % n
67 + if x == n - 1:
68 + break
69 + else:
70 + return a
71 + return None
72 +
73 +
74 +def small_factor(n):
75 + for q in TRIAL_PRIMES:
76 + if q * q > n:
77 + return None
78 + if n % q == 0:
79 + return q
80 + return None
81 +
82 +
83 +def main():
84 + t0 = time.time()
85 + m = 59
86 + qs = [int(p) for p in primes_upto(m)]
87 + # choose the largest odd L whose greedy covering leaves <= UMAX holes
88 + for UMAX in (16, 12, 8):
89 + chosen = None
90 + for L in range(91, 300, 2):
91 + classes, unc = greedy_cover_partial(qs, L)
92 + if len(unc) <= UMAX:
93 + chosen = (L, classes, unc)
94 + if chosen is None:
95 + continue
96 + L, classes, unc = chosen
97 + print(f"UMAX={UMAX}: L={L}, holes={len(unc)} at {unc}")
98 + x, P = crt(classes)
99 + t_max = (MR_LIMIT - x - L - 1) // P
100 + hit = None
101 + for t in range(1, int(t_max) + 1):
102 + N = x + t * P
103 + if not is_prime(N):
104 + continue
105 + if not is_prime(N + L + 1):
106 + continue
107 + if all(not is_prime(N + i) for i in unc):
108 + hit = (t, N)
109 + break
110 + if hit:
111 + break
112 + print(f" no t <= {t_max} works; relaxing")
113 + assert hit, "hybrid construction failed at all UMAX levels"
114 + t, N = hit
115 + gap = L + 1
116 + print(f"t={t} N={N} ({len(str(N))} digits)")
117 + print(f"CERTIFIED GAP {gap}, merit {gap/log(N):.4f}, csg {gap/log(N)**2:.5f} "
118 + f"[{time.time()-t0:.1f}s]")
119 +
120 + # build interior certificates
121 + interior = {}
122 + for i in range(1, L + 1):
123 + cov = [p for p, a in classes.items() if i % p == a]
124 + if cov:
125 + q = min(cov)
126 + assert (N + i) % q == 0 and N + i > q
127 + interior[str(i)] = {"type": "covering_factor", "q": q}
128 + else:
129 + f = small_factor(N + i)
130 + if f:
131 + interior[str(i)] = {"type": "trial_factor", "q": f}
132 + else:
133 + a = mr_witness_for(N + i)
134 + assert a is not None, f"no witness for position {i}?!"
135 + interior[str(i)] = {"type": "mr_witness", "a": a}
136 + n_types = {}
137 + for v in interior.values():
138 + n_types[v["type"]] = n_types.get(v["type"], 0) + 1
139 + print("interior certificate types:", n_types)
140 +
141 + cert = {
142 + "title": "Certified prime desert via hybrid covering system",
143 + "author": "Simon-Pierre Boucher — contact@spboucher.ai",
144 + "date": "2026-08-06",
145 + "claim": f"N and N+{gap} are consecutive primes (gap exactly {gap}).",
146 + "N": str(N),
147 + "gap": gap,
148 + "merit": round(gap / log(N), 5),
149 + "csg_ratio": round(gap / log(N) ** 2, 6),
150 + "construction": {
151 + "prime_set_max": m,
152 + "primorial_P": str(crt(classes)[1]),
153 + "crt_residue_x": str(x),
154 + "shift_t": t,
155 + "classes_a_p": {str(p): a for p, a in sorted(classes.items())},
156 + "holes": unc,
157 + "note": "positions with a covering factor are composite for EVERY "
158 + "shift t (proven); hole positions are certified composite "
159 + "for THIS t by explicit factor or strong MR witness",
160 + },
161 + "interior_certificates": interior,
162 + "endpoint_primality": {
163 + "method": "deterministic Miller-Rabin, bases {2..37} (12 bases), "
164 + "valid for n < 3.317e24 (Sorenson-Webster 2015)",
165 + "validity_bound": str(MR_LIMIT),
166 + "both_endpoints_below_bound": bool(N + gap < MR_LIMIT),
167 + },
168 + "epistemic_status": "CERTIFIED construction; NOT a record (see records.md: "
169 + "merit record 41.94, largest maximal gap 1854). "
170 + "Baseline with same primes: classic primorial gives gap 61; "
171 + "pure covering (attempt 1) gave gap 90.",
172 + }
173 + with open(CERTS / "desert_certificate.json", "w") as f:
174 + json.dump(cert, f, indent=2)
175 + print("certificate written:", CERTS / "desert_certificate.json")
176 +
177 +
178 +if __name__ == "__main__":
179 + main()
added src/test_core.py +93 −0
@@ -0,0 +1,93 @@
1 +#!/usr/bin/env python3
2 +# =============================================================================
3 +# test_core.py — Phase 0 validation: DO NOT proceed while any test fails
4 +# Author: Simon-Pierre Boucher — contact@spboucher.ai
5 +# =============================================================================
6 +# Validates every core building block against independently known values:
7 +# pi(10^k), maximal prime gaps below 10^6, strong-pseudoprime traps,
8 +# Carmichael numbers, cross-agreement sieve vs MR vs BPSW.
9 +# =============================================================================
10 +
11 +import sys
12 +import time
13 +
14 +import numpy as np
15 +
16 +from core import (bpsw, gaps_in_range, is_prime, prime_count, primes_upto,
17 + sieve_segment)
18 +
19 +FAILURES = []
20 +
21 +
22 +def check(name, got, expected):
23 + ok = got == expected
24 + print(f" [{'PASS' if ok else 'FAIL'}] {name}: got {got}, expected {expected}")
25 + if not ok:
26 + FAILURES.append(name)
27 +
28 +
29 +t0 = time.time()
30 +print("== pi(x) against known values ==")
31 +check("pi(10^2)", len(primes_upto(10**2)), 25)
32 +check("pi(10^4)", len(primes_upto(10**4)), 1229)
33 +check("pi(10^6)", len(primes_upto(10**6)), 78498)
34 +check("pi(10^7)", len(primes_upto(10**7)), 664579)
35 +check("pi(10^8) [segmented]", prime_count(10**8), 5761455)
36 +
37 +print("== segmented sieve == full sieve on random windows ==")
38 +full = primes_upto(2_000_000)
39 +for lo, hi in [(0, 1000), (999_000, 1_001_000), (1_500_000, 1_600_000)]:
40 + seg = sieve_segment(lo, hi)
41 + ref = full[(full >= lo) & (full < hi)]
42 + check(f"segment [{lo},{hi})", int(np.array_equal(seg, ref)), 1)
43 +
44 +print("== maximal prime gaps (known table, gaps starting below 10^6) ==")
45 +# (gap, prime after which it occurs) — classical table (OEIS A005250/A002386)
46 +KNOWN_MAXIMAL = [(4, 7), (6, 23), (8, 89), (14, 113), (18, 523), (20, 887),
47 + (22, 1129), (34, 1327), (36, 9551), (44, 15683), (52, 19609),
48 + (72, 31397), (86, 155921), (96, 360653), (112, 370261),
49 + (114, 492113)]
50 +best = 0
51 +found = []
52 +for p, g in gaps_in_range(2, 1_400_000):
53 + if g > best:
54 + best = g
55 + found.append((g, p))
56 +found = [(g, p) for g, p in found if g >= 4 and p < 10**6]
57 +check("maximal gaps < 10^6", found, KNOWN_MAXIMAL)
58 +
59 +print("== Miller-Rabin deterministic: traps and cross-check ==")
60 +# strong pseudoprimes base 2 must be caught
61 +for n in [2047, 3277, 4033, 4681, 8321, 15841, 29341]:
62 + check(f"spsp(2) {n} composite", is_prime(n), False)
63 +# Carmichael numbers
64 +for n in [561, 1105, 1729, 41041, 825265, 321197185]:
65 + check(f"Carmichael {n} composite", is_prime(n), False)
66 +# known primes, including large
67 +check("2^31-1 prime", is_prime(2**31 - 1), True)
68 +check("2^61-1 prime", is_prime(2**61 - 1), True)
69 +check("2^67-1 composite (Mersenne)", is_prime(2**67 - 1), False)
70 +check("10^18+9 prime", is_prime(10**18 + 9), True)
71 +# cross-check MR vs sieve over [0, 40000)
72 +ref_set = set(primes_upto(40000).tolist())
73 +mismatch = sum(1 for n in range(40000) if is_prime(n) != (n in ref_set))
74 +check("MR == sieve on [0,40000)", mismatch, 0)
75 +
76 +print("== BPSW: same traps + cross-check ==")
77 +for n in [2047, 3277, 561, 1105, 1729, 41041]:
78 + check(f"BPSW {n} composite", bpsw(n), False)
79 +check("BPSW 2^61-1 prime", bpsw(2**61 - 1), True)
80 +check("BPSW 10^18+9 prime", bpsw(10**18 + 9), True)
81 +mismatch = sum(1 for n in range(3, 40000) if bpsw(n) != (n in ref_set))
82 +check("BPSW == sieve on [3,40000)", mismatch, 0)
83 +# random 64-bit cross-check BPSW vs deterministic MR (fixed seed)
84 +rng = np.random.default_rng(42)
85 +vals = [int(x) for x in rng.integers(2**40, 2**62, size=300)]
86 +mismatch = sum(1 for n in vals if bpsw(n) != is_prime(n))
87 +check("BPSW == det-MR on 300 random 62-bit (seed 42)", mismatch, 0)
88 +
89 +print(f"\nTotal time: {time.time()-t0:.1f}s")
90 +if FAILURES:
91 + print(f"FAILED: {FAILURES}")
92 + sys.exit(1)
93 +print("ALL TESTS PASSED")
added src/worker_c4.py +112 −0
@@ -0,0 +1,112 @@
1 +#!/usr/bin/env python3
2 +# =============================================================================
3 +# worker_c4.py — Cycle 3: one distributed chunk of the C4' scan
4 +# Author: Simon-Pierre Boucher — contact@spboucher.ai
5 +# =============================================================================
6 +# Scans gaps whose END prime lies in [lo, hi) and emits MERGEABLE partial sums:
7 +# S1/S2 (lag-1/lag-2 Pearson streaming sums), gap histogram, D = N2-N4,
8 +# chunk-local record gaps. A 5000-wide sieve buffer below lo recovers the
9 +# two gaps preceding the first counted one (max gap < 1e12 is ~540, and any
10 +# 540-window below 1e12 contains a prime, so the buffer always suffices).
11 +# Deterministic. Usage: python3 worker_c4.py <lo> <hi> <out.json>
12 +# =============================================================================
13 +
14 +import json
15 +import platform
16 +import sys
17 +import time
18 +
19 +import numpy as np
20 +
21 +from core import primes_upto, sieve_segment
22 +
23 +BUFFER = 5000
24 +SEG = 50_000_000
25 +
26 +
27 +def run(lo, hi, out_path, sieve_fn=None, tag=""):
28 + sieve = sieve_fn or sieve_segment
29 + t0 = time.time()
30 + base = primes_upto(int(hi ** 0.5) + 1)
31 + start = 2 if lo <= 2 else lo - BUFFER
32 +
33 + hist = np.zeros(6000, dtype=np.int64)
34 + S1 = dict(n=0, x=0.0, y=0.0, xx=0.0, yy=0.0, xy=0.0)
35 + S2 = dict(n=0, x=0.0, y=0.0, xx=0.0, yy=0.0, xy=0.0)
36 + D = 0
37 + records = []
38 + best = 0
39 + n_gaps = 0
40 + prev = None
41 + tail = [] # up to 2 gaps immediately preceding current segment's first gap
42 + s = start
43 + while s < hi:
44 + e = min(s + SEG, hi)
45 + primes = sieve(s, e, base)
46 + s = e
47 + if len(primes) == 0:
48 + continue
49 + if prev is not None:
50 + primes = np.concatenate(([prev], primes))
51 + if len(primes) >= 2:
52 + gaps = np.diff(primes)
53 + ends = primes[1:]
54 + starts = primes[:-1]
55 + inr = ends >= lo # gaps counted by this worker (end prime >= lo)
56 + gin = gaps[inr]
57 + n_gaps += len(gin)
58 + hist += np.bincount(gin, minlength=len(hist))[: len(hist)]
59 + D += int((gin == 2).sum() - (gin == 4).sum())
60 + # pair sums: pair (g_{n-1}, g_n) / (g_{n-2}, g_n) counted at g_n
61 + gl = np.concatenate((np.array(tail, dtype=gaps.dtype), gaps))
62 + k = len(tail)
63 + idx = np.nonzero(inr)[0] + k # positions of counted gaps in gl
64 + i1 = idx[idx >= 1]
65 + a = gl[i1 - 1].astype(np.float64); b = gl[i1].astype(np.float64)
66 + S1["n"] += len(a)
67 + S1["x"] += float(a.sum()); S1["y"] += float(b.sum())
68 + S1["xx"] += float((a * a).sum()); S1["yy"] += float((b * b).sum())
69 + S1["xy"] += float((a * b).sum())
70 + i2 = idx[idx >= 2]
71 + a2 = gl[i2 - 2].astype(np.float64); b2 = gl[i2].astype(np.float64)
72 + S2["n"] += len(a2)
73 + S2["x"] += float(a2.sum()); S2["y"] += float(b2.sum())
74 + S2["xx"] += float((a2 * a2).sum()); S2["yy"] += float((b2 * b2).sum())
75 + S2["xy"] += float((a2 * b2).sum())
76 + # chunk-local records among counted gaps
77 + if len(gin) and int(gin.max()) > best:
78 + sin = starts[inr]
79 + for i in np.nonzero(gin > best)[0]:
80 + g, p = int(gin[i]), int(sin[i])
81 + if g > best:
82 + best = g
83 + records.append([g, p])
84 + tail = [int(v) for v in gaps[-2:]]
85 + prev = primes[-1]
86 +
87 + nz = np.nonzero(hist)[0]
88 + out = {
89 + "lo": lo, "hi": hi, "n_gaps": int(n_gaps), "D": D,
90 + "S1": S1, "S2": S2,
91 + "hist": {int(g): int(hist[g]) for g in nz},
92 + "records": records,
93 + "host": platform.node() + tag,
94 + "seconds": round(time.time() - t0, 1),
95 + }
96 + with open(out_path, "w") as f:
97 + json.dump(out, f)
98 + print(f"done [{lo},{hi}) on {out['host']} in {out['seconds']}s", flush=True)
99 +
100 +
101 +def main():
102 + lo, hi = int(float(sys.argv[1])), int(float(sys.argv[2]))
103 + out_path = sys.argv[3]
104 + if len(sys.argv) > 4 and sys.argv[4] == "gpu":
105 + from gpu_sieve import sieve_segment_gpu
106 + run(lo, hi, out_path, sieve_fn=sieve_segment_gpu, tag="/gpu")
107 + else:
108 + run(lo, hi, out_path)
109 +
110 +
111 +if __name__ == "__main__":
112 + main()
113