Paper v4: figures (scaling law, race, histogram, certificate map), Gallagher/Brent/Rankin/Lemke Oliver-Soundararajan math and references; cycle 5 launch infra (SEG env, idempotent workers)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
9 changed files +377 −12
added
paper/figs/fig_certificate.pdf
+0 −0
Binary file not shown.
added
paper/figs/fig_hist.pdf
+0 −0
Binary file not shown.
added
paper/figs/fig_race.pdf
+0 −0
Binary file not shown.
added
paper/figs/fig_scaling.pdf
+0 −0
Binary file not shown.
modified
paper/main.pdf
+0 −0
Binary file not shown.
modified
paper/main.tex
+130 −10
@@ -13,6 +13,7 @@ | ||
| 13 | 13 | \usepackage{algpseudocode} |
| 14 | 14 | \usepackage[hidelinks]{hyperref} |
| 15 | 15 | \usepackage{xcolor} |
| 16 | +\usepackage{graphicx} | |
| 16 | 17 | |
| 17 | 18 | \newtheorem{theorem}{Theorem} |
| 18 | 19 | \newtheorem{lemma}{Lemma} |
@@ -33,7 +34,7 @@ Certified Deserts, a Refuted Race Conjecture, and an Empirical Scaling Law\\ | ||
| 33 | 34 | for Consecutive-Gap Correlations} |
| 34 | 35 | \author{Simon-Pierre Boucher\\ |
| 35 | 36 | \texttt{contact@spboucher.ai}} |
| 36 | −\date{August 6, 2026 --- working draft v3} | |
| 37 | +\date{August 6, 2026 --- working draft v4} | |
| 37 | 38 | |
| 38 | 39 | \begin{document} |
| 39 | 40 | \maketitle |
@@ -141,13 +142,47 @@ Cram\'er's heuristic suggests $\limsup \mathrm{CSG} = 1$; Granville's correction | ||
| 141 | 142 | $\limsup \ge 2e^{-\gamma} \approx 1.1229$. The largest CSG value ever observed is |
| 142 | 143 | $0.9206\ldots$ (Nyman's gap 1132 after 1693182318746371). |
| 143 | 144 | |
| 144 | −\paragraph{Cram\'er random model.} | |
| 145 | −In the Cram\'er model each integer $n$ is independently ``prime'' with probability | |
| 146 | −$1/\ln n$; consecutive gaps near $x$ are then asymptotically i.i.d.\ | |
| 147 | −$\mathrm{Exponential}(\lambda)$ up to discretization, | |
| 148 | −so \emph{any} statistic that vanishes for independent gaps (such as the correlations of | |
| 149 | −Section~\ref{sec:c4}) measures genuinely arithmetic structure. We use this as a | |
| 150 | −significance filter (the ``Cram\'er/Maier test''). | |
| 145 | +\paragraph{Cram\'er random model and Gallagher's theorem.} | |
| 146 | +In the Cram\'er model \cite{Cramer} each integer $n$ is independently ``prime'' with | |
| 147 | +probability $1/\ln n$; consecutive gaps near $x$ are then asymptotically i.i.d.\ | |
| 148 | +$\mathrm{Exponential}(\lambda)$ up to discretization. Remarkably, this exponential law | |
| 149 | +is not merely a toy: Gallagher \cite{Gallagher} proved that the uniform Hardy--Littlewood | |
| 150 | +conjecture implies that the number of primes in the short interval $(x,\, x + \eta \ln x]$ | |
| 151 | +converges to a Poisson law, | |
| 152 | +\begin{equation} | |
| 153 | +\#\bigl\{ p \in (x,\, x + \eta \ln x] \bigr\} \;\xrightarrow{d}\; \mathrm{Poisson}(\eta), | |
| 154 | +\qquad\text{equivalently}\qquad | |
| 155 | +\Pr\!\Bigl( \frac{g}{\ln x} > t \Bigr) \;\to\; e^{-t}, | |
| 156 | +\label{eq:gallagher} | |
| 157 | +\end{equation} | |
| 158 | +because the singular series averages to one: | |
| 159 | +$\sum_{\mathcal H \subset [1,h],\, |\mathcal H| = k} \Sing(\mathcal H) \sim h^k / k!$. | |
| 160 | +Thus the \emph{marginal} exponential shape of gaps carries no arithmetic information at | |
| 161 | +leading order --- but joint statistics of \emph{consecutive} gaps do, since Gallagher's | |
| 162 | +averaging does not force independence at the $1/\lambda$ correction level. Any statistic | |
| 163 | +that vanishes for independent gaps (such as the correlations of Section~\ref{sec:c4}) | |
| 164 | +therefore measures genuinely arithmetic structure. We use this as a significance filter | |
| 165 | +(the ``Cram\'er/Maier test'', cf.\ Maier's matrix method \cite{Maier} which shows the | |
| 166 | +Cram\'er model's predictions fail in short intervals). | |
| 167 | + | |
| 168 | +\paragraph{Refined single-gap model.} | |
| 169 | +The standard first-order heuristic for the count of gaps of size exactly $g$ is | |
| 170 | +\begin{equation} | |
| 171 | +N(g, x) \;\approx\; C(g) \int_2^x e^{-g/\ln t}\, \frac{dt}{(\ln t)^2}, | |
| 172 | +\label{eq:single-gap} | |
| 173 | +\end{equation} | |
| 174 | +(interior compositeness absorbed into the exponential); Brent \cite{Brent} showed that | |
| 175 | +accurate gap counts require the full inclusion--exclusion over interior prime patterns, | |
| 176 | +\begin{equation} | |
| 177 | +N(g, x) \;\approx\; \sum_{k \ge 0} (-1)^k | |
| 178 | +\sum_{\substack{0 < h_1 < \cdots < h_k < g}} | |
| 179 | +\Sing\bigl(\{0, h_1, \dots, h_k, g\}\bigr) \int_2^x \frac{dt}{(\ln t)^{k+2}}, | |
| 180 | +\label{eq:brent} | |
| 181 | +\end{equation} | |
| 182 | +truncated at moderate $k$. Our own measurements confirm the necessity: the naive | |
| 183 | +\eqref{eq:single-gap} exhibits structured residuals (Section~\ref{sec:lessons}), and its | |
| 184 | +pairwise analogue underpredicts the gap correlation threefold | |
| 185 | +(Section~\ref{sec:model}). | |
| 151 | 186 | |
| 152 | 187 | \section{Methodological protocol} |
| 153 | 188 | \label{sec:protocol} |
@@ -383,6 +418,17 @@ $D = +26, +359, +55$, suggesting: | ||
| 383 | 418 | $N(2, x) > N(4, x)$ for all $x \ge 10^6$. |
| 384 | 419 | \end{conjecture} |
| 385 | 420 | |
| 421 | +\begin{figure}[t] | |
| 422 | +\centering | |
| 423 | +\includegraphics[width=\textwidth]{figs/fig_race.pdf} | |
| 424 | +\caption{The 2-vs-4 race $D(x) = N(2,x) - N(4,x)$ on a symlog scale. Blue: exact | |
| 425 | +$D$ sampled at $\sim$60 log-spaced points per decade up to $10^{10}$ (recomputed for | |
| 426 | +this figure); orange diamonds: campaign checkpoints to $10^{13}$. The lead changes | |
| 427 | +indefinitely; the walk stays within its diffusive envelope | |
| 428 | +$O(\sqrt{N(2,x)+N(4,x)})$.} | |
| 429 | +\label{fig:race} | |
| 430 | +\end{figure} | |
| 431 | + | |
| 386 | 432 | \begin{result}[\statuslbl{refuted with minimal counterexample}] |
| 387 | 433 | \label{res:race} |
| 388 | 434 | A full-resolution scan (every gap event inspected via cumulative sums) shows |
@@ -423,6 +469,21 @@ Rubinstein--Sarnak analysis for $\pi(x; 4, 1)$ vs $\pi(x; 4, 3)$). | ||
| 423 | 469 | \section{Result II: a certified prime desert from a hybrid covering system} |
| 424 | 470 | \label{sec:desert} |
| 425 | 471 | |
| 472 | +Explicit long runs of composites go back to Westzynthius \cite{West}, Erd\H{o}s | |
| 473 | +\cite{Erdos35} and Rankin \cite{Rankin}, whose residue-shifting refinements of the | |
| 474 | +primorial construction give | |
| 475 | +\begin{equation} | |
| 476 | +G(x) \;\ge\; (c + o(1))\, | |
| 477 | +\frac{\ln x \,\ln\ln x \,\ln\ln\ln\ln x}{(\ln\ln\ln x)^2} | |
| 478 | +\end{equation} | |
| 479 | +for the maximal gap $G(x)$ below $x$ (Rankin's form; the FGKMT/Maynard breakthroughs | |
| 480 | +\cite{FGKMT} improved the exponent structure of the triple-log factor and settled the | |
| 481 | +Erd\H{o}s \$10{,}000 problem). Our construction is a computational descendant of exactly | |
| 482 | +this residue-shifting idea, with two modern twists: the greedy choice of residues is | |
| 483 | +\emph{certified per position}, and a controlled number of uncovered positions | |
| 484 | +(``holes'') is permitted and certified a posteriori --- which is what lifts the achieved | |
| 485 | +length from the classical regime to $260$ at 25 digits. | |
| 486 | + | |
| 426 | 487 | \begin{result}[\statuslbl{certified}] |
| 427 | 488 | \label{res:desert} |
| 428 | 489 | Let $N = 1{,}116{,}336{,}781{,}708{,}038{,}449{,}369{,}693$ (25 digits). Then $N$ and |
@@ -469,6 +530,18 @@ endpoints, search over the shift $t$: at $t = 580$, | ||
| 469 | 530 | Hence every $N + i$, $1 \le i \le 259$, carries an explicit compositeness certificate |
| 470 | 531 | and both endpoints an unconditional primality certificate: the gap is exactly 260. |
| 471 | 532 | |
| 533 | +\begin{figure}[t] | |
| 534 | +\centering | |
| 535 | +\includegraphics[width=\textwidth]{figs/fig_certificate.pdf} | |
| 536 | +\caption{Certificate map of the 260-desert. Each interior position $i$ carries an | |
| 537 | +explicit compositeness certificate for $N+i$: blue dots — smallest covering prime | |
| 538 | +(valid for \emph{every} CRT shift $t$, Lemma 2); orange diamonds — explicit trial | |
| 539 | +factor at the 15 holes; aqua square — the one hole certified by a strong Miller--Rabin | |
| 540 | +witness (plotted at an arbitrary height: a witness proves compositeness without | |
| 541 | +exhibiting a factor).} | |
| 542 | +\label{fig:certificate} | |
| 543 | +\end{figure} | |
| 544 | + | |
| 472 | 545 | \paragraph{Endpoint-search heuristic.} |
| 473 | 546 | $N$ and $N + L + 1$ are constructed coprime to all $p \le 59$, so by the standard |
| 474 | 547 | sieve heuristic each is prime with probability |
@@ -520,6 +593,19 @@ $10^{13}$ & $-0.54264$ & $-0.25247$ \\ | ||
| 520 | 593 | \end{tabular} |
| 521 | 594 | \end{center} |
| 522 | 595 | |
| 596 | +\begin{figure}[t] | |
| 597 | +\centering | |
| 598 | +\includegraphics[width=\textwidth]{figs/fig_scaling.pdf} | |
| 599 | +\caption{Left: the scaling law. Circles: the 16 measured checkpoints for lag-1 (blue) | |
| 600 | +and lag-2 (aqua) with their two-parameter fits $c + d/\lambda$; orange diamonds: the | |
| 601 | +three out-of-sample predictions later confirmed by measurement ($10^{12}$, $10^{13}$ | |
| 602 | +lag-1; $10^{13}$ lag-2); gray squares: the standing predictions at $10^{14}$ and | |
| 603 | +$10^{15}$. Right: the same lag-1 data against the first-order Hardy--Littlewood | |
| 604 | +triple-correlation model \eqref{eq:triple-model}, which has the right sign and drift | |
| 605 | +form but $\approx 32\%$ of the magnitude.} | |
| 606 | +\label{fig:scaling} | |
| 607 | +\end{figure} | |
| 608 | + | |
| 523 | 609 | \begin{conjecture}[C4$'$; \statuslbl{conjectured}, one prediction confirmed] |
| 524 | 610 | \label{conj:c4p} |
| 525 | 611 | With $\lambda = \lnx$, |
@@ -611,8 +697,13 @@ out-of-sample ones in (i). | ||
| 611 | 697 | |
| 612 | 698 | We have not located these constants in the literature. The qualitative anticorrelation |
| 613 | 699 | of consecutive gaps is known empirically, and $c, c_2$ may be derivable from existing |
| 614 | −conditional results (e.g.\ pair-correlation heuristics); we make no novelty claim | |
| 615 | −pending a proper literature pass. | |
| 700 | +conditional results; we make no novelty claim pending a proper literature pass. The | |
| 701 | +closest well-studied phenomenon is the Lemke Oliver--Soundararajan bias | |
| 702 | +\cite{LOS}: consecutive primes are measurably biased against repeating their residue | |
| 703 | +class mod $q$, with a $1/\ln x$-order correction governed by the same triple singular | |
| 704 | +series that appears in \eqref{eq:triple-model} --- our $\rho(x)$ can be viewed as the | |
| 705 | +gap-space aggregate of those pattern biases, and their conditional analysis is the | |
| 706 | +natural route to a proof of Conjecture~\ref{conj:c4p}. | |
| 616 | 707 | |
| 617 | 708 | \section{Result IV: the first-order Hardy--Littlewood model fails quantitatively} |
| 618 | 709 | \label{sec:model} |
@@ -666,6 +757,16 @@ $\operatorname{argmax}_g N(g, x) = 6$ at every checkpoint from $10^6$ to $10^{12 | ||
| 666 | 757 | (\statuslbl{verified}), consistent with the Odlyzko--Rubinstein--Wolf conjecture that 6 |
| 667 | 758 | is the champion from $\approx 947$ to $\approx 1.7\times 10^{35}$. |
| 668 | 759 | |
| 760 | +\begin{figure}[t] | |
| 761 | +\centering | |
| 762 | +\includegraphics[width=\textwidth]{figs/fig_hist.pdf} | |
| 763 | +\caption{The gap histogram $N(g, 10^8)$ (log scale). Multiples of 6 (blue) are strict | |
| 764 | +local maxima --- the factor $(p-1)/(p-2) = 2$ at $p = 3$ in the singular series | |
| 765 | +\eqref{eq:singular-pair}; the jumping champion is $g = 6$ at every checkpoint from | |
| 766 | +$10^6$ to $10^{13}$.} | |
| 767 | +\label{fig:hist} | |
| 768 | +\end{figure} | |
| 769 | + | |
| 669 | 770 | \paragraph{Local dominance of multiples of 6.} |
| 670 | 771 | By \eqref{eq:singular-pair}, $3 \mid g$ multiplies $C(g)$ by |
| 671 | 772 | $(3-1)/(3-2) = 2$, so multiples of 6 should be strict local maxima of |
@@ -784,6 +885,25 @@ Re-verification of the two headline certificates requires only a stock Python~3: | ||
| 784 | 885 | \emph{Lucas pseudoprimes}, Math.~Comp.~35 (1980), 1391--1417. |
| 785 | 886 | \bibitem{RS} M.~Rubinstein, P.~Sarnak, \emph{Chebyshev's bias}, |
| 786 | 887 | Experiment.~Math.~3 (1994), 173--197. |
| 888 | +\bibitem{Cramer} H.~Cram\'er, \emph{On the order of magnitude of the difference between | |
| 889 | +consecutive prime numbers}, Acta Arith.~2 (1936), 23--46. | |
| 890 | +\bibitem{Gallagher} P.~X.~Gallagher, \emph{On the distribution of primes in short | |
| 891 | +intervals}, Mathematika 23 (1976), 4--9. | |
| 892 | +\bibitem{Maier} H.~Maier, \emph{Primes in short intervals}, Michigan Math.~J.~32 | |
| 893 | +(1985), 221--225. | |
| 894 | +\bibitem{Brent} R.~P.~Brent, \emph{The distribution of small gaps between successive | |
| 895 | +primes}, Math.~Comp.~28 (1974), 315--324. | |
| 896 | +\bibitem{LOS} R.~J.~Lemke Oliver, K.~Soundararajan, \emph{Unexpected biases in the | |
| 897 | +distribution of consecutive primes}, Proc.~Natl.~Acad.~Sci.~113 (2016), E4446--E4454. | |
| 898 | +\bibitem{West} E.~Westzynthius, \emph{\"Uber die Verteilung der Zahlen, die zu den $n$ | |
| 899 | +ersten Primzahlen teilerfremd sind}, Comm.~Phys.-Math.~Soc.~Sci.~Fenn.~5 (1931), 1--37. | |
| 900 | +\bibitem{Erdos35} P.~Erd\H{o}s, \emph{On the difference of consecutive primes}, | |
| 901 | +Q.~J.~Math.~6 (1935), 124--128. | |
| 902 | +\bibitem{Rankin} R.~A.~Rankin, \emph{The difference between consecutive prime numbers}, | |
| 903 | +J.~London Math.~Soc.~13 (1938), 242--247. | |
| 904 | +\bibitem{OeSHP} T.~Oliveira e Silva, S.~Herzog, S.~Pardi, \emph{Empirical verification | |
| 905 | +of the even Goldbach conjecture and computation of prime gaps up to $4\times 10^{18}$}, | |
| 906 | +Math.~Comp.~83 (2014), 2033--2060. | |
| 787 | 907 | \bibitem{Granville} A.~Granville, \emph{Harald Cram\'er and the distribution of prime |
| 788 | 908 | numbers}, Scand.~Actuar.~J.~1 (1995), 12--28. |
| 789 | 909 | \bibitem{A005250} OEIS Foundation, \emph{Sequences A005250, A002386 (record gaps |
added
src/figures.py
+239 −0
@@ -0,0 +1,239 @@ | ||
| 1 | +#!/usr/bin/env python3 | |
| 2 | +# ============================================================================= | |
| 3 | +# figures.py — paper figures from existing campaign data | |
| 4 | +# Author: Simon-Pierre Boucher — contact@spboucher.ai | |
| 5 | +# ============================================================================= | |
| 6 | +# Generates (into paper/figs/): | |
| 7 | +# fig_scaling.pdf rho·lnx vs lnx: data, fit, confirmed predictions, HL model | |
| 8 | +# fig_race.pdf D(x) = N2-N4 lead changes (fine trace to 1e10 + checkpoints) | |
| 9 | +# fig_hist.pdf N(g, 1e8) with the mod-6 comb structure | |
| 10 | +# fig_certificate.pdf the 260-gap desert certificate map | |
| 11 | +# Palette (validated, dataviz method): blue #2a78d6, orange #eb6834, aqua #1baf7a | |
| 12 | +# Deterministic; the race trace is recomputed locally (scan to 1e10, ~40 s). | |
| 13 | +# ============================================================================= | |
| 14 | + | |
| 15 | +import json | |
| 16 | +from math import log | |
| 17 | +from pathlib import Path | |
| 18 | + | |
| 19 | +import matplotlib | |
| 20 | +matplotlib.use("Agg") | |
| 21 | +import matplotlib.pyplot as plt | |
| 22 | +import numpy as np | |
| 23 | + | |
| 24 | +from core import primes_upto, sieve_segment | |
| 25 | + | |
| 26 | +ROOT = Path(__file__).resolve().parent.parent | |
| 27 | +FIGS = ROOT / "paper" / "figs" | |
| 28 | +FIGS.mkdir(exist_ok=True) | |
| 29 | +DATA = ROOT / "data" | |
| 30 | + | |
| 31 | +BLUE, ORANGE, AQUA = "#2a78d6", "#eb6834", "#1baf7a" | |
| 32 | +INK, MUTED = "#0b0b0b", "#52514e" | |
| 33 | + | |
| 34 | +plt.rcParams.update({ | |
| 35 | + "font.size": 8.5, "axes.labelsize": 9, "axes.titlesize": 9.5, | |
| 36 | + "axes.spines.top": False, "axes.spines.right": False, | |
| 37 | + "axes.linewidth": 0.7, "xtick.major.width": 0.7, "ytick.major.width": 0.7, | |
| 38 | + "grid.linewidth": 0.5, "grid.alpha": 0.25, "legend.frameon": False, | |
| 39 | + "figure.dpi": 150, "text.color": INK, "axes.edgecolor": MUTED, | |
| 40 | + "xtick.color": MUTED, "ytick.color": MUTED, "axes.labelcolor": INK, | |
| 41 | +}) | |
| 42 | + | |
| 43 | + | |
| 44 | +def load_checkpoints(): | |
| 45 | + c2 = json.load(open(DATA / "cycle2_c4_4e10_M3U96a.json"))["checkpoints"] | |
| 46 | + c4 = json.load(open(DATA / "cycle4_c4_1e13.json"))["checkpoints"] | |
| 47 | + pts = [(float(x), v["rho1_lnx"], v["rho2_lnx"], v["D_N2_minus_N4"]) | |
| 48 | + for x, v in c2.items() if 1e8 <= float(x) < 1e10] | |
| 49 | + pts += [(float(x), v["rho1_lnx"], v["rho2_lnx"], v["D_N2_minus_N4"]) | |
| 50 | + for x, v in c4.items()] | |
| 51 | + return sorted(pts) | |
| 52 | + | |
| 53 | + | |
| 54 | +def fit_cd(lams, vals): | |
| 55 | + A = np.vstack([np.ones_like(lams), 1 / lams]).T | |
| 56 | + (c, d), *_ = np.linalg.lstsq(A, vals, rcond=None) | |
| 57 | + return c, d | |
| 58 | + | |
| 59 | + | |
| 60 | +def fig_scaling(pts): | |
| 61 | + lam = np.array([log(x) for x, *_ in pts]) | |
| 62 | + v1 = np.array([p[1] for p in pts]) | |
| 63 | + v2 = np.array([p[2] for p in pts]) | |
| 64 | + c1, d1 = fit_cd(lam, v1) | |
| 65 | + c2_, d2_ = fit_cd(lam, v2) | |
| 66 | + model = json.load(open(DATA / "model_c4.json"))["predictions"] | |
| 67 | + ml = np.array([m["lambda"] for m in model.values()]) | |
| 68 | + mv = np.array([m["rho_times_lambda"] for m in model.values()]) | |
| 69 | + | |
| 70 | + fig, (ax, axb) = plt.subplots(1, 2, figsize=(6.4, 2.7), | |
| 71 | + gridspec_kw={"width_ratios": [3, 2]}) | |
| 72 | + ll = np.linspace(18.0, 34.6, 200) | |
| 73 | + # panel (a): data + fits + confirmed predictions | |
| 74 | + ax.plot(ll, c1 + d1 / ll, color=BLUE, lw=1.4) | |
| 75 | + ax.plot(ll, c2_ + d2_ / ll, color=AQUA, lw=1.4) | |
| 76 | + ax.plot(lam, v1, "o", ms=3.4, color=BLUE, mfc="white", mew=1.1) | |
| 77 | + ax.plot(lam, v2, "o", ms=3.4, color=AQUA, mfc="white", mew=1.1) | |
| 78 | + # confirmed out-of-sample predictions (1e12 and 1e13 for lag1; 1e13 for lag2) | |
| 79 | + for X, obs in ((1e12, -0.54756), (1e13, -0.54264)): | |
| 80 | + ax.plot(log(X), obs, "D", ms=5.5, color=ORANGE, mfc="none", mew=1.4) | |
| 81 | + ax.plot(log(1e13), -0.25247, "D", ms=5.5, color=ORANGE, mfc="none", mew=1.4) | |
| 82 | + for X, txt in ((1e14, -0.5387), (1e15, -0.5350)): | |
| 83 | + ax.plot(log(X), txt, "s", ms=4.5, color=MUTED, mfc="none", mew=1.0) | |
| 84 | + ax.annotate("lag 1: $\\rho\\,\\ln x = c + d/\\ln x$", (18.4, -0.525), | |
| 85 | + color=BLUE, fontsize=8) | |
| 86 | + ax.annotate("lag 2", (19.8, -0.272), color=AQUA, fontsize=8) | |
| 87 | + ax.annotate("confirmed predictions", (29.2, -0.572), | |
| 88 | + color=ORANGE, fontsize=7.5, ha="center") | |
| 89 | + ax.annotate("", xy=(log(1e12) + 0.1, -0.553), xytext=(28.2, -0.568), | |
| 90 | + arrowprops=dict(arrowstyle="->", color=ORANGE, lw=0.9)) | |
| 91 | + ax.annotate("", xy=(log(1e13) - 0.1, -0.547), xytext=(30.0, -0.568), | |
| 92 | + arrowprops=dict(arrowstyle="->", color=ORANGE, lw=0.9)) | |
| 93 | + ax.annotate("next targets", (31.6, -0.528), color=MUTED, fontsize=7.5) | |
| 94 | + ax.set_xlabel("$\\ln x$") | |
| 95 | + ax.set_ylabel("$\\rho(x)\\,\\ln x$") | |
| 96 | + ax.grid(True) | |
| 97 | + ax.set_xlim(17.8, 35.4) | |
| 98 | + for X in (1e8, 1e10, 1e12, 1e14): | |
| 99 | + ax.axvline(log(X), color=MUTED, lw=0.4, alpha=0.18) | |
| 100 | + ax.text(log(X), ax.get_ylim()[1], f"$10^{{{int(round(log(X,10)))}}}$", | |
| 101 | + fontsize=6.5, color=MUTED, ha="center", va="bottom") | |
| 102 | + # panel (b): observed vs first-order HL model | |
| 103 | + axb.plot(lam, v1, "o", ms=3.2, color=BLUE, mfc="white", mew=1.0, label="observed (lag 1)") | |
| 104 | + axb.plot(ll, c1 + d1 / ll, color=BLUE, lw=1.2) | |
| 105 | + axb.plot(ml, mv, "--", color=ORANGE, lw=1.4, label="first-order HL model") | |
| 106 | + axb.axhline(0, color=MUTED, lw=0.6) | |
| 107 | + axb.set_xlim(17.8, 95) | |
| 108 | + axb.set_ylim(-0.65, 0.02) | |
| 109 | + axb.set_xlabel("$\\lambda = \\ln x$") | |
| 110 | + axb.grid(True) | |
| 111 | + axb.annotate("observed", (23, -0.52), color=BLUE, fontsize=8) | |
| 112 | + axb.annotate("first-order HL model\n($\\approx 32\\%$ of magnitude)", | |
| 113 | + (40, -0.115), color=ORANGE, fontsize=8) | |
| 114 | + fig.tight_layout() | |
| 115 | + fig.savefig(FIGS / "fig_scaling.pdf") | |
| 116 | + plt.close(fig) | |
| 117 | + print("fig_scaling.pdf") | |
| 118 | + | |
| 119 | + | |
| 120 | +def race_trace(limit=10**10, per_decade=60): | |
| 121 | + """D(x) sampled at ~per_decade log-spaced points per decade, exact.""" | |
| 122 | + targets = np.unique(np.round(np.logspace(6, np.log10(limit), | |
| 123 | + int(per_decade * (np.log10(limit) - 6)))).astype(np.int64)) | |
| 124 | + base = primes_upto(int(limit ** 0.5) + 1) | |
| 125 | + D = 0 | |
| 126 | + out_x, out_d = [], [] | |
| 127 | + prev = None | |
| 128 | + ti = 0 | |
| 129 | + for lo in range(2, limit, 50_000_000): | |
| 130 | + hi = min(lo + 50_000_000, limit) | |
| 131 | + primes = sieve_segment(lo, hi, base) | |
| 132 | + if prev is not None: | |
| 133 | + primes = np.concatenate(([prev], primes)) | |
| 134 | + gaps = np.diff(primes) | |
| 135 | + ends = primes[1:] | |
| 136 | + delta = (gaps == 2).astype(np.int64) - (gaps == 4).astype(np.int64) | |
| 137 | + run = D + np.cumsum(delta) | |
| 138 | + while ti < len(targets) and targets[ti] < hi: | |
| 139 | + j = np.searchsorted(ends, targets[ti], side="right") - 1 | |
| 140 | + out_x.append(int(targets[ti])) | |
| 141 | + out_d.append(int(run[j]) if j >= 0 else D) | |
| 142 | + ti += 1 | |
| 143 | + D = int(run[-1]) | |
| 144 | + prev = primes[-1] | |
| 145 | + return np.array(out_x, dtype=float), np.array(out_d, dtype=float) | |
| 146 | + | |
| 147 | + | |
| 148 | +def fig_race(pts): | |
| 149 | + xs, ds = race_trace() | |
| 150 | + fig, ax = plt.subplots(figsize=(6.4, 2.5)) | |
| 151 | + ax.axhline(0, color=MUTED, lw=0.7) | |
| 152 | + ax.plot(xs, ds, color=BLUE, lw=1.1) | |
| 153 | + cx = [p[0] for p in pts if p[0] >= 2e10] | |
| 154 | + cd = [p[3] for p in pts if p[0] >= 2e10] | |
| 155 | + ax.plot(cx, cd, "D", ms=4.5, color=ORANGE, mfc="none", mew=1.2) | |
| 156 | + ax.set_xscale("log") | |
| 157 | + ax.set_yscale("symlog", linthresh=200) | |
| 158 | + ax.set_xlim(1e6, 2e13) | |
| 159 | + ax.set_ylim(-2.5e5, 6e4) | |
| 160 | + ax.set_xlabel("$x$") | |
| 161 | + ax.set_ylabel("$D(x) = N(2,x) - N(4,x)$") | |
| 162 | + ax.grid(True, which="major") | |
| 163 | + ax.annotate("first tie after $10^6$\n$x = 80{,}966{,}861$", | |
| 164 | + xy=(8.1e7, 0), xytext=(2.5e6, -2600), fontsize=7.5, color=INK, | |
| 165 | + arrowprops=dict(arrowstyle="->", color=MUTED, lw=0.8)) | |
| 166 | + ax.annotate("fine trace (every gap, sampled)", (2.6e6, 6000), color=BLUE, fontsize=8) | |
| 167 | + ax.annotate("campaign checkpoints", (2.2e11, -12000), color=ORANGE, fontsize=8) | |
| 168 | + fig.tight_layout() | |
| 169 | + fig.savefig(FIGS / "fig_race.pdf") | |
| 170 | + plt.close(fig) | |
| 171 | + print("fig_race.pdf") | |
| 172 | + | |
| 173 | + | |
| 174 | +def fig_hist(): | |
| 175 | + rows = np.loadtxt(DATA / "gap_histogram_1e8.csv", delimiter=",", skiprows=1) | |
| 176 | + g, cnt = rows[:, 0].astype(int), rows[:, 1] | |
| 177 | + m = (g >= 2) & (g <= 150) & (g % 2 == 0) | |
| 178 | + g, cnt = g[m], cnt[m] | |
| 179 | + six = g % 6 == 0 | |
| 180 | + fig, ax = plt.subplots(figsize=(6.4, 2.5)) | |
| 181 | + ax.vlines(g[~six], 1, cnt[~six], color=MUTED, lw=1.6, alpha=0.75) | |
| 182 | + ax.vlines(g[six], 1, cnt[six], color=BLUE, lw=1.9) | |
| 183 | + ax.set_yscale("log") | |
| 184 | + ax.set_xlabel("gap $g$") | |
| 185 | + ax.set_ylabel("$N(g, 10^8)$") | |
| 186 | + ax.grid(True, axis="y") | |
| 187 | + ax.annotate("multiples of 6", (66, 2.2e5), color=BLUE, fontsize=8.5) | |
| 188 | + ax.annotate("other even gaps", (80, 1.1e4), color=MUTED, fontsize=8.5) | |
| 189 | + ax.annotate("jumping champion $g = 6$", xy=(6, 8.8e5), xytext=(16, 1.5e6), | |
| 190 | + fontsize=7.5, color=INK, | |
| 191 | + arrowprops=dict(arrowstyle="->", color=MUTED, lw=0.8)) | |
| 192 | + ax.set_xlim(0, 152) | |
| 193 | + ax.set_ylim(1, 4e6) | |
| 194 | + fig.tight_layout() | |
| 195 | + fig.savefig(FIGS / "fig_hist.pdf") | |
| 196 | + plt.close(fig) | |
| 197 | + print("fig_hist.pdf") | |
| 198 | + | |
| 199 | + | |
| 200 | +def fig_certificate(): | |
| 201 | + cert = json.load(open(ROOT / "certs" / "desert_certificate.json")) | |
| 202 | + interior = cert["interior_certificates"] | |
| 203 | + cov_x, cov_y, tf_x, tf_y, mr_x = [], [], [], [], [] | |
| 204 | + for i_str, c in interior.items(): | |
| 205 | + i = int(i_str) | |
| 206 | + if c["type"] == "covering_factor": | |
| 207 | + cov_x.append(i); cov_y.append(c["q"]) | |
| 208 | + elif c["type"] == "trial_factor": | |
| 209 | + tf_x.append(i); tf_y.append(c["q"]) | |
| 210 | + else: | |
| 211 | + mr_x.append(i) | |
| 212 | + fig, ax = plt.subplots(figsize=(6.4, 2.6)) | |
| 213 | + ax.plot(cov_x, cov_y, "o", ms=2.6, color=BLUE, mew=0) | |
| 214 | + ax.plot(tf_x, tf_y, "D", ms=4.5, color=ORANGE, mfc="none", mew=1.2) | |
| 215 | + top = max(tf_y) * 3 | |
| 216 | + ax.plot(mr_x, [top] * len(mr_x), "s", ms=5, color=AQUA, mfc="none", mew=1.3) | |
| 217 | + ax.set_yscale("log") | |
| 218 | + ax.set_xlabel("position $i$ in the desert ($N + i$, $1 \\leq i \\leq 259$)") | |
| 219 | + ax.set_ylabel("certifying prime factor $q_i$") | |
| 220 | + ax.grid(True, axis="y") | |
| 221 | + ax.annotate("covering system (proven for every shift $t$)", (4, 230), | |
| 222 | + color=BLUE, fontsize=8) | |
| 223 | + ax.annotate("holes: explicit trial factor", (138, 2.6e4), | |
| 224 | + color=ORANGE, fontsize=8) | |
| 225 | + ax.annotate("hole: strong MR witness", (mr_x[0] + 8, top * 0.75), | |
| 226 | + color=AQUA, fontsize=8) | |
| 227 | + ax.set_xlim(0, 260) | |
| 228 | + fig.tight_layout() | |
| 229 | + fig.savefig(FIGS / "fig_certificate.pdf") | |
| 230 | + plt.close(fig) | |
| 231 | + print("fig_certificate.pdf") | |
| 232 | + | |
| 233 | + | |
| 234 | +if __name__ == "__main__": | |
| 235 | + pts = load_checkpoints() | |
| 236 | + fig_scaling(pts) | |
| 237 | + fig_hist() | |
| 238 | + fig_certificate() | |
| 239 | + fig_race(pts) # last: recomputes a 1e10 scan (~40 s) | |
modified
src/merge_c4.py
+2 −1
@@ -21,7 +21,8 @@ import numpy as np | ||
| 21 | 21 | DATA = Path(__file__).resolve().parent.parent / "data" |
| 22 | 22 | CHECKPOINTS = [10**10, 2 * 10**10, 4 * 10**10, |
| 23 | 23 | 10**11, 2 * 10**11, 4 * 10**11, 10**12, |
| 24 | − 2 * 10**12, 4 * 10**12, 10**13] | |
| 24 | + 2 * 10**12, 4 * 10**12, 10**13, | |
| 25 | + 2 * 10**13, 4 * 10**13, 10**14] | |
| 25 | 26 | |
| 26 | 27 | |
| 27 | 28 | def pearson(S): |
modified
src/worker_c4.py
+6 −1
@@ -20,8 +20,9 @@ import numpy as np | ||
| 20 | 20 | |
| 21 | 21 | from core import primes_upto, sieve_segment |
| 22 | 22 | |
| 23 | +import os | |
| 23 | 24 | BUFFER = 5000 |
| 24 | −SEG = 50_000_000 | |
| 25 | +SEG = int(os.environ.get("C4_SEG", 50_000_000)) | |
| 25 | 26 | |
| 26 | 27 | |
| 27 | 28 | def run(lo, hi, out_path, sieve_fn=None, tag=""): |
@@ -101,6 +102,10 @@ def run(lo, hi, out_path, sieve_fn=None, tag=""): | ||
| 101 | 102 | def main(): |
| 102 | 103 | lo, hi = int(float(sys.argv[1])), int(float(sys.argv[2])) |
| 103 | 104 | out_path = sys.argv[3] |
| 105 | + from pathlib import Path as _P | |
| 106 | + if _P(out_path).exists(): | |
| 107 | + print(f"skip existing {out_path}", flush=True) | |
| 108 | + return | |
| 104 | 109 | if len(sys.argv) > 4 and sys.argv[4] == "gpu": |
| 105 | 110 | from gpu_sieve import sieve_segment_gpu |
| 106 | 111 | run(lo, hi, out_path, sieve_fn=sieve_segment_gpu, tag="/gpu") |
| 107 | 112 | |