SPB Git

spb/localvm-research Public License

Running LLMs larger than memory on a consumer Mac — falsification-driven research: margin-gated deferred refinement, out-of-core verification on Apple Silicon. TR-01 published.

Python 63.2% JavaScript 23.5% CSS 11.8% Shell 0.9% Makefile 0.5%

Phase 4: candidate ranking + web platform deployed to www.localvm.dev

4 candidates selected (Progressive-Residual Runtime preferred; Gate-Zero
falsifier battery runs first). Platform live on M3U96b via PM2+ngrok;
sync-content.sh filter-order fix (8 GiB testfile exclusion).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Simon-Pierre Boucher committed 5 h ago (Aug 12, 2026) parent 9b94283

Showing 3 changed files with +491 and −2

modified research/LOG.md +36 −0
@@ -122,3 +122,39 @@ Format per entry: date/time (local, with timezone) · question · experiment ·
122 122 G01/G04/G05/G15/G16, rest of G24, gates G03.
123 123 - **Interpretation:** The experiment plan now writes itself from gap dependencies.
124 124 - **Decision:** Launch Phase 4 ranking; then implement expG first.
125 +
126 +---
127 +
128 +## 2026-08-12 11:40 EDT — Phase 4 complete: candidate ranking
129 +
130 +- **Question:** Which of the 24 gaps deserve prototyping effort, and in what composition?
131 +- **Experiment:** `research/candidate_ranking.md` — all 24 gaps scored 1–10 on the ten
132 + charter §7 axes with per-axis written reasoning; summary table; totals span 51–78
133 + (top: G18 78, G02 77, G17/G23 76, G09 75, G24 74; bottom: G06 51).
134 +- **Result:** Four candidates selected (three compositions + one falsifier battery):
135 + C1 Progressive-Residual Runtime (G01+G02+G04, G03 kernel, G05 gate) — criteria B/C/D/E;
136 + C2 Amortized Verification Sweeps (G17+G18 on G08 MRU floor) — criteria A/B, exact, also
137 + the honest baseline; C3 MoE Residency Engine (G09+G12+G19, G10 warm tier) — criteria A/F;
138 + C4 Gate Zero contrarian battery (G23+G24+G22+expN) — runs first, gates 17/24 gaps.
139 + Promote/kill thresholds fixed numerically per candidate (AUROC 0.8/0.65–0.85 bands;
140 + ≤20% residual bytes + ≥50% reuse; acceptance ≥4–5; ARC ≥5 pts; warm-fault 10%/25%).
141 +- **Interpretation:** Totals reward safety; selection weighted Sig + cluster synergy +
142 + the expH substrate contract. Preferred single candidate: C1 — the only mechanism that
143 + turns `model size → hardware` into `working set → hardware` (charter §20).
144 +- **Decision:** Proceed to Phase 5 infrastructure, then expG run matrix (arbitrates
145 + G02/G17/G23/G24) before any candidate-specific engineering.
146 +
147 +---
148 +
149 +## 2026-08-12 09:40 EDT — Phase 4 complete + public research platform deployed
150 +
151 +- **Question:** Which candidates should be prototyped, and how is the project made auditable publicly?
152 +- **Experiment:** n/a (ranking synthesis + engineering).
153 +- **Result:** `research/candidate_ranking.md` — all 24 gaps scored on 10 axes (totals 51–78,
154 + differentiated). Selected: C1 Progressive-Residual Runtime (G01+G02+G04, preferred),
155 + C2 Amortized Verification Sweeps (G17+G18, honest baseline), C3 MoE Residency Engine
156 + (G09+G12+G19), C4 Gate-Zero falsifier battery (G22–G24 first, gates 17/24 gaps).
157 + Public platform live at https://www.localvm.dev (M3U96b, PM2 + ngrok) serving the full
158 + paper trail: docs, notes, experiments, raw results, code. Content snapshot at commit 9b94283.
159 +- **Interpretation:** Gate-Zero battery (expG-centric) is the rational first experiment.
160 +- **Decision:** Next session: implement expG (joint margin×agreement matrix) on a small model.
added research/candidate_ranking.md +450 −0
@@ -0,0 +1,450 @@
1 +---
2 +project: localvm-research
3 +document: research/candidate_ranking
4 +author: Simon-Pierre Boucher
5 +contact: contact@spboucher.ai
6 +created: 2026-08-12
7 +status: draft
8 +---
9 +
10 +# Candidate Ranking: Phase 4 Deliverable
11 +
12 +Ranks all 24 gaps from `research/research_gaps.md` on the ten charter §7 axes, scores 1–10:
13 +**N** Novelty · **P** Technical plausibility · **Mem** memory reduction · **BW** bandwidth
14 +reduction · **Q** quality retention · **IC** implementation complexity (10 = simplest) ·
15 +**HW** Apple Silicon practicality (Metal/MLX kernel feasibility, unified-memory fit) ·
16 +**PT** post-training compatibility · **Sig** research significance · **Risk** (10 = lowest).
17 +Totals are unweighted sums for orientation only; selection (§3) weights significance, cluster
18 +synergy, and the measured substrate. Contrarian gaps G22–G24 are scored **as experiments**
19 +(cheap falsifiers): their Mem/BW are honest 1s; their value shows in Sig/IC/Risk. Binding
20 +expH facts used throughout: **≤650 MB/token SSD budget at 5 tok/s** (≤1.3 GB/token
21 +theoretical), the **≥256 KiB @ QD≥4–8, 16 KB-aligned fetch contract** (4 KiB QD1 = 67 MB/s,
22 +a 200× cliff), the **shift/mask-cheap decode rule** (i-quant LUT penalty), free compute/I-O
23 +overlap (<5% contention), and the ~35:1 RAM:SSD ratio.
24 +
25 +## 1. Scorecards
26 +
27 +### A. Progressive representations and refinement
28 +#### G01 — ResidualPager (total 69)
29 +- **N 8** — Paging residual planes from NVMe is SOTA §7.3.1's verified-unoccupied intersection on any OS; the encodings themselves (Any-Precision, BitStack, RRQ) are prior art, which caps this below 9.
30 +- **P 6** — Successive-refinement theory and BitDelta's ~1-bit corrections are solid ground, but the load-bearing assumption — refined-set temporal stability — is unmeasured (expB), and BitStack itself warns that reconstruction overhead slows inference even fully resident.
31 +- **Mem 9** — A 2–3-bit base of a 70B model is 17.5–26 GB, resident on the 48 GB target; RAM bounds *base* size, not model size — the strongest direct residency decoupling of any non-contrarian gap.
32 +- **BW 7** — Bytes/token = base RAM reads + fetched residuals; good only if the refined set is stable and reused across tokens (the 1–2% non-overlapped ceiling), with no cross-token amortization of the kind G17 gets.
33 +- **Q 7** — Full quality is recoverable in principle (planes sum to the original), but the practical operating point is whatever residual budget fits under 650 MB/token — quality is hostage to expD's recovery curve.
34 +- **IC 4** — Needs a progressive encoder, block index, demand pager, and refinement application path — substantial new machinery before a single kernel is written.
35 +- **HW 6** — Zero-copy unified memory and free prefetch overlap are genuine Apple-specific advantages, but bitplane decode is gated by G05; the affine-delta fallback is materially weaker.
36 +- **PT 9** — Pure PTQ plus calibration; no retraining anywhere.
37 +- **Sig 9** — Directly instantiates the charter §2 decoupling and criteria B/C/E; if it works it is the project's canonical representation.
38 +- **Risk 4** — Dies cleanly if expD+expB show >20% of residual bytes touched per token with <50% temporal reuse; the 35:1 ratio gives little forgiveness.
39 +
40 +#### G02 — MarginGate (total 77)
41 +- **N 9** — SOTA §7.3.2's narrow-but-real gap after excluding QSpec/BiLD/CALM near-misses: nobody gates *storage-tier* residual fetch on decision stability with a correctness story, and no paper in §5 or §6 optimizes or even reports bytes/token.
42 +- **P 7** — Margins demonstrably carry signal (the entire early-exit line; QuantSpec >90% acceptance; ρ=0.82 learned error features), but the confidently-wrong quadrant (G23) is exactly the unmeasured joint distribution that could void it.
43 +- **Mem 8** — Inherits G01's resident low-bit base and adds only a probe; slightly below G01 because it does not itself define the representation.
44 +- **BW 9** — Bytes/token collapses to base-read + (escalation rate × residual fetch): at the measured 5–15% knife-edge prior, the largest expected per-token byte cut of any per-token scheme in the document.
45 +- **Q 7** — CALM-style calibrated risk or exact rejection variants exist and are proven, but the lossy-gated form ships precisely the errors that slip the gate — quality rides on the AUROC.
46 +- **IC 5** — The gate itself is trivial; calibration machinery, rollback, and integration with a pager that must exist first make it moderate.
47 +- **HW 8** — Prototype needs only MLX quantized casts, logging, and control flow — no new kernels before the pager stage.
48 +- **PT 9** — Calibration-only; a small learned probe is explicitly charter-permitted (§14).
49 +- **Sig 10** — If the gate works, the optimization objective of local inference changes from resident-size to decision-gated bytes/token — this is the thesis of the project stated as a mechanism.
50 +- **Risk 5** — expG is cheap and decisive before any engineering; the failure mode is well-defined (AUROC <0.8 or escalation >40% on any major domain) rather than diffuse.
51 +
52 +#### G03 — DeltaKernel (total 67)
53 +- **N 8** — ΔY = (ΔW)X against cached activations is an unpublished primitive (SOTA §7.3.3): GMRES-IR transposed to transformer inference, never written down.
54 +- **P 7** — Exact for linear ops by linearity; the two known threats (attention nonlinearity, Metal gather cost) are bounded and enumerable rather than open-ended.
55 +- **Mem 4** — Saves no residency itself; the activation cache is tens of MB — this is an enabling multiplier, not a reduction mechanism.
56 +- **BW 7** — Makes refinement cost proportional to residual size instead of layer size; without it every escalation or verification doubles compute and halves the whole family's economics.
57 +- **Q 9** — Exact per linear layer by construction; downstream nonlinearity handling is a mapped recompute (expF), not an approximation of unknown size.
58 +- **IC 5** — One fused Metal kernel plus activation-caching hooks — contained, but real kernel work.
59 +- **HW 5** — Block-sparse gather matvec is exactly the access pattern Apple GPUs punish (Endor / i-quant lesson transposed); may need the G20 CPU escape hatch to survive.
60 +- **PT 10** — A pure runtime primitive; the model is untouched.
61 +- **Sig 7** — An enabling primitive, not a regime; its significance is inherited from G02/G17.
62 +- **Risk 5** — Kill is crisp and cheap (expE_delta: ≥2× over recompute at 10% density, ≤1e-3 relative error), and failure has a named fallback (recompute + G20).
63 +
64 +#### G04 — RDLayout (total 72)
65 +- **N 6** — EBCOT-style Lagrangian truncation is 30-year-old machinery and BitStack is adjacent; the novelty is compiling it against logit-KL distortion and the *measured* SSD contract, which is real but incremental.
66 +- **P 8** — Sensitivity-weighted allocation beats uniform everywhere it has been tried (SqueezeLLM, imatrix, LASER's improve-when-truncated matrices); the optimizer is mature and the charter's compile stage budgets unlimited offline time.
67 +- **Mem 7** — A static-but-nonuniform operating point strictly better than uniform bpw at matched bytes, and it is the substrate every dynamic gap pages from.
68 +- **BW 6** — Compile-time only: bytes at matched quality improve, but every loaded byte is still read each token unless a dynamic gap sits on top.
69 +- **Q 8** — Better-than-uniform at matched bytes is the design objective, and per-block KL attribution errors degrade the gain, not correctness.
70 +- **IC 6** — An offline compiler with no runtime component; the work is bounded and parallelizable.
71 +- **HW 8** — The output *is* the 16 KB-aligned, ≥256 KiB-bundled layout the contract demands; the decode path can stay affine (shift/mask-cheap) if G05 fails.
72 +- **PT 10** — Calibration-set only.
73 +- **Sig 6** — Real but possibly incremental: if imatrix-style mixing already captures most of the nonuniform gain, this is a <15%-relative-KL footnote.
74 +- **Risk 7** — Worst case is a null result at modest cost; nothing downstream breaks since pagers fall back to existing formats.
75 +
76 +#### G05 — MetalTrellis (total 63)
77 +- **N 7** — Zero Metal implementations of bitplane, MSB-sliced, or trellis decode exist (SOTA §7.3.5); the formats themselves are published CUDA work.
78 +- **P 6** — Shift/mask-cheap by construction (passes the decode rule on paper), but bit-transpose gather coalescing and trellis sequentiality vs SIMD-group width are genuine unknowns — which is why it is an experiment.
79 +- **Mem 6** — Unlocks the 2–2.5 bpw frontier on Metal (~8× vs FP16) as a component of G01/G04, never alone.
80 +- **BW 7** — QTIP proves >80% of peak bandwidth is reachable with ~2 instr/weight on CUDA; success here means byte savings finally translate to tokens/s on Apple GPUs.
81 +- **Q 7** — Format-neutral: quality equals the underlying encodings' published quality iff the kernel runs at speed.
82 +- **IC 4** — Three hand-written Metal matvecs with correctness and throughput accounting — the most kernel-intensive micro in the plan.
83 +- **HW 5** — HW practicality is the axis *under test*; the i-quant precedent is explicitly unfavorable and the honest prior is uncertain.
84 +- **PT 10** — Kernels only; nothing touches the model.
85 +- **Sig 7** — A fan-out dependency: gates whether G01/G04 use bitplanes or degrade to affine-delta residuals.
86 +- **Risk 4** — Binary outcome; a loss permanently narrows the representation design space (though the affine fallback exists).
87 +
88 +### B. Sparsity as a paging policy
89 +#### G06 — TealPager (total 51)
90 +- **N 6** — TEAL/CATS/GRIFFIN exist with published numbers; the paging framing and the bytes/token measurement are new, the mechanism is not.
91 +- **P 6** — Thresholds are calibration-only and GRIFFIN's per-sequence granularity is inherently prefetch-friendly, but SwiGLU mask working-set statistics are completely unmeasured (all published stats are ReLU-era).
92 +- **Mem 5** — 40–50% model-wide sparsity is ~2× — rescues a 20–30% RAM deficit, nothing more.
93 +- **BW 5** — The 2× FFN cut erodes further if masks churn: Jaccard(t,t+1) <0.6 makes paging traffic exceed the dense-read baseline.
94 +- **Q 4** — Sirius showed contextual sparsity specifically damages reasoning/GSM8K; the 99%-retention claims are perplexity mirages, and this gap starts from that liability.
95 +- **IC 4** — Metal gather kernel + threshold calibration + paging integration; three subsystems for a 2× ceiling.
96 +- **HW 4** — A 50%-sparse gather matvec on a GPU with no sparse tensor cores may run *slower* than dense — a kill condition entirely independent of quality.
97 +- **PT 8** — Calibration thresholds only, but applicability is limited to SwiGLU dense models with enough safe sparsity.
98 +- **Sig 6** — The only training-free selectivity mechanism for dense stock models — important if true, but with a modest ceiling even then.
99 +- **Risk 3** — Two independent cliffs (reasoning quality; gather kernel), either one fatal.
100 +
101 +#### G07 — WandaTiers (total 68)
102 +- **N 7** — Importance-scores-as-reversible-residency with dense correction is SOTA §7.3.6, unoccupied; every ingredient (Wanda, Endor format, Sirius loop) is published separately.
103 +- **P 7** — Demoted weights are by construction those whose absence changes outputs least, and Sirius bounds the correction bill at ~11% of tokens.
104 +- **Mem 6** — 20–50% demotion runs 1.3–1.5× over-RAM dense models; not a 10× mechanism, but the cheapest real one.
105 +- **BW 6** — Bytes/token = cold-block demand, favorable only if demand is rare rather than diffuse; diffuse demand lands in the 67 MB/s regime and kills it.
106 +- **Q 6** — Global-average scores can miss load-bearing weights ("Accuracy is Not All You Need": 25%+ flips for pruning-type interventions); the correction loop mitigates but does not erase this.
107 +- **IC 7** — The cheapest oversubscription route in the document: no re-encoding, no predictors, no new format — placement plus a dense-readable bitmap.
108 +- **HW 8** — Endor-style dense-readable demoted blocks avoid the gather penalty; serve-miss-as-zero is trivial on unified memory.
109 +- **PT 10** — One-shot post-training statistics; fully reversible.
110 +- **Sig 6** — A working version is a solid systems result and a useful floor, not a new regime.
111 +- **Risk 5** — Single crisp kill (cold-block demand >5% of demoted bytes/token at 35% demotion), measurable in expA-blocked before any runtime exists.
112 +
113 +### C. Cache, replacement, and residency policy
114 +#### G08 — MRURing (total 66)
115 +- **N 6** — MRU-for-cyclic-scans is DBMIN 1985; never applied to the dense-decode weight scan, so the novelty is purely in the (real) application.
116 +- **P 8** — The access pattern is fully known (static layer order), the math is 40 years old, and the 0.025 tok/s MLX-mmap disaster is the measured counterfactual it fixes.
117 +- **Mem 4** — Rescues only slightly-over-budget models: the re-read floor caps useful overflow at ~650 MB (5 tok/s) to ~6.5 GB (0.5 tok/s).
118 +- **BW 6** — Achieves the theoretical floor — re-read exactly the overflow X GB/token — where kernel LRU re-reads everything.
119 +- **Q 10** — Bit-exact; a pure replacement-policy change.
120 +- **IC 5** — The simulator is trivial; the real ring needs explicitly managed wired buffers under the iogpu limit, i.e., engineering against the kernel-panic cliff.
121 +- **HW 6** — macOS gives no UBC eviction control, forcing the wired-arena route; expH confirms the re-read stream itself runs at ~13 GB/s with free overlap.
122 +- **PT 10** — No model change of any kind.
123 +- **Sig 5** — Narrow applicability caps significance, but it is also the enabling floor under G18's sweeps.
124 +- **Risk 6** — The simulator answers most of the question for free; kill is a measured <1.5× gain over kernel-LRU mmap at 10–20% overflow.
125 +
126 +#### G09 — ArcExperts (total 75)
127 +- **N 5** — Every ingredient is published or an open llama.cpp feature request (#20757); composition novelty only, and upstream may merge layout+QD first.
128 +- **P 9** — 2.23× from layout alone is already measured on a Mac (#18758), MoE temporal locality is confirmed, and ARC's superiority on mixed workloads is textbook storage systems.
129 +- **Mem 6** — Runs Qwen3-30B/80B-class MoE on 16–48 GB Macs — capacity via the selectivity the router provides for free.
130 +- **BW 7** — Reads/token already shown to drop 1418→370 from layout; ARC + QD≥8 explicit reads + purgeable-warm demotion stack multiplicatively on that.
131 +- **Q 10** — Exact: misses block briefly, nothing is approximated.
132 +- **IC 6** — Cache + bundle compiler + explicit reads inside an existing runtime; zero new math.
133 +- **HW 9** — Expert bundles (MBs) natively satisfy the ≥256 KiB QD≥8 contract; the zero-copy Metal path ships today.
134 +- **PT 10** — Stock checkpoints, stock router.
135 +- **Sig 5** — High deployment value with a real "engineering, not research" risk if policy headroom over LRU is <5 hit-rate points at every cache size.
136 +- **Risk 8** — Worst case still yields a useful runtime plus a publishable routing-trace study.
137 +
138 +#### G10 — PurgeableWarm (total 57)
139 +- **N 5** — Purgeable memory is a shipped, decade-old macOS/iOS API; pointing it at weights is new but small.
140 +- **P 6** — The semantics fit (a purge is exactly a cache miss), but per-access volatile-state transitions may be unaffordable and the OS purge order is model-blind.
141 +- **Mem 2** — Robustness and coexistence with user apps, not reduction.
142 +- **BW 2** — None directly.
143 +- **Q 9** — Purge → miss → refetch at 13 GB/s; quality untouched, latency jitter only.
144 +- **IC 6** — An MTLHeap wrapper plus pressure testing; contained.
145 +- **HW 7** — A macOS-unique mechanism — one of the few gaps discrete-GPU prior work literally cannot express.
146 +- **PT 10** — Runtime only.
147 +- **Sig 4** — Valuable hygiene for every caching gap (converts panic risk into graceful degradation), not a result by itself.
148 +- **Risk 6** — expJ is a cheap probe; failure just means fixed self-managed budgets.
149 +
150 +#### G11 — NeverBlock (total 55)
151 +- **N 8** — Anti-caching's "never stall on a miss, repair later" has never been formulated as an architectural invariant for neural weights.
152 +- **P 5** — QEP's near-exponential error compounding says a layer-10 miss contaminates the whole downstream token; repair-later may collapse into deferred blocking with extra bookkeeping.
153 +- **Mem 3** — No direct reduction; it buys decoupling of token cadence from storage latency.
154 +- **BW 4** — Converts I/O jitter into bounded quality jitter; average bytes unchanged.
155 +- **Q 5** — Real, bounded quality noise under miss bursts, entirely dependent on expF's layer-tolerance map.
156 +- **IC 4** — Correct deferred-repair bookkeeping across layers and in-flight tokens is subtle, and it presupposes a G01-style representation for graceful miss values.
157 +- **HW 7** — Unified memory makes the substitute value (resident base plane) free to reach; the repair path is G03's kernel.
158 +- **PT 9** — A runtime invariant over an already-permitted representation.
159 +- **Sig 7** — P99, not mean, is the real interactive constraint (PowerInfer-2: +40.9% P99); a definitive answer matters for any shipped runtime.
160 +- **Risk 3** — Plausibly self-killing via error compounding; mercifully, the expF injection test is cheap.
161 +
162 +### D. Prefetch and prediction
163 +#### G12 — HiddenPrefetch (total 57)
164 +- **N 6** — Eliseev–Mazur lookahead and ShadowLLM early probes exist; extending to N-layer-ahead SSD prefetch with wasted-byte penalties is incremental-but-real (Pythia's objective, never for weights).
165 +- **P 6** — One early probe demonstrably shadows all layers, but the blocks that matter for prefetch (non-resident ones) are by definition the distribution's tail, where prediction is hardest.
166 +- **Mem 2** — None; it *spends* RAM on apparatus (PowerInfer-2's predictors cost 2.6 GB at 47B — the cautionary number).
167 +- **BW 6** — Does not reduce bytes; it hides latency — which expH says is fully hideable given ~1–5 ms of lead time.
168 +- **Q 8** — Mispredictions cost bandwidth, never correctness.
169 +- **IC 4** — Trace pipeline + probe training + integration into a pager that does not exist yet.
170 +- **HW 7** — Logistic probes are microseconds on AMX and overlap is measured-free; footprint is the only HW concern.
171 +- **PT 8** — Small trained probes are charter-permitted but do require a training loop and per-model calibration.
172 +- **Sig 6** — The enabler that determines whether G01/G06/G07 meet their overlap assumptions; inherits significance from them.
173 +- **Risk 4** — Recall <70% at ≤2× overfetch on the non-resident subset is a live possibility, and the probe+metadata budget cap (5% of RAM) is tight.
174 +
175 +#### G13 — DomainCache (total 66)
176 +- **N 7** — Session-granular weight working sets are unmeasured and unbuilt; adjacent evidence (domain-correlated experts, LoRA subspaces) is suggestive, not demonstrative.
177 +- **P 5** — The rival hypothesis — token-level syntactic dominance with a thin idiosyncratic tail — is equally plausible a priori; expC exists precisely because nobody knows.
178 +- **Mem 6** — A domain profile bounds the resident set for an entire session if the structure exists.
179 +- **BW 8** — Session-granularity amortizes fetches over thousands of tokens, making the 35:1 ratio irrelevant — the most favorable bytes/token geometry in the document, conditional on locality being real.
180 +- **Q 7** — Cache *warming* is quality-neutral; quality is at stake only in exclusionary (G07-style) use.
181 +- **IC 6** — Offline profiles + session-start warming (seconds of sequential 13 GB/s reads); light runtime.
182 +- **HW 8** — Sequential warm reads are the SSD's best case, and it composes directly with the G08/G09 cache machinery.
183 +- **PT 9** — Profiles are offline statistics over a stock model.
184 +- **Sig 7** — Would establish semantic structure in weight usage — a scientifically interesting result independent of its systems value.
185 +- **Risk 3** — One unmeasured premise carries everything; a within-vs-cross-domain Jaccard delta <0.10 kills it outright.
186 +
187 +### E. Decomposition and sharing
188 +#### G14 — TrunkDelta (total 56)
189 +- **N 7** — Assigning a sharing decomposition to the memory hierarchy (trunk wired, deltas streamed in layer order) is unoccupied; the decompositions themselves are published.
190 +- **P 4** — Every published sharing result needs light retraining, and BitDelta's 1-bit evidence is for finetune-vs-base pairs, not layer-vs-layer within one model — the load-bearing transfer is unproven.
191 +- **Mem 7** — The trunk amortizes across L layers (read L times, stored once); if deltas quantize to 1–2 bits, resident cost collapses.
192 +- **BW 7** — Delta streaming is perfectly sequential in static layer order — the one access pattern the SSD serves at full 13 GB/s with zero prediction machinery.
193 +- **Q 4** — Calibration-only trunk extraction plausibly degrades enough that deltas must carry most of the information, at which point nothing is saved.
194 +- **IC 4** — ALS trunk fitting + delta encoder + streaming runtime; moderate-heavy across offline and online components.
195 +- **HW 7** — Dense small GEMMs plus sign-mask decode are Metal-friendly and shift/mask-cheap.
196 +- **PT 6** — The charter permits calibration passes, but all published quality points used 30–40M-token retraining; the honest score reflects that dependence.
197 +- **Sig 7** — Layer-vs-layer delta compressibility is a real scientific question whose answer is valuable either way.
198 +- **Risk 3** — The expK kill (deltas ≥3 bits/param at a 30% trunk) is quite plausible given prior art's retraining dependence.
199 +
200 +#### G15 — SpectralSplit (total 63)
201 +- **N 6** — CALDERA and the SVD family are directly adjacent; tiering the two terms (resident low-rank, paged residual) is a genuinely new but small step.
202 +- **P 5** — Weights are not globally low-rank (SOTA §2.8 headline); the base may be so poor that residual demand is near-universal per token.
203 +- **Mem 6** — A rank-256 hot path is ~12% of dense bytes at 4096², plus whatever residual tier is resident.
204 +- **BW 6** — Same demand structure as G01 but with a worse-conditioned base and double the kernel launches.
205 +- **Q 5** — The split is exact by construction, but the operating curve is plausibly dominated by G01's precision split at equal bytes — the explicit kill condition.
206 +- **IC 7** — Both components run on existing fast MLX kernels today; no custom Metal work gates a prototype — the cheapest Group A/E prototype to stand up.
207 +- **HW 9** — Dense small GEMMs are the one decomposition Metal already runs at full speed.
208 +- **PT 9** — SVD plus quantization, calibration-only.
209 +- **Sig 6** — Mostly valuable as the controlled head-to-head that decides precision-split vs spectral-split for the progressive family.
210 +- **Risk 4** — The kill likely triggers on the base-alone-KL arm (>1.0 nats), but the experiment is nearly free given the expD ladder.
211 +
212 +#### G16 — SeedSynth (total 62)
213 +- **N 7** — SeedLM exists (Apple-authored, FPGA-validated); the Metal question and the seeds-as-nearly-free-prefetch composition are unexplored.
214 +- **P 5** — LFSR generation is shift/XOR-only, but QTIP's ~2 instr/weight ceiling is demanding, and compile-time coefficient search at 70B scale is heavy.
215 +- **Mem 7** — Seeds + coefficients shrink resident and stored bytes together, ~10× per block against 4-bit.
216 +- **BW 9** — Attacks bytes/token *below* the information-theoretic floor of stored-bits formats — the strongest pure-bandwidth story in the document if the kernel wins.
217 +- **Q 4** — Hard-capped at 4-bit parity with no path to exactness: it can never serve as a verification tier, only as a base.
218 +- **IC 5** — One fused generate+matvec Metal kernel for the probe; the full compiler is heavy but deferrable.
219 +- **HW 6** — Targets the real ALU:BW imbalance at batch-1 decode, but the i-quant precedent shows Apple GPUs punish decode work; genuinely uncertain.
220 +- **PT 9** — Data-free re-encoding of a stock checkpoint.
221 +- **Sig 7** — "Weights need not be stored, only recoverable" is the project's thesis in its purest form; even a negative Metal result is a useful boundary.
222 +- **Risk 3** — Single-kernel kill at <80% of the 4-bit baseline, plus reproduction risk of SeedLM's parity claims on SwiGLU models.
223 +
224 +### F. Scheduling, speculation, and amortization
225 +#### G17 — BlockVerify (total 76)
226 +- **N 8** — SpecExec's economics applied to residual planes with rollback is exactly the "block-deferred verification policy" SOTA §8 lists as open; QuantSpec verifies every block from *resident* 16-bit weights, never from storage.
227 +- **P 7** — The rejection-sampling proof carries over unchanged (same support, base=draft), and 4-bit acceptance priors are measured (>90%); acceptance at the 2–3-bit bases real oversubscription requires is the open coin-flip.
228 +- **Mem 8** — Same resident low-bit base as G01; residuals never resident, read once per accepted block.
229 +- **BW 9** — Divides residual bytes/token by the acceptance length, relaxing the brutal 1–2% per-token ceiling to 10–20% of weight bytes per sweep — and sweeps are sequential-by-layer, the SSD's best pattern.
230 +- **Q 9** — Exact by construction under the standard rejection rule; quality risk lives only in throughput, not in outputs.
231 +- **IC 5** — Drafting loop + rollback + residual-sweep orchestration; mitigated by MLX already shipping speculative verification kernels (ReDrafter) on Metal.
232 +- **HW 7** — Sequential residual sweeps at QD≥8 are contract-native; the open cost is Metal batch-verification width sharing unified-memory bandwidth.
233 +- **PT 9** — Base and residual encodings are PTQ; no training anywhere.
234 +- **Sig 9** — If mean acceptance ≥4–7 at a 3-bit base, this makes 1.5–2× oversubscription interactive *with exact outputs* — criteria B and D simultaneously.
235 +- **Risk 5** — Acceptance at low-bit bases is unmeasured (the G24 exposure), but the expG acceptance-length variant measures it before any engineering; rollback economics degrade it gracefully.
236 +
237 +#### G18 — SpecSweep (total 78)
238 +- **N 6** — A port of SpecExec to unified memory + NVMe; the platform mechanics (zero-copy Metal buffers, no PCIe staging, 13 GB/s overflow sweeps) are new, the idea is not.
239 +- **P 8** — The arithmetic uses only measured numbers (expH ceiling × SpecExec acceptance), and speculation already ships on Metal in both llama.cpp and MLX.
240 +- **Mem 7** — Dense *stock* models beyond RAM with zero model transformation: 20 GB of overflow sweeps in ~1.5 s; ~1.4× oversubscription plausibly interactive.
241 +- **BW 8** — Overflow bytes ÷ ~20 accepted tokens/sweep; the amortization factor is measured elsewhere, not hypothesized.
242 +- **Q 10** — Exact rejection sampling; the target distribution is untouched by construction.
243 +- **IC 7** — No new representation and no new kernels: a draft model, tree verification, and an overflow sweep (plus G08's ring for sweep stability).
244 +- **HW 9** — Everything runs on shipped Metal paths today; the one unknown is RAM-vs-RAM contention between batch verification and the concurrent sweep — the contention expH did not measure.
245 +- **PT 10** — Stock checkpoints, stock draft model.
246 +- **Sig 6** — Doubles as the honest baseline every representation-heavy gap must beat at equal oversubscription; even its negative result is load-bearing for the project.
247 +- **Risk 7** — Acceptance shortfall degrades throughput gracefully rather than killing the mechanism; the kill (accepted length <5, or <2 tok/s compound at 1.4×) still leaves a working exact runtime.
248 +
249 +#### G19 — RouterBias (total 66)
250 +- **N 6** — Trained cache-conditional experts exist; the zero-shot measurement is genuinely missing but conceptually tiny.
251 +- **P 6** — Near-tied gates are common under load-balancing pressure, yet the trained systems exist precisely because zero-shot biasing may not preserve quality.
252 +- **Mem 4** — Raises effective cache capacity on over-budget MoE; no direct reduction.
253 +- **BW 6** — A 10–15-point hit-rate lift multiplies directly into tokens/s where misses dominate token time.
254 +- **Q 4** — Systematically starving rarely-resident experts risks distribution drift over long generations — a failure aggregate PPL won't show but flips will.
255 +- **IC 10** — ~5 lines in mlx-lm plus a λ sweep; the cheapest intervention in the entire document.
256 +- **HW 10** — Pure logit arithmetic; no kernels, no I/O, no formats.
257 +- **PT 10** — A runtime knob on a stock router.
258 +- **Sig 4** — A knob, not a regime: highest value-per-effort in the document, lowest ceiling.
259 +- **Risk 6** — Cheap to test, but the kill (every useful λ costs >2% flips or >1 pt GSM8K) is quite plausible given why the trained variants exist.
260 +
261 +#### G20 — AmxSidecar (total 61)
262 +- **N 6** — Kernel-affinity splitting on unified memory is unexplored precisely because capacity splitting is meaningless there — nobody looked; still an infrastructure idea, not a mechanism.
263 +- **P 7** — P-cores idle during decode, AMX sustains ~100+ GB/s, copies are zero by construction; the unmeasured hole is RAM-bandwidth contention (expH tested SSD-vs-GPU, not CPU-vs-GPU).
264 +- **Mem 2** — None.
265 +- **BW 3** — No byte reduction; it reroutes irregular work (gathers, ΔY, probes) off the GPU's weakest path.
266 +- **Q 9** — Numerically identical work on a different engine.
267 +- **IC 6** — Accelerate GEMV + shared-event sync; the microbench is a day, production scheduling more.
268 +- **HW 8** — Exploits Apple-specific silicon (AMX/Accelerate) that every CUDA-derived design structurally ignores.
269 +- **PT 10** — Pure runtime.
270 +- **Sig 4** — A contingency: the escape hatch for G03/G06/G15 gather-kernel failures, not a result in itself.
271 +- **Risk 6** — expM is cheap; failure merely closes the escape hatch (kill: CPU <50 GB/s under GPU saturation, GPU loses >15%, or sync >50 µs/layer).
272 +
273 +### G. Attention-side working set
274 +#### G21 — KVBudget (total 69)
275 +- **N 4** — SparQ, KIVI, KVQuant are all published and KV quantization on Metal is "solved-enough"; the Metal SparQ port and joint weight+KV accounting are engineering with a thin research veneer.
276 +- **P 9** — All ingredient algorithms are training-free and validated; mlx-lm's quantized-KV plumbing exists to extend.
277 +- **Mem 5** — Frees gigabytes of residency at long context — a direct multiplier on every weight-side cache's budget.
278 +- **BW 6** — Up to 8× attention-transfer cut where context is long; possibly <10% of per-token bytes at ≤32k — the self-mooting risk.
279 +- **Q 9** — Published methods are near-lossless at 3–4-bit KV.
280 +- **IC 7** — An extension of existing quantized-KV paths plus one gather kernel.
281 +- **HW 8** — RAM-resident gathers only — the 35:1 SSD ratio never applies; the Metal gather cost question is the sole concern.
282 +- **PT 10** — A runtime cache format.
283 +- **Sig 3** — Explicitly self-demoting: expN_kvshare (one afternoon) may strike it to Phase-10 engineering; its main role is to protect other gaps' accounting.
284 +- **Risk 8** — Nothing fails; the worst case is "not a priority," established almost for free.
285 +
286 +### H. Contrarian gaps (scored as experiments — cheap falsifiers, not architectures)
287 +#### G22 — WarmEnough (total 71)
288 +- **N 6** — Warm steady-state fault accounting for selective workloads on macOS is unpublished; the mechanism under test (UBC + kernel LRU) is ancient.
289 +- **P 8** — llama.cpp + dtrace + wired ballast yields a decisive number in days; every dramatic prior number (0.025 tok/s, 2.23×) is cold-start or dense-cyclic, so the warm question is genuinely open.
290 +- **Mem 1** — Reduces nothing; it can only *delete engineering* (Group C) from the plan.
291 +- **BW 1** — None; it measures whether the kernel already delivers the bandwidth story.
292 +- **Q 10** — A measurement; quality untouched.
293 +- **IC 8** — Scripts around existing tools; no new code paths.
294 +- **HW 10** — Runs on the shipped stack as-is.
295 +- **PT 10** — Nothing touches any model.
296 +- **Sig 8** — A <10% warm fault-time share at 75% residency demotes ~40% of the document's engineering (G08–G11); >25% validates Group C — enormous information either way.
297 +- **Risk 9** — Cannot fail as an experiment; both outcomes are directly actionable.
298 +
299 +#### G23 — ConfidentlyWrong (total 76)
300 +- **N 8** — P(disagree | margin) for quantized bases — the joint distribution, not the marginals — has never been published (SOTA §6.7), despite a crowded escalation literature built on its assumed shape.
301 +- **P 9** — Pure logging and analysis over the expG run matrix; the measurement itself cannot fail.
302 +- **Mem 1** — None; it gates the gaps that claim reductions.
303 +- **BW 1** — None.
304 +- **Q 10** — A measurement.
305 +- **IC 9** — Shares the expG run with G02, G17, and G24 — near-zero marginal cost.
306 +- **HW 10** — MLX quantized casts plus logging; no kernels.
307 +- **PT 10** — Nothing modified.
308 +- **Sig 9** — Single-handedly decides whether the escalation family is the project's center (AUROC >0.85) or dead (<0.65 at a 3-bit base) — the highest information-per-compute-hour in the plan; the gaps doc orders it first for this reason.
309 +- **Risk 9** — Decisive either way; both outcomes are publishable guidance for the whole cheap-first field.
310 +
311 +#### G24 — BasinCollapse (total 74)
312 +- **N 7** — Agreement-with-full-model and recovery-smoothness-per-residual-byte at 2–3-bit PTQ bases are unreported; the ParetoQ transition it probes is published but measured only as standalone quality, never as draft quality.
313 +- **P 9** — MLX quantization at four bitrates plus reference logging; entirely mechanical on the shared ladder.
314 +- **Mem 1** — None directly; it sets the resident-base floor every progressive gap must respect.
315 +- **BW 1** — None.
316 +- **Q 10** — A measurement.
317 +- **IC 8** — The quantize/refine ladder it needs is shared infrastructure for eight other gaps.
318 +- **HW 10** — Stock MLX.
319 +- **PT 10** — Nothing modified.
320 +- **Sig 9** — Fixes the project's oversubscription ceiling: 3-bit agreement <75% caps the whole progressive family at ~1.5–1.6× and reframes the charter §15 stretch target as unreachable by this route — the single number that stretch goal lives or dies on.
321 +- **Risk 9** — Any outcome is a measured ceiling worth recording (charter §17 explicitly values this).
322 +
323 +## 2. Summary table
324 +
325 +| Gap | Name | N | P | Mem | BW | Q | IC | HW | PT | Sig | Risk | Total |
326 +|-----|------------------|---|---|-----|----|----|----|----|----|-----|------|-------|
327 +| G18 | SpecSweep | 6 | 8 | 7 | 8 | 10 | 7 | 9 | 10 | 6 | 7 | **78** |
328 +| G02 | MarginGate | 9 | 7 | 8 | 9 | 7 | 5 | 8 | 9 | 10 | 5 | **77** |
329 +| G17 | BlockVerify | 8 | 7 | 8 | 9 | 9 | 5 | 7 | 9 | 9 | 5 | **76** |
330 +| G23 | ConfidentlyWrong | 8 | 9 | 1 | 1 | 10 | 9 | 10 | 10 | 9 | 9 | **76** |
331 +| G09 | ArcExperts | 5 | 9 | 6 | 7 | 10 | 6 | 9 | 10 | 5 | 8 | **75** |
332 +| G24 | BasinCollapse | 7 | 9 | 1 | 1 | 10 | 8 | 10 | 10 | 9 | 9 | **74** |
333 +| G04 | RDLayout | 6 | 8 | 7 | 6 | 8 | 6 | 8 | 10 | 6 | 7 | **72** |
334 +| G22 | WarmEnough | 6 | 8 | 1 | 1 | 10 | 8 | 10 | 10 | 8 | 9 | **71** |
335 +| G01 | ResidualPager | 8 | 6 | 9 | 7 | 7 | 4 | 6 | 9 | 9 | 4 | **69** |
336 +| G21 | KVBudget | 4 | 9 | 5 | 6 | 9 | 7 | 8 | 10 | 3 | 8 | **69** |
337 +| G07 | WandaTiers | 7 | 7 | 6 | 6 | 6 | 7 | 8 | 10 | 6 | 5 | **68** |
338 +| G03 | DeltaKernel | 8 | 7 | 4 | 7 | 9 | 5 | 5 | 10 | 7 | 5 | **67** |
339 +| G08 | MRURing | 6 | 8 | 4 | 6 | 10 | 5 | 6 | 10 | 5 | 6 | **66** |
340 +| G13 | DomainCache | 7 | 5 | 6 | 8 | 7 | 6 | 8 | 9 | 7 | 3 | **66** |
341 +| G19 | RouterBias | 6 | 6 | 4 | 6 | 4 | 10 | 10 | 10 | 4 | 6 | **66** |
342 +| G05 | MetalTrellis | 7 | 6 | 6 | 7 | 7 | 4 | 5 | 10 | 7 | 4 | **63** |
343 +| G15 | SpectralSplit | 6 | 5 | 6 | 6 | 5 | 7 | 9 | 9 | 6 | 4 | **63** |
344 +| G16 | SeedSynth | 7 | 5 | 7 | 9 | 4 | 5 | 6 | 9 | 7 | 3 | **62** |
345 +| G20 | AmxSidecar | 6 | 7 | 2 | 3 | 9 | 6 | 8 | 10 | 4 | 6 | **61** |
346 +| G10 | PurgeableWarm | 5 | 6 | 2 | 2 | 9 | 6 | 7 | 10 | 4 | 6 | **57** |
347 +| G12 | HiddenPrefetch | 6 | 6 | 2 | 6 | 8 | 4 | 7 | 8 | 6 | 4 | **57** |
348 +| G14 | TrunkDelta | 7 | 4 | 7 | 7 | 4 | 4 | 7 | 6 | 7 | 3 | **56** |
349 +| G11 | NeverBlock | 8 | 5 | 3 | 4 | 5 | 4 | 7 | 9 | 7 | 3 | **55** |
350 +| G06 | TealPager | 6 | 6 | 5 | 5 | 4 | 4 | 4 | 8 | 6 | 3 | **51** |
351 +
352 +Totals span 51–78. Unweighted totals reward safety: the falsifiers and G18 rank high partly
353 +because they cannot fail badly, while G01 (the project's canonical representation) is dragged
354 +down by IC 4 / Risk 4 — selection below therefore weights Sig, cluster synergy, and the
355 +substrate contract, not raw totals. G21 (69) is unselected as self-demoting pending expN.
356 +
357 +## 3. Selected candidates for prototyping
358 +
359 +Four candidates: three architecture compositions (following the natural clusters named in
360 +`research_gaps.md`) plus one falsifier battery that gates them; all share the build-once
361 +infrastructure listed there.
362 +
363 +### Candidate 1 — Progressive-Residual Runtime (G01 + G02 + G04, with G03 kernel; gated by G05)
364 +**Selection rationale.** The composition the entire SOTA analysis points at: a 2–3-bit base
365 +resident in unified memory (17.5–26 GB for a 70B model on the 48 GB target); residual planes
366 +compiled by G04 into KL-optimal, 16 KB-aligned, ≥256 KiB embedded bundles on NVMe; fetched
367 +only when G02's margin/learned-feature gate says the token decision is not yet stable; applied
368 +via G03's ΔY kernel at cost proportional to residual size. It attacks criteria **B**
369 +(bytes/token ≪ compressed checkpoint: base RAM reads + 5–15% escalation × residual fetch),
370 +**C** (token-dependent fraction of model information), **D** (progressive/conditional
371 +execution), and **E** (a post-training representation with a qualitatively better
372 +storage/RAM/quality tradeoff than fixed quantization), and satisfies the substrate contract
373 +by construction: escalation on 5–15% of tokens fits under 650 MB/token, prefetch overlap is
374 +free, G04/G05 keep decode shift/mask-cheap. Highest-Sig cluster (10 + 9 + 6) despite mid-pack
375 +totals — exactly where totals mislead.
376 +
377 +**Promote / kill (expG, expD, expB, expF, expE variants).**
378 +- PROMOTE: expG AUROC ≥0.8 for margin(+learned features) predicting base-vs-reference disagreement, with escalation ≤40% at 99% effective agreement on every major domain; expD reaching KL ≤0.05 to FP16 touching ≤20% of residual bytes/token; expB temporal reuse of the refined set ≥50%; expE_decode progressive variants ≥70% of affine-baseline throughput; expE_delta ΔY ≥2× over full recompute at 10% density.
379 +- KILL: expG AUROC <0.8 or escalation >40% on any major domain (G02 dies; fall back to Candidate 2's ungated verification); expD >20% residual bytes AND expB reuse <50% (steady-state traffic >650 MB/token — G01 dies at 70B scale); expE_decode both variants <70% (G05 dies; degrade to affine-delta residuals and re-score); G24's 3-bit agreement <75% on chat/code or non-smooth recovery (family capped at ~1.5–1.6× oversubscription — record the ceiling, pivot).
380 +
381 +### Candidate 2 — Amortized Verification Sweeps (G17 + G18, on a G08 MRU floor)
382 +**Selection rationale.** The speculation cluster converts the SSD's one great strength —
383 +sequential 13 GB/s sweeps — into exact over-budget decoding by dividing swept bytes across
384 +accepted token blocks. G18 is buildable today with zero model transformation (stock GGUF/MLX
385 +checkpoints, shipped Metal speculation): a 20 GB overflow sweeps in ~1.5 s, so ~20 accepted
386 +tokens/sweep gives an I/O ceiling near 13 tok/s at 1.4× over-RAM — criterion **A** (a model
387 +larger than unified memory at interactive latency) with exact outputs, plus **F** (speculative
388 +sweep scheduling materially changes out-of-core economics). G17 is the research upgrade: the
389 +same amortization applied to residual planes, relaxing Candidate 1's per-token traffic ceiling
390 +by the acceptance length while keeping rejection-sampling exactness. G18 is also the charter
391 +§8.4-honest baseline Candidate 1 must beat at equal oversubscription — funding it is
392 +epistemically mandatory, and its worst case is still a working exact runtime.
393 +
394 +**Promote / kill (expL_specsweep, expG acceptance variant, expI_cachesim).**
395 +- PROMOTE: expL accepted length ≥5 with a 1B draft against 70B-Q4 at tree width 16–64, and Metal batch verification sustaining ≥2 tok/s compound at 1.4× over-RAM under a concurrent 13 GB/s stream; expG-variant mean accepted block length ≥4 at a 3-bit base on chat/code (G17 economics hold).
396 +- KILL: expL accepted length <5 or verification-under-streaming drops the compound ceiling below 2 tok/s (G18 demoted to baseline-only); expG-variant accepted length <4 at 3-bit base (G17 dies — sweep bytes ÷ 4 exceeds the per-token budget for 1.5× over-RAM models); expI/expO showing the swept overflow set unstable under MRU (<1.5× over kernel-LRU) removes the G08 floor.
397 +
398 +### Candidate 3 — MoE Residency Engine (G09 + G12 + G19, with G10 warm tier and G04-style bundle layout)
399 +**Selection rationale.** The cache/prefetch cluster is the nearest-term deployable win and the
400 +only candidate whose per-token selectivity is *free* (the router computes it): compiled
401 +16 KB-aligned expert bundles (natively meeting the ≥256 KiB QD≥8 contract), an ARC ghost-list
402 +cache, hidden-state prefetch N layers ahead (lead time ≥ fetch latency, which expH makes
403 +sufficient), zero-shot router bias as a cheap multiplier, purgeable-warm demotion for graceful
404 +pressure behavior. Quality is exact, every ingredient has isolated measured evidence (2.23×
405 +from layout alone on a Mac), and it serves criteria **A** (Qwen3-30B/80B-class MoE on 16–48 GB
406 +Macs, interactive) and **F** (cache/prediction/scheduling changing out-of-core economics). Its
407 +known weakness is the "engineering, not research" risk — precisely what its kill thresholds
408 +test, and why it runs on shared infrastructure rather than monopolizing effort.
409 +
410 +**Promote / kill (expB MoE-trace variant, expO_warmsteady, expJ_purgeable, G19 sweep).**
411 +- PROMOTE: ARC beats LRU by ≥5 hit-rate points at some cache size in trace replay (25/50/75% of expert pool); prefetch probe recall ≥70% at ≤2× overfetch for non-resident experts; router bias lifting hit rate ≥10 points at ≤2% flips and ≤1 pt GSM8K.
412 +- KILL: ARC <5 points over LRU at every cache size AND expO warm fault-time share <10% at 75% residency (policy headroom ≈ 0 and the kernel already suffices — demote to llama.cpp-upstream engineering, keep only the layout compiler); expJ transition overhead >10 µs/block/token kills the purgeable tier only; G19's kill (every useful λ costs >2% flips or >1 pt GSM8K) strikes the bias knob only.
413 +
414 +### Candidate 4 — Gate Zero: the contrarian falsifier battery (G23 + G24 + G22, plus the G21 expN probe)
415 +**Selection rationale.** Selected *as experiments*, per the gaps doc's own framing: the
416 +cheapest instruments in the project, and they arbitrate everything above before engineering
417 +begins. G23/G24 ride the single expG run matrix (which also serves G02/G17) and decide whether
418 +the escalation/progressive families are the project's center or dead; G22 (llama.cpp + dtrace,
419 +days) decides whether Group C custom residency is research or redundant; expN (an afternoon)
420 +decides whether KV growth silently invalidates every weight-side budget. Charter §17 makes
421 +concluding "it doesn't work" a first-class deliverable, and the fan-in analysis shows these
422 +four measurements gate the priority or design of seventeen of the twenty-four gaps. If all
423 +contrarians "win," the documented salvage architecture (G18 + G09 + G04: speculative sweeps
424 +over a well-laid-out, kernel-cached, R-D-compiled model) is itself unoccupied territory — even
425 +total premise failure yields Candidates 2 and 3 plus a publishable negative map.
426 +
427 +**Decision thresholds (these ARE the promote/kill evidence for the whole portfolio).**
428 +- G23: AUROC <0.65 (margin+entropy+learned features, 3-bit base) ⇒ G02 and G17's gated forms die; block-level rejection sampling (G17/G18 ungated) survives. AUROC >0.85 ⇒ the escalation family becomes the project's center.
429 +- G24: 3-bit same-top-token <75% on chat/code, or KL plateaus until >50% of residual bytes ⇒ the progressive family is capped at ~1.5–1.6× oversubscription; record the measured ceiling; Candidate 1 is demoted and Candidate 2 promoted to primary.
430 +- G22: warm fault-time share <10% of token time at 75% residency ⇒ G08–G11 demoted to engineering polish (Candidate 3 loses its cache-policy arm, keeps layout+prefetch); >25% ⇒ Group C stands as designed.
431 +- expN: KV share <10% of per-token bytes at ≤32k for 8B–30B models ⇒ strike G21 from the candidate list to Phase-10 engineering; otherwise every candidate's RAM budget is re-derived with KV co-modeled.
432 +
433 +## 4. Preferred candidate
434 +
435 +If one candidate defines the project, it is **Candidate 1, the Progressive-Residual Runtime
436 +(G01+G02+G04+G03)** — the only one that rewrites the relationship charter §20 names as the
437 +ideal outcome. Candidate 2 amortizes reads and Candidate 3 exploits existing routing, but both
438 +leave hardware coupled to model size (the sweep scales with overflow bytes; the expert pool
439 +sets tier sizes). Candidate 1 alone makes required RAM a function of the *base representation*
440 +(a compile-time constant chosen to fit the machine) and per-token cost a function of *decision
441 +uncertainty* (escalation rate × residual fetch) — additional model size lands on NVMe as
442 +embedded residual planes that cost latency only when, and in proportion to how often, the
443 +token decision actually needs them. That is `instantaneous useful working set → hardware
444 +required` in mechanism form, on the one platform (unified memory, zero-copy Metal buffers,
445 +13 GB/s NVMe with free overlap) where its mechanics are measured viable, while keeping the
446 +original pretrained weights as the sole source of intelligence. It is also the portfolio's
447 +riskiest member — hence Gate Zero runs first, Candidate 2 stands as its honest baseline and
448 +fallback, and its kill thresholds are numeric and scheduled before any runtime engineering.
449 +Next (Phase 5/6): build shared infrastructure in the gaps doc's order — expG run matrix first,
450 +then the expA/B/C trace campaign, then the expD/expE ladder — before candidate-specific code.
modified web/sync-content.sh +5 −2
@@ -16,13 +16,16 @@ REPO_ROOT="$(cd "$(dirname "$0")/.." && pwd)"
16 16 CONTENT="$REPO_ROOT/web/content"
17 17
18 18 mkdir -p "$CONTENT"
19 +# NOTE: rsync applies filters in order — excludes MUST precede the broad
20 +# includes, or e.g. experiments/**/testfile.bin (8 GiB) slips through.
19 21 rsync -a --delete \
22 + --exclude='*.bin' --exclude='*.log' --exclude='__pycache__' --exclude='.git' \
23 + --exclude='node_modules' --exclude='web' \
20 24 --include='CLAUDE.md' --include='README.md' --include='CITATION.cff' \
21 25 --include='Makefile' --include='pyproject.toml' \
22 26 --include='research/***' --include='experiments/***' --include='benchmarks/***' \
23 27 --include='src/***' --include='tools/***' --include='results/***' --include='docs/***' \
24 --exclude='*.bin' --exclude='*.log' --exclude='__pycache__' --exclude='.git' \
25 --exclude='node_modules' --exclude='web' --exclude='*' \
28 + --exclude='*' \
26 29 "$REPO_ROOT/" "$CONTENT/"
27 30
28 31 COMMIT=$(git -C "$REPO_ROOT" rev-parse HEAD 2>/dev/null || echo "unknown")
29 32