project: localvm-research document: expB_token_stability/analysis author: Simon-Pierre Boucher contact: contact@spboucher.ai created: 2026-08-12 status: reviewed
Analysis — expB_token_stability
Run: results/expB_token_stability/20260812T045224Z/ · consumes expA's trace
(Qwen3-1.7B, 6 144 positions × 28 layers), 64-neuron sets covering 95% of
per-token FFN energy.
text
Hypothesis / Falsification
Hoped: Jaccard(t,t+1) ≥ 0.5 with slow decay; 128-token union ≤80%.
Kill: Jaccard(t,t+1) < 0.3 OR 128-token union ≥95%.
Result — KILL CRITERION TRIGGERED (union), stickiness ILLUSORY
Jaccard(t,t+δ): 0.690 (δ=1) → 0.688 (δ=16) → 0.686 (δ=32) — flat.
Random-set null at matched sizes: 0.640. Real stickiness beyond
set-size artifact: +0.05 — negligible. (The registered ≥0.5 threshold
is technically met at 0.69, but the null exposes it as a size effect:
sets covering ~77% of blocks overlap ~64% by chance alone.)
Union working set: 98.0% of blocks after 8 tokens, 99.3% after 32,
99.9% after 128 — the entire FFN is touched within a few tokens.
Domain locality (expC-lite): within-domain = across-domain = 1.0 —
every trajectory's union covers everything; no domain-specific caches.
Interpretation
1. NEGATIVE, decisively: for dense SwiGLU models, per-token FFN energy
working sets are neither sparse at pageable granularity (expA) nor
temporally reusable (here) nor domain-clustered (here). The
DejaVu / LLM-in-a-flash paradigm does NOT transfer to modern
non-ReLU architectures without ReLUfication retraining — now
measured, not just suspected from the literature. G06 (TealPager),
G12 (HiddenPrefetch for dense FFNs) and G13 (DomainCache) are dead
on this model family; G07 (WandaTiers) survives only for STATIC
(input-independent) importance tiering.
2. Scope note (honesty): we measured activation-ENERGY sets on a 1.7B
dense model. Residual-relevance sets (which bytes fix a decision)
could differ, and MoE models have architectural, not emergent,
sparsity — C3 is untouched by this result. But the prior for any
per-token dynamic weight selection on dense SwiGLU is now strongly
unfavorable.
3. PIVOT (with expF, expA): all three fine-grained routes for cutting
C1's per-token escalation bytes are closed — layers (expF), blocks
(expA), temporal caching (expB). What remains is BATCH AMORTIZATION,
which expH makes attractive: the full 1 GB residual streams
sequentially in ~80 ms at 13 GB/s; a margin-gated queue that defers
low-margin tokens and refines them in periodic sweeps shares one
sequential residual pass across many tokens. C1 thereby converges
with C2 (Amortized Verification Sweeps) — the evidence has selected
the candidate.
Next experiment
Prototype candidate: margin-gated deferred-refinement runtime
(C1→C2 merge): 4-bit resident base generates optimistically; low-margin
tokens queue; a periodic sequential residual sweep verifies/corrects
(speculative-decoding-style rollback on flips). Measure end-to-end
bytes/token, tok/s, and agreement on this Mac. Also rerun expG at 8B to
check margin-signal scaling before committing the prototype design.