--- project: localvm-research document: research/notes/sparsity_pruning author: Simon-Pierre Boucher contact: contact@spboucher.ai created: 2026-08-11 status: draft --- # Activation Sparsity & Weight Pruning — Research Notes (charter §4.2, §4.3) Deep literature scan, 2022–2026. Focus: which sparsity/pruning mechanisms decouple `total model size` from `resident model size` and `bytes read per token`, and which of them actually map to Apple Silicon unified memory + Metal + internal NVMe. --- ## 1. Landscape overview Two families, one shared premise: - **Activation sparsity (§4.2):** for a given input, only a small input-dependent subset of FFN neurons / attention heads meaningfully contributes to the output. If that subset can be *predicted cheaply before the matmul*, the untouched weights never need to be read — a **bandwidth** and (with an offload tier) a **capacity** win, with no permanent model change. - **Weight pruning (§4.3):** some weights contribute little for *all* inputs and can be removed permanently (one-shot unstructured: SparseGPT, Wanda; structured: SliceGPT, LLM-Pruner, Sheared-LLaMA, Minitron; depth: ShortGPT). A **capacity** win, but paid for with permanent quality loss and (for unstructured) very poor hardware realizability. Key historical arc: 1. **2022–2023 (ReLU era):** the "lazy neuron" observation — in ReLU transformers <10% of FFN neurons fire per token (3.0% for T5-Base). DejaVu turns this into *contextual sparsity* with per-layer predictors (up to 85% sparsity, 2× over FasterTransformer on OPT-175B). 2. **2023–2024 (offload era):** PowerInfer (hot/cold neurons split across GPU/CPU), **LLM in a flash** (Apple: neurons paged from flash on demand), PowerInfer-2 / Ripple (smartphones, UFS storage). Sparsity becomes a **cache/paging policy**, not just a FLOP saver. This is the closest prior art to the localvm-research goal. 3. **2024–2026 (SwiGLU problem & training-free era):** modern models (Llama-2/3, Mistral, Qwen) use SwiGLU — activations are no longer exactly zero. Responses: *ReLUfication* retraining (ReLU Strikes Back, ProSparse, TurboSparse dReLU, Q-Sparse), or *training-free thresholding* of near-zero activations (CATS, TEAL, GRIFFIN, DIP, R-Sparse, SparseInfer, statistical calibration). Sirius (NeurIPS 2024) shows the quality cost is concentrated in reasoning tasks and is *recoverable by correction*. 4. **Unresolved:** essentially all fast implementations are CUDA (Triton kernels, GPU/CPU split) or Android/UFS. **No system implements predictor-driven sparse weight paging on macOS/Metal/Apple NVMe.** PowerInfer's own README lists "Metal backend for sparse inference on macOS" as *planned, not implemented* (macOS today = CPU only, "limited" gains). The most important conceptual distinction for us: pruning research asks *"which weights can we delete?"*; offload-sparsity research asks *"which weights must be resident right now?"* The second question converts pruning from a destructive transform into a **cache-residency policy** — weights are demoted to NVMe, not destroyed, and recovered when the input distribution needs them. Only a handful of systems (LLM in a flash, PowerInfer-2, M2Cache, DIP, Ripple) partially do this, none on a Mac. --- ## 2. Techniques and systems ### 2.1 DejaVu — contextual sparsity (ICML 2023 oral) - **Mechanism.** Hypothesis: small, input-dependent sets of attention heads and FFN neurons reproduce the dense output for each input. Trains small MLP predictors per layer; exploits slowly-changing hidden states across layers to predict layer ℓ's sparsity from layer ℓ−1's input (asynchronous "lookahead" hides predictor latency). - **Memory / bandwidth.** No resident-memory reduction (full model stays in GPU HBM). Bandwidth: up to 85% contextual sparsity ⇒ proportionally fewer weight bytes streamed from HBM per token. Contextual sparsity has up to 7× better efficiency–accuracy trade-off than static sparsity. - **Quality.** "Without compromising model quality" on OPT-175B; >2× latency reduction vs FasterTransformer, >6× vs HuggingFace. - **Predictor overhead.** Small per-layer MLPs; hidden by async lookahead on A100s. - **Hardware assumptions.** OPT (ReLU FFN), multi-A100 discrete-GPU serving. The GPU-resident-model assumption gives it no capacity benefit; irrelevant as-is for a memory-constrained Mac, but the *predictability result* (85%) is foundational. - **Limitation.** ReLU-dependent; whole model must fit in accelerator memory; predictors trained per model. - **Extension for us.** Reuse the "slowly changing hidden states" property as a *prefetch signal* for NVMe reads rather than a FLOP-skip signal. ### 2.2 PowerInfer — hot/cold neuron split on one consumer GPU (SOSP 2024) - **Mechanism.** Neuron activations follow a **power-law**: a small "hot" set is activated across almost all inputs, the "cold" majority is input-specific. Hot neurons preloaded to GPU VRAM; cold neurons computed on CPU; adaptive per-layer predictors + neuron-aware sparse operators. - **Numbers.** OPT-175B-class models on one RTX 4090: 13.2 tok/s average, 29.08 peak — only 18% below an A100. Requires ReLU-family models (ReluLLaMA, ProSparse-LLaMA, Bamboo, TurboSparse); explicitly does **not** support vanilla Llama/Mistral/Qwen. - **Memory / bandwidth.** GPU-resident set ≪ model size (capacity win via CPU DRAM as second tier); bandwidth win from computing only predicted-active neurons. - **Hardware assumptions.** **Discrete GPU + PCIe + separate CPU DRAM.** This split is *meaningless on Apple unified memory* — there is no "GPU VRAM vs CPU RAM" distinction; the Mac analogue is RAM (hot) vs NVMe (cold), i.e., exactly the LLM-in-a-flash setting. - **macOS status (verified on repo).** Runs on Apple M chips CPU-only with "limited" improvement; Metal sparse backend is a *planned feature that never shipped*. Project still active (SmallThinker 2025, Tiiny AI Pocket Lab Jan 2026) but the Mac gap remains. - **Extension.** Port the hot/cold *statistical* insight to a RAM/NVMe hierarchy with Metal gather-matvec kernels; hot set pinned in wired memory, cold set demand-paged. ### 2.3 LLM in a flash (Apple, ACL 2024) — deepest dive, closest prior art #1 *Alizadeh et al., arXiv 2312.11514. The only major paper that ran sparse weight paging on actual Apple hardware.* **What exactly they did:** - **Hardware:** Apple **M1 Max** (1TB SSD) and **M2 Ultra** (2TB SSD), CPU (fp32) and Metal GPU (fp16) paths; plus Linux RTX 4090 (bf16). Memory budget: ~**half the model size** in DRAM. - **Models:** OPT-6.7B, sparsified/ReLUfied Falcon-7B, Persimmon-8B, Phi-2, FATReLU Llama-2-7B. All ReLU-family FFNs. - **Predictor:** low-rank predictor per FFN layer (OPT-6.7B: rank 128 for layers 1–28, rank 1024 for last 4). Trained on 10k C4 samples, 2 epochs, ~4 h/predictor on A100. Cost: <2.4% of non-embedding weights/FLOPs; ~5% false negatives, 7% false positives; 2.75% (CPU) / 4.8% (GPU) of compute time. - **Windowing:** keep the union of neurons activated in the last k=4–5 tokens in DRAM; per new token only load the *delta*. OPT-6.7B at k=4: each token touches **2.4% of FFN neurons**; FFN occupies only 15.5% of DRAM; total DRAM ≈ **52.1% of model size**. Falcon-7B: 3.1% of FFN neurons/token, DRAM ≈ 52.9%. - **Row–column bundling:** store up-projection column i together with down-projection row i so one neuron = one contiguous 2·d_model read → doubles chunk size. Measured effective flash throughput: ~1.25 GiB/s (predictor+windowing) → **~2.25 GiB/s with bundling**, vs >6 GiB/s sequential on M1 Max. Sweet spot: ≥32 KiB random reads across **32 threads**. - **Bytes/token (OPT-6.7B):** naive 13.4 GB → predictor only 6.7 GB → **predictor+windowing 0.2 GB per token**. (This is the single most important measured number for our charter §8.3.) - **Speedups:** OPT-6.7B total per-token latency 669 ms CPU (4.75× vs naive), 565 ms Metal/M1 (4.23×), 305 ms Metal/M2 Ultra (7.44×), 84 ms CUDA (26.4×); +speculative → 37×. Runs models up to **2× DRAM size**. **What they did NOT do (our opening):** - Attention weights + embeddings kept **permanently in DRAM** (~19–32% of model) — only FFN weights are paged. No paging of attention, no KV-cache tiering. - Predicated on **ReLU sparsity**; dense SwiGLU models "not addressed" (their words: the method "is constructed on the foundation of sparsified networks"). - Only ~2× DRAM oversubscription demonstrated on ≤8B models — not 4–10×, not 30–70B on a Mac. - Single-batch, greedy decoding only; no power measurement; no quantization co-design (fp16 neurons on flash — 4-bit bundles would quadruple effective neuron throughput). - **No code release.** Nothing in MLX/llama.cpp implements it today. - Negative result they report: bundling by *co-activation* (closest coactivated neighbor) failed — hot neurons got loaded repeatedly. (Ripple later solved this with global placement.) ### 2.4 PowerInfer-2 — deepest dive, closest prior art #2 *Xue et al., arXiv 2406.06282. First 47B model on a smartphone.* **What exactly they did:** - **Hardware:** OnePlus 12 (24 GB DRAM, 19 GB usable, UFS 4.0, Snapdragon 8 Gen 3) and OnePlus Ace 2 (16 GB, UFS 3.1). UFS 4.0: ~4 GB/s sequential (512 KB), **~1 GB/s at 4 KB random**, 850 MB/s over larger ranges. - **Models:** TurboSparse-Mixtral-47B (dReLU MoE, ~3B activated params/token), TurboSparse-Mistral-7B, Bamboo-7B, sparse Llama-13B, Qwen2-7B; SiLU Mistral-7B as a stress case. - **Neuron cluster abstraction:** groups of same-layer FFN neurons with similar activation statistics. Hot (frequently active) clusters → large, dense, NPU-computed; cold clusters → small, sparse, CPU-computed. Prefill: NPU dense matmuls while a CPU core streams weights. Decode: NPU handles ~70% (hot dense part), CPU the sparse remainder. - **Segmented neuron cache:** three DRAM regions — (1) pinned attention + KV, (2) hot region (cluster-granularity LRU), (3) cold region (neuron-granularity LRU); ratios adapt to batch size. - **I/O pipeline:** neuron-cluster-level pipelining overlaps compute on cached clusters with UFS reads of missing ones. Cold neurons stored as **Gate-Up-Down bundles** (~80% co-activation across the three matrices). For 4-bit models: two-phase read — 4 KB Gate first, Up/Down 4 KB fetched *only if* Gate output ≠ 0. - **Results:** 11.68 tok/s decoding for TurboSparse-Mixtral-47B on 24 GB phone (up to 27.8–29× vs llama.cpp; 3.84× vs "LLMFlash" i.e. LLM-in-a-flash-style baseline). With ~50% FFN offload on a 7B: ~11.1 tok/s vs ~14.5 in-memory — offload costs only ~23%. - **Overhead accounting (7 GB budget, 47B model):** 1 GB non-FFN weights + **2.6 GB predictors** + 2.7 GB quantization scales + 0.3 GB runtime = 6.6 GB, leaving only 400 MB of neuron cache (1.8% of FFN weights) — still runs. Predictor DRAM cost is *large*, a real design lesson. **What they did NOT do:** - Rooted **Android only**; "iOS portability" claimed but never validated. Nothing on macOS/Metal/ANE. - Depends on **dReLU ReLUfied models** (TurboSparse = SFT-retrained Mistral/Mixtral, ~150B tokens); on stock SiLU models speedup drops to 2.4× (vs 4.6× ReLU). Not post-training-only. - Offline per-device planner required; no cross-device generalization; high tail latency (P99 +40.9% vs mean). - Never tested Apple-class NVMe (6+ GB/s vs their 1 GB/s random UFS): **the Mac's storage is ~4–6× faster than the storage this system was designed around** — its economics should transfer favorably, but nobody has built it. ### 2.5 Ripple / Neuralink — co-activation-aware flash layout (2024) - **Mechanism.** Neurons that fire together are placed together in flash ("neuron co-activation linking"), converting many small random reads into fewer large sequential reads — directly attacking the IOPS bound that limits smartphone (and, less severely, Mac) storage. - **Relevance.** Solves exactly the negative result LLM in a flash reported for co-activation bundling, via offline global placement optimization. An offline "layout compiler" of this kind belongs in our compile stage (§14 of charter). Hardware: Android/UFS; no Mac port. ### 2.6 ShadowLLM — better importance predictors (EMNLP 2024) - **Mechanism.** Single early-layer predictor "shadows" the whole model instead of per-layer predictors; goes beyond magnitude-based criteria for head/neuron importance. - **Numbers.** >15% end-to-end accuracy improvement over DejaVu-style criteria at equal sparsity; up to 20% speedup over DejaVu; validated on OPT/Llama-2 up to 30B. Code: abdelfattah-lab/shadow_llm. - **Hardware.** CUDA. Predictor design is transferable; a single-point predictor is attractive on Mac because it gives *maximum I/O prefetch lead time* (predict at layer 0 → prefetch layer 30). ### 2.7 ReLUfication line: ReLU Strikes Back → ProSparse → TurboSparse → Q-Sparse - **ReLU Strikes Back (Apple, ICLR 2024):** swapping SiLU/GELU→ReLU and fine-tuning has negligible quality impact while enabling up to ~3× less computation/weight transfer at the memory-bound decode step. Establishes that sparsity is *recoverable post-training* with modest fine-tuning. - **ProSparse (2402.13516):** activation substitution + progressive L1 regularization + threshold shifting → **89.3% / 88.8% / 87.9%** activation sparsity on LLaMA2-7B/13B/MiniCPM-1B at parity with the Swish originals. - **TurboSparse (2406.05955):** **dReLU** + data-mix continued training (~150B tokens) → Mistral-7B activates only **2.5B** params/token; Mixtral-47B activates **4.3B**; 2–5× decode speedup; 11 tok/s on phones (feeds PowerInfer-2). - **Q-Sparse (2407.10969, NeurIPS 2024):** top-K activation sparsification with STE during training; full activation sparsity, inference-optimal scaling law for sparse LLMs; works with BitNet 1.58-bit. Training-time method — violates our post-training-only constraint but maps the ceiling. - **Caveat for us:** all of these need GPU-scale fine-tuning (out of scope to *produce*, but the checkpoints — ProSparse-LLaMA, Bamboo-7B, TurboSparse-Mistral/Mixtral — are downloadable and are ideal *test vehicles* on the Mac). ### 2.8 Training-free SwiGLU sparsity: CATS, TEAL, GRIFFIN, DIP, and successors - **CATS (COLM 2024):** thresholds the **gate output** of SwiGLU blocks (per-layer thresholds from calibration distributions). ~**99% of base task performance at 50% FFN activation sparsity** on Mistral-7B/Llama2-7B without fine-tuning; but only sparsifies Wup/Wdown ⇒ ~25% model-wide sparsity. Custom GPU kernel gives wall-clock gains. - **TEAL (ICLR 2025):** magnitude-thresholds **hidden states model-wide** (every matrix incl. attention), exploiting zero-mean unimodal activation distributions. **40–50% model-wide sparsity with minimal degradation** on Llama-2/-3/Mistral 7B–70B; 1.53×/1.8× decode speedup at 40/50% via Triton gather kernels; composes with weight quantization. **CUDA-only kernels** — the natural first Metal port target. - **GRIFFIN (ICML 2024):** "flocking" — within a sequence, tokens activate largely the *same* FF neurons. Selects experts once per sequence from the prompt, no training/calibration, works on many non-ReLU activations. **50% of FF parameters with little-to-no degradation** (generation + classification), lower latency. Sequence-level selection = coarse, *prefetch-friendly* granularity (one NVMe read burst per sequence, not per token). - **DIP — Dynamic Input Pruning with Cache-Aware Masking (Qualcomm, 2412.01380):** predictor-free magnitude sparsification of SwiGLU + optional LoRA recovery + **cache-aware masking**: the sparsity mask is chosen considering *what is already in the DRAM cache*, trading tiny accuracy for large cache-hit-rate gains. Phi-3-Medium under mobile DRAM limits: **46% less memory, 40% more throughput, <0.1 ppl loss** vs streaming the dense model. This is the first explicit "sparsity-as-cache-policy" formulation — conceptually the closest paper to our §4.3 key question, but on simulated mobile constraints, CUDA/simulation, no Mac. - **Others (2024–2026):** SparseInfer (training-free sign-bit activation prediction); Post-Training Statistical Calibration (2412.07174); R-Sparse (rank-aware, training-free, attention+FFN); La RoSA (layerwise rotation before sparsification); Spark Transformer (Google, NeurIPS 2025: restores FFN+attention sparsity with low-cost top-k predictor); CETT/“Universal Properties” (2509.00454: sparsity potential *grows with model size*, first diffusion-LLM study); Fast Forward (2602.00397: predictive FFN sparsity for *prefill*); tree-structured FFN dynamic sparsity at scale (2604.08565); flexible N:M *activation* sparsity benchmarking for next-gen accelerators (2509.22166). ### 2.9 SparQ Attention — the attention-side analogue (ICML 2024) - Fetches only the KV-cache rows whose keys matter for the current query (rank the query's large components, approximate scores, fetch top keys). **Up to 8× reduction in attention data transfer** with negligible loss on Llama-2/3, Mistral, Gemma, Pythia; no fine-tuning. - Relevance: our working-set question applies to KV as well as weights; SparQ shows demand-driven fetching works for attention state. Complements FFN-side sparsity (weights dominate at short context, KV at long context). ### 2.10 One-shot weight pruning: SparseGPT & Wanda (verified numbers) - **SparseGPT (ICML 2023):** layer-wise sparse regression with approximate Hessian-based weight updates; prunes OPT-175B/BLOOM-176B in <4.5 h on one GPU; 50–60% unstructured with small ppl increase at 100B+ scale. - **Wanda (ICLR 2024):** score = |weight| × ‖input activation‖, per-output-row, no weight update; ~300× faster to compute than SparseGPT. - **WikiText-2 perplexity (dense → magnitude / SparseGPT / Wanda at 50% unstructured):** - LLaMA-7B: 5.68 → 17.29 / 7.22 / 7.26 - LLaMA-13B: 5.09 → 20.21 / 6.21 / 6.15 - LLaMA-65B: 3.56 → 5.90 / 4.57 / 4.57 - LLaMA-2-70B: 3.12 → 4.98 / 3.98 / 3.98 - 2:4 structured is much worse at small scale (LLaMA-2-7B: 5.12 → 11.02 Wanda). - **Interpretation:** (i) one-shot 50% is nearly free only at ≥65B scale; at 7B it costs ~25–28% ppl. (ii) Hardware-friendly 2:4 patterns are the most damaging. (iii) End-to-end gain is meager even on NVIDIA (Wanda reports 1.24× e2e with 2:4 on LLaMA-7B) — and **Apple GPUs have no sparse tensor cores at all**, so unstructured weight sparsity yields *zero* bandwidth savings on a Mac unless the representation itself skips loads (which CSR-style indexing overhead largely cancels — see Endor). - **Reframing for us:** Wanda's per-input-statistics score is cheap enough to compute *online*; it can rank weights for **residency** (RAM vs NVMe) instead of deletion. ### 2.11 Structured / depth / width pruning: SliceGPT, LLM-Pruner, ShortGPT, Sheared-LLaMA, Minitron - **SliceGPT (ICLR 2024):** orthogonal-rotation + slicing; removes up to 25% of parameters with 99% (Llama-2-70B, OPT-66B) / 90% (Phi-2) zero-shot retention; produces smaller *dense* matrices — the only pruning family that trivially runs fast on Metal (it's just smaller GEMMs). - **LLM-Pruner (NeurIPS 2023):** gradient-based structural pruning + LoRA recovery; ~20% compression at moderate loss. - **ShortGPT (2403.03853):** Block Influence = 1 − cos-sim(layer input, output); middle-to-late layers (e.g. 21–29 of LLaMA-2-7B) are highly redundant; deleting whole layers beats fancier methods on many benchmarks, though generation/reasoning suffer more than classification (follow-ups: Prune&Comp, E³-Pruner, layer-pruning-limits studies 2025–26). - **Sheared-LLaMA (ICLR 2024):** Lagrangian-learned masks over layers/heads/dims + dynamic batch loading; LLaMA2-7B → 1.3B/2.7B using ~50B tokens (3% of from-scratch compute); beats equal-size models trained from scratch. - **Minitron (NVIDIA 2024):** iterative width(+depth) pruning + KD; Nemotron-4 15B → 8B/4B at up to 40× fewer tokens than from-scratch; Llama-3.1-Minitron-4B-Depth distilled on ~380B tokens. - **Assessment:** all of these *permanently* shrink the model — excellent baselines, but they answer the wrong question for us (they reduce total size, not the size↔residency coupling), and the strong ones require retraining budgets we don't have. Their *diagnostics* (Block Influence, activation-norm channel ranking) are directly reusable as residency scores. ### 2.12 Sparsity + offload hybrids: M2Cache, Endor, RAP - **M2Cache (2410.14740):** neuron-level mixed precision + three-tier cache **HBM → DRAM → SSD**; important neurons fp16, colder ones more aggressively quantized, coldest on SSD; LRU at neuron granularity. The only published system with an explicit *SSD tier in a neuron cache hierarchy*. Server GPUs, not Mac. - **Endor (2406.11674):** key observation for §4.3 — CSR-style formats for unstructured-pruned LLMs spend the saved bytes on indices, so offloaded pruned models are usually stored *dense*; proposes a hardware-friendly bitmap format so pruned weights actually transfer fewer bytes from SSD/flash. Directly applicable to any Mac NVMe streaming design. - **RAP (2505.17138):** RL-guided *runtime* elastic pruning of weights + KV under a live memory budget — "pruning as a scheduling decision," another step toward pruning-as-policy. ### 2.13 Quality reality check: Sirius - **Sirius (NeurIPS 2024):** systematic evaluation shows contextual-sparsity models hold up on prompt-understanding tasks but **significantly degrade on reasoning/deduction/knowledge (GSM8K, coding)**; yet sparse and dense models share problem-solving structure, and correcting only ~**11% of tokens** (dense-model verification, KV/hardware-efficient) restores full accuracy at ~78% of the theoretical efficiency gain. - **Implication:** any localvm design that uses aggressive sparsity should budget a dense-verification / correction path (cheap on unified memory since CPU+GPU share the cache); perplexity alone will overstate quality (CATS/TEAL-style "99% retention" claims are mostly perplexity + short benchmarks, not multi-step reasoning). --- ## 3. Relevance to localvm-research ### 3.1 How predictable are activation patterns, really? (measured numbers) - **Per-token sparsity (ReLU models):** 85% contextual sparsity (DejaVu, OPT); 2.4–3.1% of FFN neurons touched per token (LLM in a flash, OPT-6.7B/Falcon-7B); ~90% (ProSparse); <10% FFN activation in ReLU transformers generally (Lazy Neuron). - **Temporal stability:** windowing over k=4–5 tokens works — after caching the last-4-token union, the per-token *delta* of new neurons is small enough to cut bytes/token from 6.7 GB (predictor alone) to **0.2 GB** (OPT-6.7B). This is a direct, measured confirmation of our Experiment B hypothesis on real hardware. - **Spatial/structural stability:** power-law hot/cold split (PowerInfer); ~80% co-activation of Gate-Up-Down bundles per neuron (PowerInfer-2); strong cross-neuron co-activation exploitable by layout (Ripple); **flocking** — sequence-level shared expert sets (GRIFFIN, 50% FF params chosen once per prompt). - **Predictor accuracy/cost:** low-rank per-layer predictors reach ~5% FN / 7% FP at <2.4% weight/FLOP overhead (LLM in a flash); but at scale predictors get heavy — 2.6 GB DRAM for a 47B model (PowerInfer-2). Predictor-free thresholding (CATS/TEAL/DIP) eliminates this cost at some accuracy expense; ShadowLLM shows one early predictor can serve all layers (better accuracy and more prefetch lead time). - **Open question for our Experiments A–C:** all the stability numbers above are from ReLU or ReLUfied models; nobody has published Jaccard/transition statistics for *thresholded SwiGLU* working sets (TEAL/CATS-style masks) — we should measure this ourselves on Llama-3/Qwen-class models before building anything. ### 3.2 Does sparsity survive in non-ReLU models? Partially, and this is the field's central tension: - Exact zeros: essentially none in SwiGLU (DIP: "little inherent sparsity"). - *Approximate* sparsity: 40–50% model-wide prunable per token with minimal ppl loss (TEAL); 50% FFN with ~99% task retention (CATS); 50% FF via flocking (GRIFFIN); and the recoverable ceiling with fine-tuning is ~85–90% (ProSparse/TurboSparse). Universal-properties (2509.00454) finds effective sparsity *increases with model size* — good news since our targets are big. - But 40–50% ≠ 97%. For SSD paging, a 2× reduction in bytes/token is helpful yet far from the 50× that ReLU windowing achieved. Bridging options: combine threshold sparsity with quantization (TEAL composes), low-rank hot path + sparse cold residual (R-Sparse suggests the decomposition), Sirius-style correction, or accept ReLUfied checkpoints where they exist. - Quality caveat: Sirius's reasoning-degradation result means our quality metrics (§8.2) must include GSM8K/coding-style tasks, not just perplexity. ### 3.3 What maps to Apple Silicon, and what doesn't | Assumption in prior work | Apple Silicon reality | |---|---| | GPU VRAM vs CPU DRAM split (DejaVu, PowerInfer) | No split — unified memory. The only meaningful hierarchy is **RAM vs NVMe** | | PCIe transfer cost motivates hot/cold placement | Zero-copy CPU↔GPU; placement = *residency*, not device | | 2:4 sparse tensor cores (Wanda/SparseGPT speedups) | **Absent on Apple GPUs** — N:M weight sparsity gives no free lunch; savings must come from avoided *loads* in custom Metal kernels | | UFS 4.0: ~4 GB/s seq, ~1 GB/s 4K random (PowerInfer-2) | Apple NVMe: >6 GiB/s seq measured (M1 Max); ~2.25 GiB/s effective sparse reads at 32 KiB×32 threads — **the storage is 2–6× better than what PowerInfer-2 was built for** | | Linux O_DIRECT/io_uring | macOS F_NOCACHE + many-threaded pread; APFS page cache behaves differently (our Experiment H) | | Triton/CUDA gather kernels (TEAL, CATS, DejaVu) | **No Metal equivalents exist anywhere** (PowerInfer's Metal sparse backend: never shipped) | | Rooted Android, offline per-device planner (PowerInfer-2) | Full user control of macOS; can pin memory (wired), use mmap, run calibration at "compile" time | ### 3.4 The §4.3 key question: pruning as a cache policy — prior-art verdict Searched explicitly for "discarded weights live on SSD and are recovered on demand": - **Partially done:** LLM in a flash (sparsity-driven demand paging of FFN weights from flash — but ReLU-only, ≤2× oversubscription, no code, attention pinned); PowerInfer-2 + Ripple (same idea on Android/UFS with cluster caches and layout optimization); M2Cache (neuron LRU over HBM/DRAM/SSD with mixed precision); DIP (**mask chosen as a function of cache contents** — pruning literally becomes the cache policy, but only simulated mobile constraints); RAP (runtime pruning under memory budget, no SSD recovery); On-Demand Multi-Task Sparsity (2511.19986, edge, task-level sparse deltas from flash); VLM in a flash (2511.18692, neuron chunking for I/O-efficient VLM sparsification — shows the line is alive in 2025–26). - **Not done anywhere (verified gap):** (a) using a **one-shot pruning importance score (Wanda/SparseGPT-style) as the *static tier-assignment* policy** — "pruned" weights demoted to NVMe in an Endor-style dense-readable format and *re-materialized* when a cheap online statistic (input norms, gate outputs, flocking profile) says they matter, restoring the dense model's quality ceiling instead of accepting permanent 7B-scale pruning damage; (b) any of this on **macOS / Metal / unified memory / Apple NVMe**; (c) combination with **layer-level** granularity (ShortGPT-scored cold layers paged in only when a router detects they're needed); (d) 4-bit-quantized neuron bundles on flash (PowerInfer-2 does 4-bit on UFS; LLM in a flash used fp16 — nobody did quantized bundles on Apple NVMe with Metal decode). ### 3.5 What's unexplored (candidate experiment seeds) 1. **Rebuild the LLM-in-a-flash measurement stack on our M5 Max** (Experiment H + E): reproduce the 32 KiB×32-thread flash-read curve, then measure TEAL-style 50% SwiGLU masks as a *paging* policy (not a FLOP policy) on Llama-3-8B/Qwen-14B class models. Nobody has published bytes/token for thresholded-SwiGLU paging. 2. **Working-set statistics for SwiGLU masks** (Experiments A–C): Jaccard(t, t+1), window-union growth curves, cross-prompt domain overlap — the numbers exist only for ReLU models. 3. **Wanda-score residency tiers + Sirius-style correction**: keep top-p% weights (by |W|·‖x‖ calibration) resident, stream the rest on demand from NVMe, dense-verify occasionally. This composes three verified results into a system nobody has built, on hardware (fast NVMe + unified memory + Metal) that is *more* favorable than anything prior work targeted. --- ## Sources - Deja Vu: Contextual Sparsity for Efficient LLMs at Inference Time — https://arxiv.org/abs/2310.17157 (accessed 2026-08-11) - Deja Vu (OpenReview, ICML 2023) — https://openreview.net/forum?id=wIPIhHd00i (accessed 2026-08-11) - PowerInfer: Fast Large Language Model Serving with a Consumer-grade GPU — https://arxiv.org/abs/2312.12456 (accessed 2026-08-11) - PowerInfer (SOSP 2024 paper PDF, IPADS/SJTU) — https://ipads.se.sjtu.edu.cn/_media/publications/song-sosp24.pdf (accessed 2026-08-11) - PowerInfer GitHub (macOS/Metal support status, supported ReLU models) — https://github.com/SJTU-IPADS/PowerInfer (accessed 2026-08-11) - PowerInfer-2: Fast Large Language Model Inference on a Smartphone — https://arxiv.org/abs/2406.06282 (accessed 2026-08-11) - PowerInfer-2 project page — https://powerinfer.ai/v2/ (accessed 2026-08-11) - LLM in a flash: Efficient Large Language Model Inference with Limited Memory (Apple) — https://arxiv.org/abs/2312.11514 (accessed 2026-08-11) - ReLU Strikes Back: Exploiting Activation Sparsity in Large Language Models (Apple, ICLR 2024) — https://arxiv.org/abs/2310.04564 (accessed 2026-08-11) - ReLU Strikes Back — Apple Machine Learning Research page — https://machinelearning.apple.com/research/relu (accessed 2026-08-11) - The Lazy Neuron Phenomenon: On Emergence of Activation Sparsity in Transformers — https://arxiv.org/abs/2210.06313 (accessed 2026-08-11) - TEAL: Training-Free Activation Sparsity in Large Language Models — https://arxiv.org/abs/2408.14690 (accessed 2026-08-11) - TEAL — Together AI blog — https://www.together.ai/blog/teal-training-free-activation-sparsity-in-large-language-models (accessed 2026-08-11) - CATS: Contextually-Aware Thresholding for Sparsity in Large Language Models (COLM 2024) — https://arxiv.org/abs/2404.08763 (accessed 2026-08-11) - CATS GitHub — https://github.com/ScalingIntelligence/CATS (accessed 2026-08-11) - GRIFFIN: Prompt-prompted Adaptive Structured Pruning for Efficient LLM Generation (ICML 2024) — https://arxiv.org/abs/2404.01365 (accessed 2026-08-11) - GRIFFIN GitHub — https://github.com/hdong920/GRIFFIN (accessed 2026-08-11) - ShadowLLM: Predictor-based Contextual Sparsity for Large Language Models (EMNLP 2024) — https://arxiv.org/abs/2406.16635 (accessed 2026-08-11) - ShadowLLM — ACL Anthology — https://aclanthology.org/2024.emnlp-main.1068/ (accessed 2026-08-11) - ProSparse: Introducing and Enhancing Intrinsic Activation Sparsity within Large Language Models — https://arxiv.org/abs/2402.13516 (accessed 2026-08-11) - Turbo Sparse: Achieving LLM SOTA Performance with Minimal Activated Parameters — https://arxiv.org/abs/2406.05955 (accessed 2026-08-11) - Q-Sparse: All Large Language Models can be Fully Sparsely-Activated (NeurIPS 2024) — https://arxiv.org/abs/2407.10969 (accessed 2026-08-11) - Sirius: Contextual Sparsity with Correction for Efficient LLMs (NeurIPS 2024) — https://arxiv.org/abs/2409.03856 (accessed 2026-08-11) - SparQ Attention: Bandwidth-Efficient LLM Inference (ICML 2024) — https://arxiv.org/abs/2312.04985 (accessed 2026-08-11) - SparseGPT: Massive Language Models Can Be Accurately Pruned in One-Shot (ICML 2023) — https://arxiv.org/abs/2301.00774 (accessed 2026-08-11) - Wanda: A Simple and Effective Pruning Approach for Large Language Models (ICLR 2024) — https://arxiv.org/abs/2306.11695 (accessed 2026-08-11) - ShortGPT: Layers in Large Language Models are More Redundant Than You Expect — https://arxiv.org/abs/2403.03853 (accessed 2026-08-11) - Sheared LLaMA: Accelerating Language Model Pre-training via Structured Pruning (ICLR 2024) — https://arxiv.org/abs/2310.06694 (accessed 2026-08-11) - Compact Language Models via Pruning and Knowledge Distillation (Minitron, NVIDIA) — https://arxiv.org/abs/2407.14679 (accessed 2026-08-11) - LLM Pruning and Distillation in Practice: The Minitron Approach — https://arxiv.org/pdf/2408.11796 (accessed 2026-08-11) - SliceGPT: Compress Large Language Models by Deleting Rows and Columns (ICLR 2024) — https://arxiv.org/abs/2401.15024 (accessed 2026-08-11) - LLM-Pruner: On the Structural Pruning of Large Language Models (NeurIPS 2023) — https://arxiv.org/abs/2305.11627 (accessed 2026-08-11) - M2Cache: Harnessing Your DRAM and SSD for Sustainable and Accessible LLM Inference with Mixed-Precision and Multi-level Caching — https://arxiv.org/abs/2410.14740 (accessed 2026-08-11) - Ripple/Neuralink: Accelerating LLM Inference on Smartphones with Correlation-Aware Neuron Management / Neuron Co-Activation Linking — https://arxiv.org/abs/2410.19274 (accessed 2026-08-11) - DIP: Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking (Qualcomm AI Research) — https://arxiv.org/abs/2412.01380 (accessed 2026-08-11) - Endor: Hardware-Friendly Sparse Format for Offloaded LLM Inference — https://arxiv.org/pdf/2406.11674 (accessed 2026-08-11) - SparseInfer: Training-free Prediction of Activation Sparsity for Fast LLM Inference — https://arxiv.org/pdf/2411.12692 (accessed 2026-08-11) - Post-Training Statistical Calibration for Higher Activation Sparsity — https://arxiv.org/pdf/2412.07174 (accessed 2026-08-11) - R-Sparse: Rank-Aware Activation Sparsity for Efficient LLM Inference — https://arxiv.org/abs/2504.19449 (accessed 2026-08-11) - Spark Transformer: Reactivating Sparsity in FFN and Attention (NeurIPS 2025) — https://arxiv.org/html/2506.06644v2 (accessed 2026-08-11) - Universal Properties of Activation Sparsity in Modern Large Language Models — https://arxiv.org/abs/2509.00454 (accessed 2026-08-11) - RAP: Runtime Adaptive Pruning for LLM Inference — https://arxiv.org/pdf/2505.17138 (accessed 2026-08-11) - DuoGPT: Training-free Dual Sparsity through Activation-aware Pruning in LLMs — https://arxiv.org/html/2506.20194 (accessed 2026-08-11) - Motivating Next-Gen Accelerators with Flexible (N:M) Activation Sparsity — https://arxiv.org/pdf/2509.22166 (accessed 2026-08-11) - VLM in a flash: I/O-Efficient Sparsification of Vision-Language Model via Neuron Chunking — https://arxiv.org/html/2511.18692 (accessed 2026-08-11) - On-Demand Multi-Task Sparsity for Efficient Large-Model Deployment on Edge Devices — https://arxiv.org/pdf/2511.19986 (accessed 2026-08-11) - Fast Forward: Accelerating LLM Prefill with Predictive FFN Sparsity — https://arxiv.org/pdf/2602.00397 (accessed 2026-08-11) - Dynamic sparsity in tree-structured feed-forward layers at scale — https://arxiv.org/pdf/2604.08565 (accessed 2026-08-11)