project: modelmap document: expH_capture_cost_frontier — hypothesis (run #1, registered before run) author: Simon-Pierre Boucher contact: contact@spboucher.ai website: https://modelmap.io created: 2026-08-12 modified: 2026-08-12 status: reviewed
Hypothesis — expH run #1 (storage formats + hook overhead)
Capture cost frontier — what can you map on which Mac (macOS-specific). Run #1 scope: (A) activation-store formats on APFS, (B) capture/hook overhead in MLX vs PyTorch-MPS on a synthetic transformer. Registered 2026-08-12 before the first benchmark execution.
Hypothesis : (A) For the SAE-shuffle access pattern (random row
batches), raw np.memmap sustains ≥ 2× the read
throughput of chunked zarr on APFS (warm cache),
because dense fp16 activations gain little from
compression while chunk decode adds latency.
(B) Activation capture (retain per-layer outputs)
costs < 2× plain inference on both MLX and
PyTorch-MPS at the 12-layer/d1024 scale, and the
device→CPU copy + disk write — not the hook/retain
mechanism — dominates capture cost.
Falsification criterion : (A) dies if zarr (any codec) matches or beats mmap
on random-batch reads (ratio < 1.25×).
(B) dies if retain-only capture overhead exceeds
2× plain inference on either backend; the "capture
is cheap on consumer Macs" premise dies at > 3×
including copy+write.
Method : Storage: 100k rows × 4096 dims fp16 (~0.8 GB) per
format (raw mmap, safetensors-mmap, zarr zstd, zarr
uncompressed); sequential write; sequential scan;
random-batch reads (4096 rows/batch, SAE-shuffle
pattern); 3 repeats/cell. Compute: synthetic
pre-norm transformer (12 layers, d_model 1024,
4 heads, seq 512, batch 8, fp16) implemented
identically in torch-MPS and MLX; modes = plain /
retain-on-device / retain+CPU-copy+mmap-write;
20 timed forwards (10 for write mode), 3 repeats,
explicit device synchronization.
Baseline / null : plain forward (no capture) per backend; sequential
scan as the reference for random-read degradation.
Result : (pending)
Interpretation : (pending — with explicit confidence level)
Next experiment : (pending)Known limitations, declared in advance. Warm-cache regime (the 0.8 GB store fits in the unified buffer cache): read numbers upper-bound the cold case; a purge/F_NOCACHE cold pass is a registered follow-up (run #2), and the sister project's cold-cache NVMe ceiling (localvm expH: ~13.1 GB/s at ≥256 KiB, QD≥4) bounds what run #2 can find. Synthetic model ≠ real checkpoint: run #3 will repeat mode timings on a real 0.5B model via mlx-lm.
Hypothesis — expH run #2 (cold-cache storage, second hardware)
Registered 2026-08-12 before the run. Host: MacLustr node M3U96a
(Mac Studio, 32 cores, 96 GB), where sudo purge empties the unified buffer
cache between repetitions — also our first cross-hardware replication point.
Hypothesis : The warm-cache format ordering survives cold:
raw mmap sustains ≥ 1.5× zarr-uncompressed on
random-batch reads with a purged cache; absolute
throughputs drop for all formats; zarr-zstd's
relative gap narrows (decode overlaps disk IO)
but does not close.
Falsification criterion : Dies if any zarr variant matches or beats mmap
cold (ratio < 1.25×) — that would rehabilitate
chunked stores and overturn run #1's design call.
Method : Same grid as run #1 Part A, store enlarged to
200k×4096 fp16 (~1.6 GB), `sudo purge` before
every timed read repetition, 3 repeats/cell,
writes measured warm (unchanged claim).
Baseline / null : run #1 warm numbers on M5 Max; sequential scan
as the random-read reference.
Result : (pending)
Interpretation : (pending)
Next experiment : (pending)Hypothesis — expH run #3 (real quantized checkpoint via mlx-lm)
Registered 2026-08-12 before the run. Local M5 Max, model = mlx-community Qwen3-0.6B (4-bit), capture by wrapping decoder layers.
Hypothesis : On a real 0.6B 4-bit model under MLX, per-layer
residual retention costs ≤ 1.15× plain prefill
(lazy-graph retention, as in run #1's synthetic
result), and retention + NumPy conversion + mmap
write stays ≤ 1.5×.
Falsification criterion : Dies at retain > 1.5× or retain+write > 2× —
would mean synthetic overheads do not transfer to
real checkpoints and all run #1 Part B numbers
must be re-derived on real models.
Method : Prefill of a fixed 1024-token prompt, batch 1,
10 timed forwards after 3 warmups, 3 repeats;
modes = plain / retain-all-layers / retain+write;
mx.eval barriers around each forward.
Baseline / null : plain prefill throughput (tok/s).
Result : (pending)
Interpretation : (pending)
Next experiment : (pending)