SPB Git

spb/modelmap Public License

Internal cartography of local LLMs on Apple Silicon — registered, gated, negative-first. Public atlas at modelmap.io.

Python 66.3% JavaScript 24.5% CSS 8.1% Shell 0.7%
7.3 KB

# project: modelmap document: Phase 1 notes — §4.9 Systems side — local mapping at scale on macOS author: Simon-Pierre Boucher contact: contact@spboucher.ai website: https://modelmap.io created: 2026-08-12 status: draft

# §4.9 — Systems side of local interpretability at scale

# Activation storage size math (framing)

Llama-class 7–8B (d_model 4096, 32 layers): fp16 residual capture = 8 KB/token/layer; all layers = 256 KB/token. MLP hidden (d_mlp 14336) is 3.5× larger per layer.

  • 10M tokens × 1 layer ≈ 80 GB — cacheable on a 1–2 TB SSD for one or two hookpoints.
  • SAE budget (500M tokens × 1 layer) ≈ 4 TB fp16 — pre-caching infeasible; on-the-fly generation mandatory.
  • Full-model capture for 100k tokens ≈ 25–90 GB — fine for probing/patching corpora (10⁴–10⁶ tokens). The Mac regime splits cleanly: small causal/probing corpora → cache everything; SAE-scale corpora → stream, never store.

# Storage formats

  • safetensors https://github.com/safetensors/safetensors — zero-copy mmap, lazy per-tensor; no appendable writes, no chunk index; free-form metadata dict (exploit for provenance embedding).
  • zarr https://github.com/zarr-developers/zarr-python — chunked, compressed, appendable N-D; ideal shape for token×layer×d_model with random chunk access; but dense fp16 activations are high-entropy → compression gains modest; reported slower than h5py/npy raw throughput. Benchmark on macOS, don't assume.
  • HDF5 — mature; single-writer pain; little advantage over zarr/mmap here.
  • Raw mmap / np.memmap — best random-access latency (the pattern SAE shuffling needs); on unified memory + Apple NVMe likely the throughput winner; needs external index + own content hashes. No published macOS numbers exist for any of these — Experiment H claims this gap.

# SAE training pipelines (streaming on limited memory)

  • SAELens https://github.com/decoderesearch/SAELens — on-the-fly buffer shuffle (approximate); optional CacheActivationsRunner (safetensors shards); buffer size = the RAM knob; no provenance schema for caches; shuffle quality is a replication variable modelmap should control.
  • EleutherAI sparsify https://github.com/EleutherAI/sparsify — TopK SAEs/transcoders with no activation caching at all; bitsandbytes 8-bit path is CUDA-only (useless on MPS).
  • Anthropic (disclosed): collect billions → full on-disk shuffle → stream (https://transformer-circuits.pub/2024/scaling-monosemanticity/). The gold standard a Mac cannot replicate. Open testable question: how much does buffer-shuffled streaming degrade feature replication vs full shuffles? Testable locally at 0.5–3B.
  • Budget: (weights + KV cache + buffer + SAE) < unified memory. 3B fp16 + 2 GB buffer + 131k SAE fits 16 GB; 8B needs 32 GB+.

# Hooks and capture paths

  • PyTorch MPS: forward hooks fire normally (eager); costs: forced syncs on .cpu() copies, operator gaps (PYTORCH_ENABLE_MPS_FALLBACK=1 silently round-trips via CPU), no float64, device_map="auto" CPU offload does not work on MPS — model must fit unified memory. fp16 numerics on MPS can differ from CUDA — a confound when comparing to published results. Hook overhead unmeasured in literature (Experiment H).
  • TransformerLens https://github.com/TransformerLensOrg/TransformerLens — HookPoints everywhere; MPS opt-in; materializes all hook activations eagerly (memory-hungry); HF checkpoints only, no GGUF.
  • MLX / mlx_lm: lazy evaluation (https://ml-explore.github.io/mlx/build/html/usage/lazy_evaluation.html) — intermediates can be retained for free during graph construction; early materialization wrecks throughput; no built-in hook API — introspection = wrapping modules (plain Python, easy). MLX runs 4-bit quantized models natively → the only realistic path to quantized-model activation capture with a Python API (llama.cpp aside). M5 neural-accelerator work: https://machinelearning.apple.com/research/exploring-llms-mlx-m5
  • mlxterp https://github.com/coairesearch/mlxterp — the one existing MLX mechinterp library: model.trace() context manager, ~196 activations/forward, composable interventions, mlx-lm model support. Very early (~11 stars); no streaming store, no SAE training, no provenance, no overhead benchmarks. Direct prior art for modelmap's capture layer — evaluate before building.

# Quantized-runtime introspection (llama.cpp)

  • eval-callback https://github.com/ggml-org/llama.cpp/tree/master/examples/eval-callback — ggml graph callback intercepting every op; can dump every intermediate tensor during inference including the Metal backend — real C-level activation extraction from actually-quantized compute.
  • imatrix — per-channel mean-squared activation statistics guiding quantization; itself a crude weight-importance map.
  • Verified absent: no interpretability tooling wraps this. A thin exporter (eval-callback → mmap store) would give Q8/Q4/Q2 activation capture no Python framework offers — key for Experiment F. Caveats: dequantized dumps op-by-op have real overhead; no token bookkeeping; graph names shift across versions — pin a commit.

# Metal capture & profiling

  • MTLCaptureManager (https://developer.apple.com/documentation/metal/mtlcapturemanager) records .gputrace (30k+ buffers per forward) — a debugging tool, not a data pipeline. Instruments' Metal System Trace for kernel timings/occupancy → Experiment H instrumentation, and for validating that Python-level capture matches GPU reality.

# The competing philosophy: NNsight + NDIF

  • https://arxiv.org/abs/2407.14561 (ICLR 2025); https://ndif.us/ ; European replication eDIF: https://arxiv.org/pdf/2508.10553
  • Deferred-execution intervention graphs shipped to a shared GPU fabric — don't run locally; send the experiment away. modelmap's antithesis and best contrast case: NDIF serves unquantized checkpoints and cannot study the quantized runtimes people actually use locally; reproducibility depends on a remote service. NNsight's deferred-graph idiom is conceptually close to MLX laziness — study for API design.

# Disk/IO on macOS

  • APFS: native sparse files (pre-sized mmap stores written out of order allocate nothing); copy-on-write clones = instant snapshot of an activation store — a genuinely nice provenance primitive; metadata-only checksums (data corruption NOT detected — store own content hashes). Refs: https://eclecticlight.co/2024/06/08/apfs-how-sparse-files-work/ ; https://eclecticlight.co/2021/03/29/sparse-files-are-common-in-apfs/
  • fio-style numbers for mmap workloads on Apple NVMe essentially unpublished — measure, don't import Linux ext4 folklore. (Sister project localvm-research already measured cold-cache NVMe: ~13.1 GB/s ceiling at ≥256 KiB blocks, QD≥4 — reuse.)

# Verified open gaps (modelmap's to claim)

  1. No published MLX-vs-MPS hook-overhead benchmarks.
  2. No macOS activation-store format benchmarks (mmap vs zarr vs safetensors on APFS).
  3. No quantized-runtime interpretability tooling (llama.cpp eval-callback unwrapped).
  4. Buffer-shuffle vs full-shuffle effect on SAE feature replication unpublished.
  5. 2026 systems-side interest appearing (https://arxiv.org/pdf/2605.11093 — model-internal observability for LLM inference) — move fast.