SPB Git

spb/localvm-research Public License

Running LLMs larger than memory on a consumer Mac — falsification-driven research: margin-gated deferred refinement, out-of-core verification on Apple Silicon. TR-01 published.

Python 63.2% JavaScript 23.5% CSS 11.8% Shell 0.9% Makefile 0.5%

# project: localvm-research document: expH_ssd_feasibility/hypothesis author: Simon-Pierre Boucher contact: contact@spboucher.ai created: 2026-08-11 modified: 2026-08-12 status: draft

# Hypothesis — expH_ssd_feasibility

text
Hypothesis
  The internal Apple NVMe SSD (AP2048Z, 2 TB) sustains enough *random* read
  bandwidth at weight-block-sized granularities (256 KB–4 MB) to stream a
  meaningful fraction of model weights per token: we expect ≥ 2 GB/s random
  reads at ≥ 1 MB blocks with F_NOCACHE (true storage path), degrading
  sharply below 64 KB, and only modest degradation under concurrent Metal
  GPU compute (unified memory contention is expected to be small relative
  to SSD ceiling).

Falsification criterion
  If uncached random reads at 1 MB blocks sustain < 500 MB/s, or collapse
  by > 50% under concurrent GPU matmul load, SSD weight streaming cannot
  supply even ~0.5 GB/token at interactive rates and candidate designs
  must not assume per-token SSD reads on this class of hardware.

Method
  Create an 8 GiB incompressible test file on the internal APFS volume.
  Measure read throughput at block sizes {4 KB, 16 KB, 64 KB, 256 KB,
  1 MB, 4 MB} × {sequential, random} × {F_NOCACHE on, off/warm cache} ×
  {1, 4, 8 reader threads}. Cold-cache condition enforced via F_NOCACHE
  (macOS has no O_DIRECT); warm condition by pre-reading the file.
  Repeat each cell 3 times, report mean/median/std. Then repeat the key
  uncached random cells while an MLX fp16 4096×4096 matmul loop saturates
  the GPU. Record hardware manifest and thermal level before/after.

Baseline
  Apple's advertised sequential throughput for this SSD class (~5-7 GB/s)
  and the sequential-read measurement from the same harness (internal
  baseline; no straw men — we compare random vs our own sequential).

Result
  Hypothesis CONFIRMED and understated. Cold random reads: 67 MB/s (4 KiB
  QD1) → 13.8 GB/s (1 MiB QD8, at the iostat-validated ~13.1 GB/s device
  ceiling). GPU load cost <5%. 4 MiB and threaded-sequential cells were
  cache-contaminated and are flagged invalid. Full numbers:
  results/expH_ssd_feasibility/20260812T034359Z/ and analysis.md.

Interpretation
  SSD streaming is viable at ~650 MB/token (realistic) if and only if
  blocks are ≥256 KiB at QD≥4. Prefetch can overlap Metal compute freely.
  See analysis.md for the full argument and the measurement-discipline
  findings (F_NOCACHE semantics).

Next experiment
  expD/expG — the substrate is not the bottleneck; whether ≤650 MB/token of
  useful bytes preserves behavior is. Secondary: rerun sequential cells with
  a 64 GiB (> RAM) file.