spb/localvm-research Public License
Running LLMs larger than memory on a consumer Mac — falsification-driven research: margin-gated deferred refinement, out-of-core verification on Apple Silicon. TR-01 published.
Python 63.2%
JavaScript 23.5%
CSS 11.8%
Shell 0.9%
Makefile 0.5%
-
candidate_01 32B scale run: charter 16.A/B demonstrated in prototype form
…
q8-32B (does not fit beside base on 48 GB) streamed at ~11.6 GB/s sequential; +0.28 nats over pure-q4 (only fitting alternative); 3.72 GB/token logical (9.4x under checkpoint); 1.69 tok/s with measured headroom. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>