SPB Git

spb/modelmap Public License

Internal cartography of local LLMs on Apple Silicon — registered, gated, negative-first. Public atlas at modelmap.io.

Python 66.3% JavaScript 24.5% CSS 8.1% Shell 0.7%

History of research/LOG.md · clear filter

  1. expC run #6: L12 handle fails replication — Level 3 abandoned (as registered)
    Only 1/4 fresh sources monotone (positive dose arm unstable); specificity
    failed to replicate (random |delta| 3.36 vs bound 1.14 — run #5's 0.45 was
    3-draw sampling luck). Halving held 4/4: necessary but not a reliable
    additive handle. interventions/v1 (Level 2) stands as the final causal
    claim. Gate record: refuse -> pass(L2) -> refuse -> refuse. Methodology
    rule adopted: specificity controls need >=10 random draws + percentile
    bound. Charts: x-axis customization for dose-response figures.
    
    Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
    Simon-Pierre Boucher committed 3 h ago (Aug 12, 2026) · 1 file changed +27
  2. expC run #5: L3 gate refused — L12 is a textbook handle anyway
    Conjunctive dose-response criterion failed: L12 passes everything
    (strictly monotone +1.95..+6.79 across -2s..+2s, halved, specific 0.45
    vs bound 1.11) but L04 is an overdose regime (any 2-sigma perturbation,
    random included, collapses the margin) and L08 is not specific at 2s.
    make_l3_mapcard.py refused v2; interventions/v1 stays Level 2. Gate
    record: refuse (r3) -> pass (r4, L2) -> refuse (r5). Run #6 registered
    ideas: layer-local dose calibration; minimal L12 single-layer L3 claim.
    
    Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
    Simon-Pierre Boucher committed 3 h ago (Aug 12, 2026) · 1 file changed +28
  3. expC run #4: BAND claim passes — first Level-2 atlas entry (causal map)
    Early-band (2-15) mean specific damage +3.23..+3.35 across six fresh
    direction sources (min +3.228 vs pre-registered bar 2.5; spread 0.12) on
    a fresh behavioral bank (baseline +4.45). Late band unstable, no claim.
    Published atlas/qwen3-0.6b-4bit/interventions/v1 (Level 2): a single
    diff-of-means agreement direction, erased at any early-band layer,
    removes ~73-75% of grammatical preference. The gate that refused run #3's
    per-layer map passed the band version. Band map leads the home page.
    
    Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
    Simon-Pierre Boucher committed 3 h ago (Aug 12, 2026) · 1 file changed +31
  4. expC run #3: publication gate refused the causal profile (as designed)
    Replication failed: mean pairwise rho 0.495 (<0.7), band claim fails on
    2/5 direction sources; make_interventions_mapcard.py exits 1 and the
    atlas stays at two entries. The refusal decomposes the object: early band
    (2-15) replicates tightly (+2.98..+3.34 of +4.24), late band (20-27) is
    estimator noise (-0.44..+2.77) -- retroactively explaining run #2's
    anti-correlation. Run #4 pre-registers the narrower BAND claim.
    Charts: interventions profile builder + dynamic y-axis (renders when a
    map passes the gate).
    
    Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
    Simon-Pierre Boucher committed 3 h ago (Aug 12, 2026) · 1 file changed +25
  5. expC run #2: survival ledger 0/2 — causal scan finds the real structure
    P1 CONFIRMED: bottom-5 skip damage was general (specificity 0.51 < random
    mean 1.44), top-5 above mean but below p95 — skip never singles out the
    probe layers. P2 FALSIFIED: direction-erasure profile anti-correlates with
    the probe profile (rho=-0.136, p=0.76). Discovery: erasing the agreement
    diff-of-means direction at ANY layer 2-15 destroys the behavior (up to
    +3.97/+4.24 at L12, random-direction controls netted); late probe-ranked
    layers carry little; L18/L22 suppressive. The load-bearing object is an
    early-constructed DIRECTION, not a late place. Tap gains an edit hook.
    
    Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
    Simon-Pierre Boucher committed 3 h ago (Aug 12, 2026) · 1 file changed +32
  6. expC run #1: agreement map FAILS causal verification — survival ledger 0/1
    Layer-skip ablation (Tap.skip): top-5 differential layers damage +2.17,
    BELOW random-5 mean +3.14 (p95 +4.85); bottom-5 (layers 0-4) largest at
    +4.99. Decodability != causal support (Hase-class dissociation, in-house).
    Atlas probes/v2 records the failed check (interventions + confidence.md);
    survival-rate ledger opened in expC analysis.
    
    Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
    Simon-Pierre Boucher committed 4 h ago (Aug 12, 2026) · 1 file changed +35
  7. expA run #2: first positive maps (agreement, arith_valid) + real noise floor
    - v2 promptsets: structure-borne, token-balanced (overlap certificates);
      capture_pooled returns mean+last-token reps in one pass
    - Strict twin gate falsified again (word_order twin acc 0.96 — subword
      statistics); gate retired, differential (real-twin) is the map
    - agreement: 25/28 signal layers (max dSel +0.379); arith_valid: real
      0.86/0.90 vs twin 0.56, strongest last-token; word_order null-dominated
    - First real noise floor: top-5 replication ~0.54 off ceiling; dataset
      shift > seed SD on only 3-8/28 layers (registered 1/3 not met)
    - atlas probes/v2 published (Level 1, per-property verdicts); site charts
      generalized to any probes entry; home leads with agreement map
    
    Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
    Simon-Pierre Boucher committed 4 h ago (Aug 12, 2026) · 1 file changed +38
  8. expA run #1: validity gate FAILED — first atlas entry is a negative result
    - src/modelmap/capture: MLX Tap layer + random-init architecture twin
    - promptsets v1 (6x240, sha256 manifest) — now the positive-control corpus
    - Result: ceiling everywhere (acc 1.000); twin ALSO at 1.00 on all 28
      layers (max sel 0.56-0.88 vs registered 0.05 gate) — probe map reads
      tokenizer+architecture prior, zero trained-model signal. Shuffled-label
      control would not have caught it; only the twin null did.
    - Published atlas/qwen3-0.6b-4bit/probes/v1 (negative_result=true,
      Level 1, replication 1.00) via publish.py gate
    - Doctrine: probe maps only as REAL-TWIN differentials from now on
    
    Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
    Simon-Pierre Boucher committed 4 h ago (Aug 12, 2026) · 1 file changed +41
  9. expH runs #2-#3: first falsified hypothesis + first quantized capture
    Run #2 (cold cache, M3U96a, purge/repeat): FALSIFIED — warm ordering
    inverts; zarr-uncompressed 0.62 GB/s beats raw-mmap 0.14 on cold random
    batches (page-fault QD1 IO vs 32MiB chunk reads). Rule: IO granularity
    decides, not the container. Store design revised; run #4 registered.
    
    Run #3 (mlx-lm Qwen3-0.6B-4bit): CONFIRMED — retain 1.004x plain prefill
    (capture free under lazy eval), retain+write 1.28x. First quantized-model
    activation capture in Python tooling.
    
    Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
    Simon-Pierre Boucher committed 4 h ago (Aug 12, 2026) · 1 file changed +37
  10. Phase 5 substrate + expH run #1: first measured result
    - mapcard v0 schema + publish gate (G23), stats core (bootstrap/FDR/
      replication), probe harness with built-in shuffled-label controls
      (unit test caught a fancy-indexing shuffle bug pre-science)
    - expH hypothesis registered before run; run #1 on M5 Max 48GB:
      (A) mmap 3.2-10.8x zarr on random-batch reads (confirmed >=2x bar);
      (B) MLX capture nearly free (1.02x), torch-MPS retain 1.22x (confirmed)
    - expA hypothesis registered (dataset-variance > seed-variance)
    
    Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
    Simon-Pierre Boucher committed 5 h ago (Aug 12, 2026) · 1 file changed +45
  11. Phase 4: candidate ranking — 24 gaps scored on ten axes, 4 candidates selected
    candidate_01 weight-only pre-screen (G06+G08/G03/G09), candidate_02
    quantization deformation atlas (G01+G02), candidate_03 localization index
    (G17+G11/G12), candidate_04 localvm working-set bridge (G24, reduced).
    Mandatory substrate scheduled first: expH, expA/expD noise floors, map
    cards v0, ablation-curves rule.
    
    Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
    Simon-Pierre Boucher committed 5 h ago (Aug 12, 2026) · 1 file changed +41
  12. Phase 3: 24 research gaps (charter §6)
    Seven clusters: quantization×internals, weight-only pre-screens,
    replication/method-agreement, cross-model coordinates, localization
    science, systems/tooling, atlas methodology + localvm bridge. Each gap
    carries the six-field block with a smallest falsifying Mac experiment.
    
    Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
    Simon-Pierre Boucher committed 5 h ago (Aug 12, 2026) · 1 file changed +31
  13. Phase 2: state-of-the-art map (charter §5)
    ~40 techniques, six families + cross-cutting instruments, eleven-field
    blocks with epistemic flags; overlap analysis, already-tried combinations,
    novelty traps, estimated Mac cost frontier (16/32/64 GB).
    
    Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
    Simon-Pierre Boucher committed 5 h ago (Aug 12, 2026) · 1 file changed +35
  14. Phase 1 sweep #1: ten-area literature survey (§4.1-4.10)
    - research/notes/: 10 theme notes with verified sources, Apple Silicon
      status, failure modes, epistemic status per technique
    - research/bibliography.md: ~200 sources with URLs + access dates
    - research/LOG.md: sweep results — quantization gap confirmed (~5 shallow
      papers), weight-only pre-screen gap open, noise floors (30% SAE seed
      overlap, 1-5% neuron universality), atlas provenance gap verified
    
    Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
    Simon-Pierre Boucher committed 5 h ago (Aug 12, 2026) · 1 file changed +55
  15. Bootstrap modelmap: charter-compliant skeleton, tooling, and modelmap.io platform
    - Repo skeleton per charter §3 with mandatory author headers everywhere
    - tools/check_headers.py (gates commits), new_experiment.py, new_map.py
    - benchmarks: hardware_manifest.py (sysctl-based) + harness skeleton
    - experiments/micro A-H scaffolded with empty seven-field blocks
    - site/: modelmap.io platform (Express, light theme, English), coherent
      with localvm-research/web — Home, Research, Experiments, Atlas (with
      confidence Levels 0-3), Results, Code, About; attribution footer
    
    Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
    Simon-Pierre Boucher committed 6 h ago (Aug 12, 2026) · 1 file changed +35