spb/modelmap Public License
Internal cartography of local LLMs on Apple Silicon — registered, gated, negative-first. Public atlas at modelmap.io.
Python 66.3%
JavaScript 24.5%
CSS 8.1%
Shell 0.7%
-
expC run #6: L12 handle fails replication — Level 3 abandoned (as registered)
…
Only 1/4 fresh sources monotone (positive dose arm unstable); specificity failed to replicate (random |delta| 3.36 vs bound 1.14 — run #5's 0.45 was 3-draw sampling luck). Halving held 4/4: necessary but not a reliable additive handle. interventions/v1 (Level 2) stands as the final causal claim. Gate record: refuse -> pass(L2) -> refuse -> refuse. Methodology rule adopted: specificity controls need >=10 random draws + percentile bound. Charts: x-axis customization for dose-response figures. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
-
expC run #5: L3 gate refused — L12 is a textbook handle anyway
…
Conjunctive dose-response criterion failed: L12 passes everything (strictly monotone +1.95..+6.79 across -2s..+2s, halved, specific 0.45 vs bound 1.11) but L04 is an overdose regime (any 2-sigma perturbation, random included, collapses the margin) and L08 is not specific at 2s. make_l3_mapcard.py refused v2; interventions/v1 stays Level 2. Gate record: refuse (r3) -> pass (r4, L2) -> refuse (r5). Run #6 registered ideas: layer-local dose calibration; minimal L12 single-layer L3 claim. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
-
expC run #4: BAND claim passes — first Level-2 atlas entry (causal map)
…
Early-band (2-15) mean specific damage +3.23..+3.35 across six fresh direction sources (min +3.228 vs pre-registered bar 2.5; spread 0.12) on a fresh behavioral bank (baseline +4.45). Late band unstable, no claim. Published atlas/qwen3-0.6b-4bit/interventions/v1 (Level 2): a single diff-of-means agreement direction, erased at any early-band layer, removes ~73-75% of grammatical preference. The gate that refused run #3's per-layer map passed the band version. Band map leads the home page. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
-
expC run #3: publication gate refused the causal profile (as designed)
…
Replication failed: mean pairwise rho 0.495 (<0.7), band claim fails on 2/5 direction sources; make_interventions_mapcard.py exits 1 and the atlas stays at two entries. The refusal decomposes the object: early band (2-15) replicates tightly (+2.98..+3.34 of +4.24), late band (20-27) is estimator noise (-0.44..+2.77) -- retroactively explaining run #2's anti-correlation. Run #4 pre-registers the narrower BAND claim. Charts: interventions profile builder + dynamic y-axis (renders when a map passes the gate). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
-
expC run #2: survival ledger 0/2 — causal scan finds the real structure
…
P1 CONFIRMED: bottom-5 skip damage was general (specificity 0.51 < random mean 1.44), top-5 above mean but below p95 — skip never singles out the probe layers. P2 FALSIFIED: direction-erasure profile anti-correlates with the probe profile (rho=-0.136, p=0.76). Discovery: erasing the agreement diff-of-means direction at ANY layer 2-15 destroys the behavior (up to +3.97/+4.24 at L12, random-direction controls netted); late probe-ranked layers carry little; L18/L22 suppressive. The load-bearing object is an early-constructed DIRECTION, not a late place. Tap gains an edit hook. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
-
expC run #1: agreement map FAILS causal verification — survival ledger 0/1
…
Layer-skip ablation (Tap.skip): top-5 differential layers damage +2.17, BELOW random-5 mean +3.14 (p95 +4.85); bottom-5 (layers 0-4) largest at +4.99. Decodability != causal support (Hase-class dissociation, in-house). Atlas probes/v2 records the failed check (interventions + confidence.md); survival-rate ledger opened in expC analysis. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
-
Bootstrap modelmap: charter-compliant skeleton, tooling, and modelmap.io platform
…
- Repo skeleton per charter §3 with mandatory author headers everywhere - tools/check_headers.py (gates commits), new_experiment.py, new_map.py - benchmarks: hardware_manifest.py (sysctl-based) + harness skeleton - experiments/micro A-H scaffolded with empty seven-field blocks - site/: modelmap.io platform (Express, light theme, English), coherent with localvm-research/web — Home, Research, Experiments, Atlas (with confidence Levels 0-3), Results, Code, About; attribution footer Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>