SPB Git

spb/modelmap Public License

Internal cartography of local LLMs on Apple Silicon — registered, gated, negative-first. Public atlas at modelmap.io.

Python 66.3% JavaScript 24.5% CSS 8.1% Shell 0.7%

# project: modelmap document: expC_causal_verification — analysis (run #1) author: Simon-Pierre Boucher contact: contact@spboucher.ai website: https://modelmap.io created: 2026-08-12 modified: 2026-08-12 status: reviewed

# Analysis — expC run #1: the agreement probe map FAILS causal verification

Run: results/expC_causal_verification/20260812T064534Z/results.json (5.7 s). Model: Qwen3-0.6B-4bit. Input map: agreement differential (probes/v2). Intervention: layer-skip ablation via the Tap layer. Hypothesis and the binary verdict criterion registered before the run.

text
Hypothesis              : the top-5 differential layers (by real−twin probe
                          selectivity) causally support agreement behavior —
                          skip damage ≥ p95 of random-5 draws AND ≥ 2× their
                          mean.
Falsification criterion : top-5 damage inside the random distribution.
Result                  : FALSIFIED — survival 0/1. Baseline grammatical
                          margin +4.63 (sanity holds: the model robustly
                          prefers correct agreement). Skip damage:
                          top-5 (layers 17,18,19,21,22) = +2.17 —
                          BELOW the random-5 mean (+3.14, p95 +4.85, 20
                          draws); bottom-5 differential (layers 0–4) = +4.99,
                          the LARGEST of all conditions.
Interpretation          : Level 1 for the negative claim. Where agreement
                          information is most linearly decodable above the
                          architecture null (late-mid layers) is NOT where
                          the computation is causally load-bearing for the
                          behavior. This is the Hase-class dissociation
                          (localization ≠ causal support — notes §4.6)
                          measured end-to-end in our own pipeline, on a
                          pre-registered binary verdict. The
                          correlational→causal survival ledger opens at 0/1.
                          Declared caveats bite exactly as registered:
                          (a) layer-skip is coarse — early-layer skips
                          plausibly cause GENERAL degradation, not
                          agreement-specific damage (bottom-5 +4.99 reads as
                          "the model breaks", not "agreement lives at layers
                          0–4"); a specificity control (margin damage
                          normalized by general perplexity damage) is the
                          registered follow-up; (b) held-out pairs = 89
                          after dedup against every v2 probe text (below the
                          planned 200 — combo pool exhausted; enlarge banks
                          next run).
Next experiment         : run #2 with (1) perplexity-normalized specificity
                          scores per skip condition, (2) direction-level
                          intervention (LEACE erasure of the agreement
                          direction in the forward pass at layer ℓ) — a
                          surgical test the probe map CAN legitimately pass
                          or fail, (3) larger held-out bank. Then the same
                          protocol on arith_valid.

# Ledger

# Correlational claim Intervention Survives?
1 agreement top-5 differential layers (probes/v2) layer-skip ablation No (0/1)
2 agreement differential layer profile (probes/v2) direction-level erasure scan No (0/2)

The published survival rate is the running fraction of this table — the charter §8.3 metric, now live.


# Analysis — expC run #2: ledger 0/2 — and the control scan finds the real structure

Run: results/expC_causal_verification/20260812T065406Z/results.json (192 held-out pairs, baseline margin +4.24, NLL 5.73). Hypotheses registered before the run.

text
Hypothesis (P1)         : run #1's bottom-5 damage is general, not
                          agreement-specific, once normalized by NLL damage.
Result (P1)             : CONFIRMED. Specificity (margin damage / NLL
                          damage): bottom-5 = 0.51 — BELOW the random-5 mean
                          (1.44); top-5 = 2.51 — above the random mean but
                          below the p95 (4.73). Run #1's verdict stands, now
                          with the confound measured: early-layer skips
                          break the model generally; nothing in the skip
                          family singles out the probe map's top layers.
Hypothesis (P2)         : per-layer direction-erasure specific damage
                          correlates with the differential probe profile
                          (Spearman ρ ≥ 0.4, perm-p < 0.05).
Result (P2)             : FALSIFIED, decisively — ρ = −0.136, perm-p 0.757.
                          Survival ledger: 0/2. The probe map's layer
                          ranking is behaviorally void at this granularity.
                          BUT the scan itself uncovered strong causal
                          structure the probe map missed: erasing the
                          layer-ℓ agreement direction (diff-of-means, with
                          random-direction controls netted out) at ANY
                          single layer in 2–15 destroys most of the margin
                          (specific damage +2.6…+4.0 of a +4.24 baseline;
                          L12: +3.97, L15: +3.75, L6: +3.71), while the
                          late layers the probes ranked highest carry
                          little (L21: +0.69, L20: +0.29) — and L18/L22
                          erasure slightly HELPS (−1.86/−1.09), suggesting
                          suppressive components.
Interpretation          : Level 0–1 (single model, single direction
                          estimate from set A, no seed replication of the
                          intervention yet — by our own doctrine this
                          causal profile is NOT publishable as an atlas
                          entry until replicated). Two lessons:
                          (1) correlational layer rankings did not survive
                          two different causal tests — the survival rate
                          the field never publishes is, so far, 0%;
                          (2) the behaviorally load-bearing object is a
                          low-dimensional DIRECTION present across
                          early-mid layers, not a "place" — consistent with
                          the linear-representation view and with why
                          decodability peaks (late, after information is
                          everywhere) diverge from causal joints (early,
                          where the direction is constructed). Random-
                          direction erasure also hurts at layers 2–9
                          (+0.9…+2.5): early residual streams are fragile
                          to ANY rank-1 deletion — netted out in the
                          specific column.
Next experiment         : run #3 — replicate the direction-erasure profile
                          (direction re-estimated on set B and on 5
                          bootstrap seeds; report profile replication rate)
                          → if stable, publish as the atlas's first
                          INTERVENTIONS map (Level toward 2–3) and make the
                          probes/v2 entry point to it as the causal
                          counterpart. Then the same scan on arith_valid.

# Analysis — expC run #3: the publication gate REFUSED the causal profile

Run: results/expC_causal_verification/20260812T065929Z/results.json. Five direction sources (set A, set B, 3 bootstraps of A), shared random-direction controls, same held-out bank (192 pairs, baseline +4.24). Hypothesis, criterion, and publication rule registered before the run.

text
Hypothesis              : the per-layer specific-damage profile is
                          estimator-stable — mean pairwise Spearman ρ ≥ 0.7
                          and early(2–15) ≥ 3× late(20–27) in EVERY source.
Result                  : FALSIFIED — mean ρ = 0.495 (min 0.176); the band
                          claim fails for setB and bootA2 (ratio 1.2).
                          make_interventions_mapcard.py refused the entry
                          (exit 1), per the registered rule. The atlas
                          stays at two entries — the gate did its job.
                          THE STABLE PART: early-band (layers 2–15) mean
                          specific damage replicates tightly across all
                          five sources (+2.98, +3.22, +3.34, +3.27, +3.28)
                          — erasing the estimated agreement direction
                          anywhere in the early-mid band reliably destroys
                          ~70–79% of the behavior, whatever the estimation
                          set. THE UNSTABLE PART: the late band (20–27)
                          swings from −0.44 to +2.77 with the estimator —
                          the late-layer "profile" is direction-estimation
                          noise, which also retroactively explains run #2's
                          anti-correlation (the probe ranking lives exactly
                          where the causal profile is noise).
Interpretation          : Level 0–1. The full profile is NOT a stable
                          object; the coarser claim ("an early-band
                          agreement direction is causally load-bearing") is
                          the replicable candidate. Publishing gates that
                          refuse are the mechanism that keeps the atlas
                          honest — this refusal is itself a process result
                          worth reporting on the site's methodology page.
Next experiment         : run #4 with the NARROWER pre-registered claim:
                          early-band (2–15) mean specific damage ≥ 2.5 on
                          fresh direction estimates (new bootstrap seeds +
                          a held-out estimation split) and a fresh
                          behavioral bank; late band explicitly excluded as
                          unstable. If it passes, publish the interventions
                          map as an early-band BAND claim (not a per-layer
                          ranking), Level 2.

# Analysis — expC run #4: the BAND claim passes — first Level-2 atlas entry

Run: results/expC_causal_verification/20260812T071954Z/results.json. Everything fresh: six direction sources (disjoint halves of A and B + two new bootstraps), new behavioral bank (8 unseen locations, 192 pairs, baseline margin +4.45). Claim, bar, and publication rule registered before the run.

text
Hypothesis              : early-band (2–15) mean specific damage ≥ 2.5 in
                          EVERY fresh source; late band reported, no claim.
Result                  : CONFIRMED — early-band means: Ahalf1 +3.35,
                          Ahalf2 +3.23, bootA +3.30, Bhalf1 +3.26,
                          Bhalf2 +3.33, bootB +3.27 (min +3.228 vs bar
                          2.5). Spread across six estimators: 0.12 — the
                          band is a tight, estimator-stable causal object.
                          Late band again unstable (+0.27…+2.83), as
                          declared; it carries no claim.
Interpretation          : Level 2, published: atlas/qwen3-0.6b-4bit/
                          interventions/v1 — the project's first Level-2
                          entry and first causal map. The claim: a single
                          linear direction (diff-of-means over correct vs
                          violated agreement), erased at ANY one layer in
                          2–15, removes ~73–75% of the model's grammatical
                          preference on held-out pairs, replicated across
                          six independent direction estimates and netted
                          against random-direction damage. NOT Level 3:
                          both agreeing methods share the diff-of-means
                          estimator — activation-addition steering is the
                          registered L3 path. Granularity discipline paid
                          off: the per-layer version of this map was
                          refused (run #3); the band version replicates.
Next experiment         : (1) L3 path: steering (add the direction) should
                          INCREASE the margin on violated-preference pairs;
                          (2) same band protocol on arith_valid; (3)
                          quantization drift of the band map (candidate_02:
                          does the early band move under Q8/FP16?);
                          (4) cross-model: does the band replicate on
                          Qwen3-1.7B (expG entry)?

# Analysis — expC run #5: Level-3 gate refused — and layer 12 is a textbook handle

Run: results/expC_causal_verification/20260812T072455Z/results.json. Steering doses α·σℓ·u at layers {4, 8, 12}, α ∈ {−2…+2}; conjunctive criterion (monotone AND halved AND specific at EVERY test layer) registered before the run. make_l3_mapcard.py refused (exit 1); v1 stays Level 2.

text
Hypothesis              : strict dose-response monotonicity + halving +
                          random-direction specificity at all of L4/L8/L12.
Result                  : FALSIFIED as a conjunction.
                          L12 — PASSES EVERYTHING, textbook: margins
                          +1.95 / +3.00 / +4.45 / +5.82 / +6.79 across
                          −2σ…+2σ (strictly monotone), halved at −2σ,
                          random-direction |Δ| = 0.45 vs bound 1.11.
                          L08 — near-monotone (+2σ dips: 5.70 < 6.26),
                          random |Δ| 2.97: not specific at this dose.
                          L04 — OVERDOSE REGIME: ±2σ both collapse the
                          margin (+0.08 / +0.04) and random directions are
                          equally destructive (|Δ| 4.40): at early layers,
                          ANY perturbation of magnitude ~2σ (σ estimated
                          from mean-pooled reps) breaks the computation —
                          the rank-1 fragility of run #2, now dose-resolved.
Interpretation          : Level 0–1. The direction is a clean, dose-
                          controlled causal handle at mid-band (L12) but
                          the conjunctive claim over the whole test set
                          fails, so the gate held v2 back — correctly.
                          Reading: "necessity everywhere in the band"
                          (erasure, L2 entry) coexists with "controllable
                          handle only where the layer tolerates
                          perturbation". Dose scale is a confound at early
                          layers: σ from mean-pooled statistics likely
                          overdoses positions at layers with different
                          norm profiles.
Next experiment         : run #6 (registered idea): layer-local dose
                          calibration (σ from position-level projections at
                          the target layer; or a dose-sweep to find each
                          layer's non-destructive range), then re-register
                          the handle claim on the sub-band that tolerates
                          calibrated doses (candidate: 10–14). Also register
                          the L12 single-layer handle claim with fresh
                          direction estimates as a minimal L3 candidate.

# Gate record (for methodology.md)

Three refusals/passes to date, all pre-registered: run #3 per-layer profile REFUSED → run #4 band claim PASSED (Level 2 published) → run #5 L3 conjunction REFUSED (v1 unchanged). The atlas never received a claim its evidence didn't carry.


# Analysis — expC run #6: the L12 handle claim fails replication — Level 3 abandoned

Run: results/expC_causal_verification/20260812T072913Z/results.json. Four fresh direction sources at L12, third fresh bank (192 pairs, baseline +4.57), fresh random-direction seeds. Registered rule: any source failing monotonicity/halving, or specificity failing, kills the L3 claim and closes the steering program for this object.

text
Hypothesis              : L12 is an estimator-stable dose-controlled handle.
Result                  : FALSIFIED. Monotone: 1/4 sources only (Ahalf2);
                          the POSITIVE dose arm is unstable (Ahalf1 +2σ dips
                          5.59<6.29; Bhalf1 +1σ dips below baseline; Bhalf2
                          +2σ < +1σ). Halving at −2σ: 4/4 — the negative
                          (erasure-like) arm is robust, again. SPECIFICITY
                          FAILED TO REPLICATE: random-direction mean |Δ| =
                          3.36 vs bound 1.14 on the fresh bank and fresh
                          random seeds (run #5's L12 value was 0.45 — with
                          only 3 random draws, that pass now reads as
                          sampling luck).
Interpretation          : Level 1 for the negative. The direction is
                          NECESSARY (v1, Level 2, band-replicated) but NOT
                          a reliable additive handle: pushing along it does
                          not control the behavior in a dose-stable,
                          direction-specific way. Per the pre-registered
                          rule, Level 3 is abandoned for this object and
                          the steering program is closed. Had run #5's L12
                          observation been published without fresh
                          re-registration, the atlas would now contain a
                          false Level-3 claim — the gate earned its keep a
                          third time.
Methodology lesson      : 3 random-direction draws are too few for a
                          specificity bound; methodology.md will require
                          ≥10 draws with a percentile bound for any
                          specificity control from now on.
Next experiment         : program closed here. Proceeding tracks: arith_valid
                          band protocol; candidate_02 (quantization drift of
                          the Level-2 band map); expG (band replication on
                          Qwen3-1.7B). The final gate record for this arc:
                          refuse → pass(L2) → refuse → refuse.