SPB Git

spb/modelmap Public License

Internal cartography of local LLMs on Apple Silicon — registered, gated, negative-first. Public atlas at modelmap.io.

Python 66.3% JavaScript 24.5% CSS 8.1% Shell 0.7%

# project: modelmap document: expC_causal_verification — hypothesis (run #1, registered before run) author: Simon-Pierre Boucher contact: contact@spboucher.ai website: https://modelmap.io created: 2026-08-12 modified: 2026-08-12 status: reviewed

# Hypothesis — expC run #1 (does the agreement probe map survive ablation?)

Causal verification pipeline — correlational→causal survival rate. Registered 2026-08-12 before the run. Input: the differential (real−twin) agreement probe map from expA run #2 (atlas/qwen3-0.6b-4bit/probes/v2). Intervention: layer-skip ablation (block contribution removed, residual passes through) via the Tap layer.

text
Hypothesis              : The top-5 layers of the agreement DIFFERENTIAL map
                          causally support agreement behavior: skipping them
                          damages the model's grammatical preference
                          (mean logit margin of the correct verb form over
                          the incorrect one, on held-out minimal pairs) more
                          than skipping 5 random non-top layers — top-5
                          damage ≥ the 95th percentile of the random-5
                          damage distribution (20 draws), and ≥ 2× its mean.
Falsification criterion : Top-5 damage inside the random-5 distribution
                          (< 95th percentile) — the correlational map fails
                          causal verification; this survival datum (0 or 1)
                          is the first entry of expC's published
                          correlational→causal survival rate, either way.
Method                  : Behavioral metric: margin = mean over minimal
                          pairs of [logit(correct is/are) − logit(incorrect)]
                          at the verb position, teacher-forced prefix
                          "The <noun(s)> <location>". Held-out pairs: fresh
                          seed, deduped against every v2 probe text.
                          Conditions: baseline (no skip); skip top-5
                          differential layers; skip 5 random layers
                          (excluding top-5), 20 draws; skip bottom-5
                          differential layers (second control). Same model
                          (Qwen3-0.6B-4bit), deterministic forwards.
Baseline / null         : baseline margin (sanity: must be > 0, i.e. the
                          model actually prefers grammatical agreement —
                          else the whole question is moot at this size);
                          random-5 and bottom-5 skip distributions.
Result                  : (pending)
Interpretation          : (pending — with explicit confidence level)
Next experiment         : (pending)

Declared caveats. Layer-skip is a coarse intervention (removes ALL of a block's computation, not the agreement direction specifically) — a confirmed result licenses "these layers causally support agreement behavior", NOT "agreement is localized to these layers" (hydra/backup effects, notes §4.4, can mask redundancy). Direction-level interventions (LEACE-style erasure in the forward pass) are the registered follow-up.


# Hypothesis — expC run #2 (specificity control + direction-level surgery)

Registered 2026-08-12 before the run, after run #1's non-confirmation. Two declared confounds get instruments: general-damage normalization for layer-skip, and a surgical direction-level intervention the probe map can legitimately pass or fail.

text
Hypothesis (P1)         : run #1's bottom-5 damage is GENERAL, not
                          agreement-specific: normalizing margin damage by
                          general damage (mean NLL increase on neutral
                          prose), the bottom-5 specificity ratio falls at or
                          below the random-5 mean ratio, and top-5 does not
                          exceed the random p95 either (the run-1 verdict
                          stands, now with the confound measured).
Hypothesis (P2)         : direction-level erasure tracks the probe map:
                          erasing the layer-ℓ agreement direction
                          (difference-in-means, estimated on agreement_A
                          mean-pooled reps) at all positions of ℓ's output,
                          minus the damage from erasing a random direction
                          at the same layer, yields a per-layer specific-
                          damage profile that correlates with the
                          differential probe profile: Spearman ρ ≥ 0.4
                          (permutation p < 0.05, 10k perms).
Falsification criterion : (P2) ρ ≤ 0 → the localization claim also dies at
                          direction level: survival ledger 0/2 and the
                          agreement map's layer structure is declared
                          behaviorally void at this granularity.
Method                  : held-out bank enlarged with 6 NEW locations
                          (target ≥ 180 pairs, deduped as before); NLL on
                          20 neutral prose sentences per condition;
                          direction erasure h' = h − ⟨h−μ, u⟩u applied to
                          every position of layer ℓ's output, u = unit
                          class-mean difference at ℓ, μ = grand mean;
                          random-direction control: 3 seeds per layer,
                          same procedure; all 28 layers scanned.
Baseline / null         : per-layer random-direction erasure damage;
                          run #1 skip conditions rerun on the enlarged bank.
Result                  : (pending)
Interpretation          : (pending)
Next experiment         : (pending)

# Hypothesis — expC run #3 (replication of the causal direction profile)

Registered 2026-08-12 before the run. Run #2 found that erasing the diff-of-means agreement direction at any layer 2–15 destroys the behavior. By our own doctrine that profile is unpublishable until it replicates across direction estimates.

text
Hypothesis              : The per-layer specific-damage profile is an
                          estimator-stable object: profiles from directions
                          estimated on (i) set A, (ii) set B (disjoint
                          nouns), (iii) 3 bootstrap resamples of set A,
                          agree pairwise at Spearman ρ ≥ 0.7 (mean over the
                          10 pairs), AND the qualitative claim holds in
                          EVERY source: mean specific damage over layers
                          2–15 ≥ 3× mean over layers 20–27.
Falsification criterion : mean pairwise ρ < 0.7 or any source violating the
                          3× band claim → the "causal direction profile" is
                          estimator noise; not publishable; expC pivots to
                          steering-based verification.
Method                  : shared random-direction controls (3 dirs/layer,
                          computed once); agreement-direction scan repeated
                          per source (5 sources × 28 layers × 192 pairs);
                          same held-out bank and margins as run #2.
Baseline / null         : shared random-direction damage per layer.
Publication rule        : if confirmed → atlas qwen3-0.6b-4bit/
                          interventions/v1 at Level 2 (probing evidences the
                          direction, erasure confirms causal load,
                          replicated across estimates — NOT Level 3, because
                          the two agreeing methods share the diff-of-means
                          estimator; an independent intervention family
                          (activation addition) is the registered L3 path).
Result                  : (pending)
Interpretation          : (pending)
Next experiment         : (pending)

# Hypothesis — expC run #4 (the narrowed BAND claim)

Registered 2026-08-12 before the run, after run #3's gate refusal showed the early band replicates (+2.98…+3.34) while the late band is estimator noise. Everything fresh: direction estimates from DISJOINT HALVES of agreement_A and agreement_B (never used whole-set estimates), two new bootstrap seeds, and a NEW held-out behavioral bank (8 locations unseen by any prior run or promptset).

text
Hypothesis              : EARLY-BAND claim only — mean specific damage over
                          layers 2–15 is ≥ 2.5 (on the fresh bank's
                          baseline-margin scale) in EVERY one of six fresh
                          direction sources (A-half-1, A-half-2, B-half-1,
                          B-half-2, bootA-fresh, bootB-fresh). The late
                          band (20–27) is REPORTED but carries no claim —
                          declared unstable by run #3.
Falsification criterion : any source with early-band mean < 2.5 → the band
                          claim fails replication too; the direction-erasure
                          program is closed for agreement at 0.6B and expC
                          pivots to steering-based verification.
Method                  : full 28-layer scan per source (map format
                          unchanged); shared random-direction controls
                          (3/layer) on the fresh bank; margins as before.
Publication rule        : pass → atlas interventions/v1 published as a BAND
                          map at Level 2: map.json carries the six profiles,
                          their mean/min, the random-direction damage, and
                          the band statistics; the chart plots mean, worst
                          source, and the raw random-direction reference.
                          Fail → refusal documented, program closed.
Result                  : (pending)
Interpretation          : (pending)
Next experiment         : (pending)

# Hypothesis — expC run #5 (Level-3 path: dose-response steering)

Registered 2026-08-12 before the run. Erasure (runs #2–#4) established necessity of the agreement direction in the early band. Steering (activation addition) is the INDEPENDENT intervention family required for Level 3: if the direction is the causal handle, pushing along it must move the behavior in the predicted direction, dose-dependently, and random directions must not.

text
Hypothesis              : At each of three early-band test layers (4, 8, 12),
                          adding δ = α·σℓ·u to every position of the layer
                          output produces a STRICTLY MONOTONE margin in
                          α ∈ {−2, −1, 0, +1, +2} (σℓ = std of activation
                          projections on u at layer ℓ), with effect size:
                          margin(α=−2) ≤ 0.5 × baseline at every test layer.
                          Specificity: the same ±2σ doses along 3 random
                          directions move the margin by < 25% of baseline
                          (mean absolute change), at every test layer.
Falsification criterion : monotonicity broken at any test layer, OR the −2σ
                          dose fails to halve the margin anywhere, OR random
                          directions move the margin ≥ 25% — steering fails,
                          the entry stays Level 2, and the diff-of-means
                          direction is declared necessary-but-not-a-handle.
Method                  : direction u and σℓ estimated from agreement_A
                          (full set — the published v1 object); behavioral
                          bank = run #4's fresh bank (never used for
                          estimation); doses applied at one layer at a time.
Publication rule        : pass → atlas interventions/v2 at LEVEL 3 (erasure
                          necessity + steering dose-response = two
                          independent intervention families + probing;
                          v1 stays as the Level-2 record). Fail → documented,
                          v1 unchanged.
Result                  : (pending)
Interpretation          : (pending)
Next experiment         : (pending)

# Hypothesis — expC run #6 (minimal L3 claim: layer 12 as a single-layer handle)

Registered 2026-08-12 before the run, after run #5's conjunction failed while L12 passed every gate. The claim is narrowed to the object that showed textbook behavior — scope discipline again, now applied to layers.

text
Hypothesis              : At LAYER 12 ONLY, the diff-of-means agreement
                          direction is a dose-controlled causal handle, and
                          this is estimator-stable: for EVERY one of four
                          fresh direction sources (disjoint halves of
                          agreement_A and agreement_B, new permutation
                          seed), on a THIRD fresh behavioral bank (8
                          never-used locations):
                          (a) margins strictly monotone in
                              α ∈ {−2,−1,0,+1,+2}·σ;
                          (b) margin(−2σ) ≤ 0.5 × baseline;
                          (c) shared specificity control at L12: 3 random
                              directions at ±2σ move the margin by < 25%
                              of baseline on average.
Falsification criterion : any source failing (a) or (b), or (c) failing —
                          the single-layer L3 claim dies and Level 3 is
                          abandoned for this object (steering program
                          closed; candidate_02/expG proceed from the
                          Level-2 band entry).
Method                  : same margin metric; direction + σ re-estimated
                          per source at L12 from half-set mean-pooled reps;
                          doses applied at L12 output, all positions.
Publication rule        : pass → interventions/v2 at LEVEL 3 with scope
                          declared as {band necessity (v1) + single-layer
                          L12 handle}; fail → documented refusal, v1
                          unchanged.
Result                  : (pending)
Interpretation          : (pending)
Next experiment         : (pending)