project: modelmap document: expC_causal_verification — hypothesis (run #1, registered before run) author: Simon-Pierre Boucher contact: contact@spboucher.ai website: https://modelmap.io created: 2026-08-12 modified: 2026-08-12 status: reviewed
Hypothesis — expC run #1 (does the agreement probe map survive ablation?)
Causal verification pipeline — correlational→causal survival rate. Registered 2026-08-12 before the run. Input: the differential (real−twin) agreement probe map from expA run #2 (atlas/qwen3-0.6b-4bit/probes/v2). Intervention: layer-skip ablation (block contribution removed, residual passes through) via the Tap layer.
Hypothesis : The top-5 layers of the agreement DIFFERENTIAL map
causally support agreement behavior: skipping them
damages the model's grammatical preference
(mean logit margin of the correct verb form over
the incorrect one, on held-out minimal pairs) more
than skipping 5 random non-top layers — top-5
damage ≥ the 95th percentile of the random-5
damage distribution (20 draws), and ≥ 2× its mean.
Falsification criterion : Top-5 damage inside the random-5 distribution
(< 95th percentile) — the correlational map fails
causal verification; this survival datum (0 or 1)
is the first entry of expC's published
correlational→causal survival rate, either way.
Method : Behavioral metric: margin = mean over minimal
pairs of [logit(correct is/are) − logit(incorrect)]
at the verb position, teacher-forced prefix
"The <noun(s)> <location>". Held-out pairs: fresh
seed, deduped against every v2 probe text.
Conditions: baseline (no skip); skip top-5
differential layers; skip 5 random layers
(excluding top-5), 20 draws; skip bottom-5
differential layers (second control). Same model
(Qwen3-0.6B-4bit), deterministic forwards.
Baseline / null : baseline margin (sanity: must be > 0, i.e. the
model actually prefers grammatical agreement —
else the whole question is moot at this size);
random-5 and bottom-5 skip distributions.
Result : (pending)
Interpretation : (pending — with explicit confidence level)
Next experiment : (pending)Declared caveats. Layer-skip is a coarse intervention (removes ALL of a block's computation, not the agreement direction specifically) — a confirmed result licenses "these layers causally support agreement behavior", NOT "agreement is localized to these layers" (hydra/backup effects, notes §4.4, can mask redundancy). Direction-level interventions (LEACE-style erasure in the forward pass) are the registered follow-up.
Hypothesis — expC run #2 (specificity control + direction-level surgery)
Registered 2026-08-12 before the run, after run #1's non-confirmation. Two declared confounds get instruments: general-damage normalization for layer-skip, and a surgical direction-level intervention the probe map can legitimately pass or fail.
Hypothesis (P1) : run #1's bottom-5 damage is GENERAL, not
agreement-specific: normalizing margin damage by
general damage (mean NLL increase on neutral
prose), the bottom-5 specificity ratio falls at or
below the random-5 mean ratio, and top-5 does not
exceed the random p95 either (the run-1 verdict
stands, now with the confound measured).
Hypothesis (P2) : direction-level erasure tracks the probe map:
erasing the layer-ℓ agreement direction
(difference-in-means, estimated on agreement_A
mean-pooled reps) at all positions of ℓ's output,
minus the damage from erasing a random direction
at the same layer, yields a per-layer specific-
damage profile that correlates with the
differential probe profile: Spearman ρ ≥ 0.4
(permutation p < 0.05, 10k perms).
Falsification criterion : (P2) ρ ≤ 0 → the localization claim also dies at
direction level: survival ledger 0/2 and the
agreement map's layer structure is declared
behaviorally void at this granularity.
Method : held-out bank enlarged with 6 NEW locations
(target ≥ 180 pairs, deduped as before); NLL on
20 neutral prose sentences per condition;
direction erasure h' = h − ⟨h−μ, u⟩u applied to
every position of layer ℓ's output, u = unit
class-mean difference at ℓ, μ = grand mean;
random-direction control: 3 seeds per layer,
same procedure; all 28 layers scanned.
Baseline / null : per-layer random-direction erasure damage;
run #1 skip conditions rerun on the enlarged bank.
Result : (pending)
Interpretation : (pending)
Next experiment : (pending)Hypothesis — expC run #3 (replication of the causal direction profile)
Registered 2026-08-12 before the run. Run #2 found that erasing the diff-of-means agreement direction at any layer 2–15 destroys the behavior. By our own doctrine that profile is unpublishable until it replicates across direction estimates.
Hypothesis : The per-layer specific-damage profile is an
estimator-stable object: profiles from directions
estimated on (i) set A, (ii) set B (disjoint
nouns), (iii) 3 bootstrap resamples of set A,
agree pairwise at Spearman ρ ≥ 0.7 (mean over the
10 pairs), AND the qualitative claim holds in
EVERY source: mean specific damage over layers
2–15 ≥ 3× mean over layers 20–27.
Falsification criterion : mean pairwise ρ < 0.7 or any source violating the
3× band claim → the "causal direction profile" is
estimator noise; not publishable; expC pivots to
steering-based verification.
Method : shared random-direction controls (3 dirs/layer,
computed once); agreement-direction scan repeated
per source (5 sources × 28 layers × 192 pairs);
same held-out bank and margins as run #2.
Baseline / null : shared random-direction damage per layer.
Publication rule : if confirmed → atlas qwen3-0.6b-4bit/
interventions/v1 at Level 2 (probing evidences the
direction, erasure confirms causal load,
replicated across estimates — NOT Level 3, because
the two agreeing methods share the diff-of-means
estimator; an independent intervention family
(activation addition) is the registered L3 path).
Result : (pending)
Interpretation : (pending)
Next experiment : (pending)Hypothesis — expC run #4 (the narrowed BAND claim)
Registered 2026-08-12 before the run, after run #3's gate refusal showed the early band replicates (+2.98…+3.34) while the late band is estimator noise. Everything fresh: direction estimates from DISJOINT HALVES of agreement_A and agreement_B (never used whole-set estimates), two new bootstrap seeds, and a NEW held-out behavioral bank (8 locations unseen by any prior run or promptset).
Hypothesis : EARLY-BAND claim only — mean specific damage over
layers 2–15 is ≥ 2.5 (on the fresh bank's
baseline-margin scale) in EVERY one of six fresh
direction sources (A-half-1, A-half-2, B-half-1,
B-half-2, bootA-fresh, bootB-fresh). The late
band (20–27) is REPORTED but carries no claim —
declared unstable by run #3.
Falsification criterion : any source with early-band mean < 2.5 → the band
claim fails replication too; the direction-erasure
program is closed for agreement at 0.6B and expC
pivots to steering-based verification.
Method : full 28-layer scan per source (map format
unchanged); shared random-direction controls
(3/layer) on the fresh bank; margins as before.
Publication rule : pass → atlas interventions/v1 published as a BAND
map at Level 2: map.json carries the six profiles,
their mean/min, the random-direction damage, and
the band statistics; the chart plots mean, worst
source, and the raw random-direction reference.
Fail → refusal documented, program closed.
Result : (pending)
Interpretation : (pending)
Next experiment : (pending)Hypothesis — expC run #5 (Level-3 path: dose-response steering)
Registered 2026-08-12 before the run. Erasure (runs #2–#4) established necessity of the agreement direction in the early band. Steering (activation addition) is the INDEPENDENT intervention family required for Level 3: if the direction is the causal handle, pushing along it must move the behavior in the predicted direction, dose-dependently, and random directions must not.
Hypothesis : At each of three early-band test layers (4, 8, 12),
adding δ = α·σℓ·u to every position of the layer
output produces a STRICTLY MONOTONE margin in
α ∈ {−2, −1, 0, +1, +2} (σℓ = std of activation
projections on u at layer ℓ), with effect size:
margin(α=−2) ≤ 0.5 × baseline at every test layer.
Specificity: the same ±2σ doses along 3 random
directions move the margin by < 25% of baseline
(mean absolute change), at every test layer.
Falsification criterion : monotonicity broken at any test layer, OR the −2σ
dose fails to halve the margin anywhere, OR random
directions move the margin ≥ 25% — steering fails,
the entry stays Level 2, and the diff-of-means
direction is declared necessary-but-not-a-handle.
Method : direction u and σℓ estimated from agreement_A
(full set — the published v1 object); behavioral
bank = run #4's fresh bank (never used for
estimation); doses applied at one layer at a time.
Publication rule : pass → atlas interventions/v2 at LEVEL 3 (erasure
necessity + steering dose-response = two
independent intervention families + probing;
v1 stays as the Level-2 record). Fail → documented,
v1 unchanged.
Result : (pending)
Interpretation : (pending)
Next experiment : (pending)Hypothesis — expC run #6 (minimal L3 claim: layer 12 as a single-layer handle)
Registered 2026-08-12 before the run, after run #5's conjunction failed while L12 passed every gate. The claim is narrowed to the object that showed textbook behavior — scope discipline again, now applied to layers.
Hypothesis : At LAYER 12 ONLY, the diff-of-means agreement
direction is a dose-controlled causal handle, and
this is estimator-stable: for EVERY one of four
fresh direction sources (disjoint halves of
agreement_A and agreement_B, new permutation
seed), on a THIRD fresh behavioral bank (8
never-used locations):
(a) margins strictly monotone in
α ∈ {−2,−1,0,+1,+2}·σ;
(b) margin(−2σ) ≤ 0.5 × baseline;
(c) shared specificity control at L12: 3 random
directions at ±2σ move the margin by < 25%
of baseline on average.
Falsification criterion : any source failing (a) or (b), or (c) failing —
the single-layer L3 claim dies and Level 3 is
abandoned for this object (steering program
closed; candidate_02/expG proceed from the
Level-2 band entry).
Method : same margin metric; direction + σ re-estimated
per source at L12 from half-set mean-pooled reps;
doses applied at L12 output, all positions.
Publication rule : pass → interventions/v2 at LEVEL 3 with scope
declared as {band necessity (v1) + single-layer
L12 handle}; fail → documented refusal, v1
unchanged.
Result : (pending)
Interpretation : (pending)
Next experiment : (pending)