project: modelmap document: expC_causal_verification — analysis (run #1) author: Simon-Pierre Boucher contact: contact@spboucher.ai website: https://modelmap.io created: 2026-08-12 modified: 2026-08-12 status: reviewed
Analysis — expC run #1: the agreement probe map FAILS causal verification
Run: results/expC_causal_verification/20260812T064534Z/results.json (5.7 s).
Model: Qwen3-0.6B-4bit. Input map: agreement differential (probes/v2).
Intervention: layer-skip ablation via the Tap layer. Hypothesis and the
binary verdict criterion registered before the run.
Hypothesis : the top-5 differential layers (by real−twin probe
selectivity) causally support agreement behavior —
skip damage ≥ p95 of random-5 draws AND ≥ 2× their
mean.
Falsification criterion : top-5 damage inside the random distribution.
Result : FALSIFIED — survival 0/1. Baseline grammatical
margin +4.63 (sanity holds: the model robustly
prefers correct agreement). Skip damage:
top-5 (layers 17,18,19,21,22) = +2.17 —
BELOW the random-5 mean (+3.14, p95 +4.85, 20
draws); bottom-5 differential (layers 0–4) = +4.99,
the LARGEST of all conditions.
Interpretation : Level 1 for the negative claim. Where agreement
information is most linearly decodable above the
architecture null (late-mid layers) is NOT where
the computation is causally load-bearing for the
behavior. This is the Hase-class dissociation
(localization ≠ causal support — notes §4.6)
measured end-to-end in our own pipeline, on a
pre-registered binary verdict. The
correlational→causal survival ledger opens at 0/1.
Declared caveats bite exactly as registered:
(a) layer-skip is coarse — early-layer skips
plausibly cause GENERAL degradation, not
agreement-specific damage (bottom-5 +4.99 reads as
"the model breaks", not "agreement lives at layers
0–4"); a specificity control (margin damage
normalized by general perplexity damage) is the
registered follow-up; (b) held-out pairs = 89
after dedup against every v2 probe text (below the
planned 200 — combo pool exhausted; enlarge banks
next run).
Next experiment : run #2 with (1) perplexity-normalized specificity
scores per skip condition, (2) direction-level
intervention (LEACE erasure of the agreement
direction in the forward pass at layer ℓ) — a
surgical test the probe map CAN legitimately pass
or fail, (3) larger held-out bank. Then the same
protocol on arith_valid.Ledger
| # | Correlational claim | Intervention | Survives? |
|---|---|---|---|
| 1 | agreement top-5 differential layers (probes/v2) | layer-skip ablation | No (0/1) |
| 2 | agreement differential layer profile (probes/v2) | direction-level erasure scan | No (0/2) |
The published survival rate is the running fraction of this table — the charter §8.3 metric, now live.
Analysis — expC run #2: ledger 0/2 — and the control scan finds the real structure
Run: results/expC_causal_verification/20260812T065406Z/results.json (192
held-out pairs, baseline margin +4.24, NLL 5.73). Hypotheses registered
before the run.
Hypothesis (P1) : run #1's bottom-5 damage is general, not
agreement-specific, once normalized by NLL damage.
Result (P1) : CONFIRMED. Specificity (margin damage / NLL
damage): bottom-5 = 0.51 — BELOW the random-5 mean
(1.44); top-5 = 2.51 — above the random mean but
below the p95 (4.73). Run #1's verdict stands, now
with the confound measured: early-layer skips
break the model generally; nothing in the skip
family singles out the probe map's top layers.
Hypothesis (P2) : per-layer direction-erasure specific damage
correlates with the differential probe profile
(Spearman ρ ≥ 0.4, perm-p < 0.05).
Result (P2) : FALSIFIED, decisively — ρ = −0.136, perm-p 0.757.
Survival ledger: 0/2. The probe map's layer
ranking is behaviorally void at this granularity.
BUT the scan itself uncovered strong causal
structure the probe map missed: erasing the
layer-ℓ agreement direction (diff-of-means, with
random-direction controls netted out) at ANY
single layer in 2–15 destroys most of the margin
(specific damage +2.6…+4.0 of a +4.24 baseline;
L12: +3.97, L15: +3.75, L6: +3.71), while the
late layers the probes ranked highest carry
little (L21: +0.69, L20: +0.29) — and L18/L22
erasure slightly HELPS (−1.86/−1.09), suggesting
suppressive components.
Interpretation : Level 0–1 (single model, single direction
estimate from set A, no seed replication of the
intervention yet — by our own doctrine this
causal profile is NOT publishable as an atlas
entry until replicated). Two lessons:
(1) correlational layer rankings did not survive
two different causal tests — the survival rate
the field never publishes is, so far, 0%;
(2) the behaviorally load-bearing object is a
low-dimensional DIRECTION present across
early-mid layers, not a "place" — consistent with
the linear-representation view and with why
decodability peaks (late, after information is
everywhere) diverge from causal joints (early,
where the direction is constructed). Random-
direction erasure also hurts at layers 2–9
(+0.9…+2.5): early residual streams are fragile
to ANY rank-1 deletion — netted out in the
specific column.
Next experiment : run #3 — replicate the direction-erasure profile
(direction re-estimated on set B and on 5
bootstrap seeds; report profile replication rate)
→ if stable, publish as the atlas's first
INTERVENTIONS map (Level toward 2–3) and make the
probes/v2 entry point to it as the causal
counterpart. Then the same scan on arith_valid.Analysis — expC run #3: the publication gate REFUSED the causal profile
Run: results/expC_causal_verification/20260812T065929Z/results.json.
Five direction sources (set A, set B, 3 bootstraps of A), shared
random-direction controls, same held-out bank (192 pairs, baseline +4.24).
Hypothesis, criterion, and publication rule registered before the run.
Hypothesis : the per-layer specific-damage profile is
estimator-stable — mean pairwise Spearman ρ ≥ 0.7
and early(2–15) ≥ 3× late(20–27) in EVERY source.
Result : FALSIFIED — mean ρ = 0.495 (min 0.176); the band
claim fails for setB and bootA2 (ratio 1.2).
make_interventions_mapcard.py refused the entry
(exit 1), per the registered rule. The atlas
stays at two entries — the gate did its job.
THE STABLE PART: early-band (layers 2–15) mean
specific damage replicates tightly across all
five sources (+2.98, +3.22, +3.34, +3.27, +3.28)
— erasing the estimated agreement direction
anywhere in the early-mid band reliably destroys
~70–79% of the behavior, whatever the estimation
set. THE UNSTABLE PART: the late band (20–27)
swings from −0.44 to +2.77 with the estimator —
the late-layer "profile" is direction-estimation
noise, which also retroactively explains run #2's
anti-correlation (the probe ranking lives exactly
where the causal profile is noise).
Interpretation : Level 0–1. The full profile is NOT a stable
object; the coarser claim ("an early-band
agreement direction is causally load-bearing") is
the replicable candidate. Publishing gates that
refuse are the mechanism that keeps the atlas
honest — this refusal is itself a process result
worth reporting on the site's methodology page.
Next experiment : run #4 with the NARROWER pre-registered claim:
early-band (2–15) mean specific damage ≥ 2.5 on
fresh direction estimates (new bootstrap seeds +
a held-out estimation split) and a fresh
behavioral bank; late band explicitly excluded as
unstable. If it passes, publish the interventions
map as an early-band BAND claim (not a per-layer
ranking), Level 2.Analysis — expC run #4: the BAND claim passes — first Level-2 atlas entry
Run: results/expC_causal_verification/20260812T071954Z/results.json.
Everything fresh: six direction sources (disjoint halves of A and B + two
new bootstraps), new behavioral bank (8 unseen locations, 192 pairs,
baseline margin +4.45). Claim, bar, and publication rule registered before
the run.
Hypothesis : early-band (2–15) mean specific damage ≥ 2.5 in
EVERY fresh source; late band reported, no claim.
Result : CONFIRMED — early-band means: Ahalf1 +3.35,
Ahalf2 +3.23, bootA +3.30, Bhalf1 +3.26,
Bhalf2 +3.33, bootB +3.27 (min +3.228 vs bar
2.5). Spread across six estimators: 0.12 — the
band is a tight, estimator-stable causal object.
Late band again unstable (+0.27…+2.83), as
declared; it carries no claim.
Interpretation : Level 2, published: atlas/qwen3-0.6b-4bit/
interventions/v1 — the project's first Level-2
entry and first causal map. The claim: a single
linear direction (diff-of-means over correct vs
violated agreement), erased at ANY one layer in
2–15, removes ~73–75% of the model's grammatical
preference on held-out pairs, replicated across
six independent direction estimates and netted
against random-direction damage. NOT Level 3:
both agreeing methods share the diff-of-means
estimator — activation-addition steering is the
registered L3 path. Granularity discipline paid
off: the per-layer version of this map was
refused (run #3); the band version replicates.
Next experiment : (1) L3 path: steering (add the direction) should
INCREASE the margin on violated-preference pairs;
(2) same band protocol on arith_valid; (3)
quantization drift of the band map (candidate_02:
does the early band move under Q8/FP16?);
(4) cross-model: does the band replicate on
Qwen3-1.7B (expG entry)?Analysis — expC run #5: Level-3 gate refused — and layer 12 is a textbook handle
Run: results/expC_causal_verification/20260812T072455Z/results.json.
Steering doses α·σℓ·u at layers {4, 8, 12}, α ∈ {−2…+2}; conjunctive
criterion (monotone AND halved AND specific at EVERY test layer) registered
before the run. make_l3_mapcard.py refused (exit 1); v1 stays Level 2.
Hypothesis : strict dose-response monotonicity + halving +
random-direction specificity at all of L4/L8/L12.
Result : FALSIFIED as a conjunction.
L12 — PASSES EVERYTHING, textbook: margins
+1.95 / +3.00 / +4.45 / +5.82 / +6.79 across
−2σ…+2σ (strictly monotone), halved at −2σ,
random-direction |Δ| = 0.45 vs bound 1.11.
L08 — near-monotone (+2σ dips: 5.70 < 6.26),
random |Δ| 2.97: not specific at this dose.
L04 — OVERDOSE REGIME: ±2σ both collapse the
margin (+0.08 / +0.04) and random directions are
equally destructive (|Δ| 4.40): at early layers,
ANY perturbation of magnitude ~2σ (σ estimated
from mean-pooled reps) breaks the computation —
the rank-1 fragility of run #2, now dose-resolved.
Interpretation : Level 0–1. The direction is a clean, dose-
controlled causal handle at mid-band (L12) but
the conjunctive claim over the whole test set
fails, so the gate held v2 back — correctly.
Reading: "necessity everywhere in the band"
(erasure, L2 entry) coexists with "controllable
handle only where the layer tolerates
perturbation". Dose scale is a confound at early
layers: σ from mean-pooled statistics likely
overdoses positions at layers with different
norm profiles.
Next experiment : run #6 (registered idea): layer-local dose
calibration (σ from position-level projections at
the target layer; or a dose-sweep to find each
layer's non-destructive range), then re-register
the handle claim on the sub-band that tolerates
calibrated doses (candidate: 10–14). Also register
the L12 single-layer handle claim with fresh
direction estimates as a minimal L3 candidate.Gate record (for methodology.md)
Three refusals/passes to date, all pre-registered: run #3 per-layer profile REFUSED → run #4 band claim PASSED (Level 2 published) → run #5 L3 conjunction REFUSED (v1 unchanged). The atlas never received a claim its evidence didn't carry.
Analysis — expC run #6: the L12 handle claim fails replication — Level 3 abandoned
Run: results/expC_causal_verification/20260812T072913Z/results.json.
Four fresh direction sources at L12, third fresh bank (192 pairs, baseline
+4.57), fresh random-direction seeds. Registered rule: any source failing
monotonicity/halving, or specificity failing, kills the L3 claim and closes
the steering program for this object.
Hypothesis : L12 is an estimator-stable dose-controlled handle.
Result : FALSIFIED. Monotone: 1/4 sources only (Ahalf2);
the POSITIVE dose arm is unstable (Ahalf1 +2σ dips
5.59<6.29; Bhalf1 +1σ dips below baseline; Bhalf2
+2σ < +1σ). Halving at −2σ: 4/4 — the negative
(erasure-like) arm is robust, again. SPECIFICITY
FAILED TO REPLICATE: random-direction mean |Δ| =
3.36 vs bound 1.14 on the fresh bank and fresh
random seeds (run #5's L12 value was 0.45 — with
only 3 random draws, that pass now reads as
sampling luck).
Interpretation : Level 1 for the negative. The direction is
NECESSARY (v1, Level 2, band-replicated) but NOT
a reliable additive handle: pushing along it does
not control the behavior in a dose-stable,
direction-specific way. Per the pre-registered
rule, Level 3 is abandoned for this object and
the steering program is closed. Had run #5's L12
observation been published without fresh
re-registration, the atlas would now contain a
false Level-3 claim — the gate earned its keep a
third time.
Methodology lesson : 3 random-direction draws are too few for a
specificity bound; methodology.md will require
≥10 draws with a percentile bound for any
specificity control from now on.
Next experiment : program closed here. Proceeding tracks: arith_valid
band protocol; candidate_02 (quantization drift of
the Level-2 band map); expG (band replication on
Qwen3-1.7B). The final gate record for this arc:
refuse → pass(L2) → refuse → refuse.