--- project: modelmap document: expC_causal_verification — analysis (run #1) author: Simon-Pierre Boucher contact: contact@spboucher.ai website: https://modelmap.io created: 2026-08-12 modified: 2026-08-12 status: reviewed --- # Analysis — expC run #1: the agreement probe map FAILS causal verification Run: `results/expC_causal_verification/20260812T064534Z/results.json` (5.7 s). Model: Qwen3-0.6B-4bit. Input map: agreement differential (probes/v2). Intervention: layer-skip ablation via the Tap layer. Hypothesis and the binary verdict criterion registered before the run. ```text Hypothesis : the top-5 differential layers (by real−twin probe selectivity) causally support agreement behavior — skip damage ≥ p95 of random-5 draws AND ≥ 2× their mean. Falsification criterion : top-5 damage inside the random distribution. Result : FALSIFIED — survival 0/1. Baseline grammatical margin +4.63 (sanity holds: the model robustly prefers correct agreement). Skip damage: top-5 (layers 17,18,19,21,22) = +2.17 — BELOW the random-5 mean (+3.14, p95 +4.85, 20 draws); bottom-5 differential (layers 0–4) = +4.99, the LARGEST of all conditions. Interpretation : Level 1 for the negative claim. Where agreement information is most linearly decodable above the architecture null (late-mid layers) is NOT where the computation is causally load-bearing for the behavior. This is the Hase-class dissociation (localization ≠ causal support — notes §4.6) measured end-to-end in our own pipeline, on a pre-registered binary verdict. The correlational→causal survival ledger opens at 0/1. Declared caveats bite exactly as registered: (a) layer-skip is coarse — early-layer skips plausibly cause GENERAL degradation, not agreement-specific damage (bottom-5 +4.99 reads as "the model breaks", not "agreement lives at layers 0–4"); a specificity control (margin damage normalized by general perplexity damage) is the registered follow-up; (b) held-out pairs = 89 after dedup against every v2 probe text (below the planned 200 — combo pool exhausted; enlarge banks next run). Next experiment : run #2 with (1) perplexity-normalized specificity scores per skip condition, (2) direction-level intervention (LEACE erasure of the agreement direction in the forward pass at layer ℓ) — a surgical test the probe map CAN legitimately pass or fail, (3) larger held-out bank. Then the same protocol on arith_valid. ``` ## Ledger | # | Correlational claim | Intervention | Survives? | |---|---|---|---| | 1 | agreement top-5 differential layers (probes/v2) | layer-skip ablation | **No** (0/1) | | 2 | agreement differential layer *profile* (probes/v2) | direction-level erasure scan | **No** (0/2) | The published survival rate is the running fraction of this table — the charter §8.3 metric, now live. --- # Analysis — expC run #2: ledger 0/2 — and the control scan finds the real structure Run: `results/expC_causal_verification/20260812T065406Z/results.json` (192 held-out pairs, baseline margin +4.24, NLL 5.73). Hypotheses registered before the run. ```text Hypothesis (P1) : run #1's bottom-5 damage is general, not agreement-specific, once normalized by NLL damage. Result (P1) : CONFIRMED. Specificity (margin damage / NLL damage): bottom-5 = 0.51 — BELOW the random-5 mean (1.44); top-5 = 2.51 — above the random mean but below the p95 (4.73). Run #1's verdict stands, now with the confound measured: early-layer skips break the model generally; nothing in the skip family singles out the probe map's top layers. Hypothesis (P2) : per-layer direction-erasure specific damage correlates with the differential probe profile (Spearman ρ ≥ 0.4, perm-p < 0.05). Result (P2) : FALSIFIED, decisively — ρ = −0.136, perm-p 0.757. Survival ledger: 0/2. The probe map's layer ranking is behaviorally void at this granularity. BUT the scan itself uncovered strong causal structure the probe map missed: erasing the layer-ℓ agreement direction (diff-of-means, with random-direction controls netted out) at ANY single layer in 2–15 destroys most of the margin (specific damage +2.6…+4.0 of a +4.24 baseline; L12: +3.97, L15: +3.75, L6: +3.71), while the late layers the probes ranked highest carry little (L21: +0.69, L20: +0.29) — and L18/L22 erasure slightly HELPS (−1.86/−1.09), suggesting suppressive components. Interpretation : Level 0–1 (single model, single direction estimate from set A, no seed replication of the intervention yet — by our own doctrine this causal profile is NOT publishable as an atlas entry until replicated). Two lessons: (1) correlational layer rankings did not survive two different causal tests — the survival rate the field never publishes is, so far, 0%; (2) the behaviorally load-bearing object is a low-dimensional DIRECTION present across early-mid layers, not a "place" — consistent with the linear-representation view and with why decodability peaks (late, after information is everywhere) diverge from causal joints (early, where the direction is constructed). Random- direction erasure also hurts at layers 2–9 (+0.9…+2.5): early residual streams are fragile to ANY rank-1 deletion — netted out in the specific column. Next experiment : run #3 — replicate the direction-erasure profile (direction re-estimated on set B and on 5 bootstrap seeds; report profile replication rate) → if stable, publish as the atlas's first INTERVENTIONS map (Level toward 2–3) and make the probes/v2 entry point to it as the causal counterpart. Then the same scan on arith_valid. ``` --- # Analysis — expC run #3: the publication gate REFUSED the causal profile Run: `results/expC_causal_verification/20260812T065929Z/results.json`. Five direction sources (set A, set B, 3 bootstraps of A), shared random-direction controls, same held-out bank (192 pairs, baseline +4.24). Hypothesis, criterion, and publication rule registered before the run. ```text Hypothesis : the per-layer specific-damage profile is estimator-stable — mean pairwise Spearman ρ ≥ 0.7 and early(2–15) ≥ 3× late(20–27) in EVERY source. Result : FALSIFIED — mean ρ = 0.495 (min 0.176); the band claim fails for setB and bootA2 (ratio 1.2). make_interventions_mapcard.py refused the entry (exit 1), per the registered rule. The atlas stays at two entries — the gate did its job. THE STABLE PART: early-band (layers 2–15) mean specific damage replicates tightly across all five sources (+2.98, +3.22, +3.34, +3.27, +3.28) — erasing the estimated agreement direction anywhere in the early-mid band reliably destroys ~70–79% of the behavior, whatever the estimation set. THE UNSTABLE PART: the late band (20–27) swings from −0.44 to +2.77 with the estimator — the late-layer "profile" is direction-estimation noise, which also retroactively explains run #2's anti-correlation (the probe ranking lives exactly where the causal profile is noise). Interpretation : Level 0–1. The full profile is NOT a stable object; the coarser claim ("an early-band agreement direction is causally load-bearing") is the replicable candidate. Publishing gates that refuse are the mechanism that keeps the atlas honest — this refusal is itself a process result worth reporting on the site's methodology page. Next experiment : run #4 with the NARROWER pre-registered claim: early-band (2–15) mean specific damage ≥ 2.5 on fresh direction estimates (new bootstrap seeds + a held-out estimation split) and a fresh behavioral bank; late band explicitly excluded as unstable. If it passes, publish the interventions map as an early-band BAND claim (not a per-layer ranking), Level 2. ``` --- # Analysis — expC run #4: the BAND claim passes — first Level-2 atlas entry Run: `results/expC_causal_verification/20260812T071954Z/results.json`. Everything fresh: six direction sources (disjoint halves of A and B + two new bootstraps), new behavioral bank (8 unseen locations, 192 pairs, baseline margin +4.45). Claim, bar, and publication rule registered before the run. ```text Hypothesis : early-band (2–15) mean specific damage ≥ 2.5 in EVERY fresh source; late band reported, no claim. Result : CONFIRMED — early-band means: Ahalf1 +3.35, Ahalf2 +3.23, bootA +3.30, Bhalf1 +3.26, Bhalf2 +3.33, bootB +3.27 (min +3.228 vs bar 2.5). Spread across six estimators: 0.12 — the band is a tight, estimator-stable causal object. Late band again unstable (+0.27…+2.83), as declared; it carries no claim. Interpretation : Level 2, published: atlas/qwen3-0.6b-4bit/ interventions/v1 — the project's first Level-2 entry and first causal map. The claim: a single linear direction (diff-of-means over correct vs violated agreement), erased at ANY one layer in 2–15, removes ~73–75% of the model's grammatical preference on held-out pairs, replicated across six independent direction estimates and netted against random-direction damage. NOT Level 3: both agreeing methods share the diff-of-means estimator — activation-addition steering is the registered L3 path. Granularity discipline paid off: the per-layer version of this map was refused (run #3); the band version replicates. Next experiment : (1) L3 path: steering (add the direction) should INCREASE the margin on violated-preference pairs; (2) same band protocol on arith_valid; (3) quantization drift of the band map (candidate_02: does the early band move under Q8/FP16?); (4) cross-model: does the band replicate on Qwen3-1.7B (expG entry)? ``` --- # Analysis — expC run #5: Level-3 gate refused — and layer 12 is a textbook handle Run: `results/expC_causal_verification/20260812T072455Z/results.json`. Steering doses α·σℓ·u at layers {4, 8, 12}, α ∈ {−2…+2}; conjunctive criterion (monotone AND halved AND specific at EVERY test layer) registered before the run. make_l3_mapcard.py refused (exit 1); v1 stays Level 2. ```text Hypothesis : strict dose-response monotonicity + halving + random-direction specificity at all of L4/L8/L12. Result : FALSIFIED as a conjunction. L12 — PASSES EVERYTHING, textbook: margins +1.95 / +3.00 / +4.45 / +5.82 / +6.79 across −2σ…+2σ (strictly monotone), halved at −2σ, random-direction |Δ| = 0.45 vs bound 1.11. L08 — near-monotone (+2σ dips: 5.70 < 6.26), random |Δ| 2.97: not specific at this dose. L04 — OVERDOSE REGIME: ±2σ both collapse the margin (+0.08 / +0.04) and random directions are equally destructive (|Δ| 4.40): at early layers, ANY perturbation of magnitude ~2σ (σ estimated from mean-pooled reps) breaks the computation — the rank-1 fragility of run #2, now dose-resolved. Interpretation : Level 0–1. The direction is a clean, dose- controlled causal handle at mid-band (L12) but the conjunctive claim over the whole test set fails, so the gate held v2 back — correctly. Reading: "necessity everywhere in the band" (erasure, L2 entry) coexists with "controllable handle only where the layer tolerates perturbation". Dose scale is a confound at early layers: σ from mean-pooled statistics likely overdoses positions at layers with different norm profiles. Next experiment : run #6 (registered idea): layer-local dose calibration (σ from position-level projections at the target layer; or a dose-sweep to find each layer's non-destructive range), then re-register the handle claim on the sub-band that tolerates calibrated doses (candidate: 10–14). Also register the L12 single-layer handle claim with fresh direction estimates as a minimal L3 candidate. ``` ## Gate record (for methodology.md) Three refusals/passes to date, all pre-registered: run #3 per-layer profile REFUSED → run #4 band claim PASSED (Level 2 published) → run #5 L3 conjunction REFUSED (v1 unchanged). The atlas never received a claim its evidence didn't carry. --- # Analysis — expC run #6: the L12 handle claim fails replication — Level 3 abandoned Run: `results/expC_causal_verification/20260812T072913Z/results.json`. Four fresh direction sources at L12, third fresh bank (192 pairs, baseline +4.57), fresh random-direction seeds. Registered rule: any source failing monotonicity/halving, or specificity failing, kills the L3 claim and closes the steering program for this object. ```text Hypothesis : L12 is an estimator-stable dose-controlled handle. Result : FALSIFIED. Monotone: 1/4 sources only (Ahalf2); the POSITIVE dose arm is unstable (Ahalf1 +2σ dips 5.59<6.29; Bhalf1 +1σ dips below baseline; Bhalf2 +2σ < +1σ). Halving at −2σ: 4/4 — the negative (erasure-like) arm is robust, again. SPECIFICITY FAILED TO REPLICATE: random-direction mean |Δ| = 3.36 vs bound 1.14 on the fresh bank and fresh random seeds (run #5's L12 value was 0.45 — with only 3 random draws, that pass now reads as sampling luck). Interpretation : Level 1 for the negative. The direction is NECESSARY (v1, Level 2, band-replicated) but NOT a reliable additive handle: pushing along it does not control the behavior in a dose-stable, direction-specific way. Per the pre-registered rule, Level 3 is abandoned for this object and the steering program is closed. Had run #5's L12 observation been published without fresh re-registration, the atlas would now contain a false Level-3 claim — the gate earned its keep a third time. Methodology lesson : 3 random-direction draws are too few for a specificity bound; methodology.md will require ≥10 draws with a percentile bound for any specificity control from now on. Next experiment : program closed here. Proceeding tracks: arith_valid band protocol; candidate_02 (quantization drift of the Level-2 band map); expG (band replication on Qwen3-1.7B). The final gate record for this arc: refuse → pass(L2) → refuse → refuse. ```