--- project: modelmap document: expC_causal_verification — hypothesis (run #1, registered before run) author: Simon-Pierre Boucher contact: contact@spboucher.ai website: https://modelmap.io created: 2026-08-12 modified: 2026-08-12 status: reviewed --- # Hypothesis — expC run #1 (does the agreement probe map survive ablation?) > Causal verification pipeline — correlational→causal survival rate. > Registered 2026-08-12 **before** the run. Input: the differential > (real−twin) agreement probe map from expA run #2 > (atlas/qwen3-0.6b-4bit/probes/v2). Intervention: layer-skip ablation > (block contribution removed, residual passes through) via the Tap layer. ```text Hypothesis : The top-5 layers of the agreement DIFFERENTIAL map causally support agreement behavior: skipping them damages the model's grammatical preference (mean logit margin of the correct verb form over the incorrect one, on held-out minimal pairs) more than skipping 5 random non-top layers — top-5 damage ≥ the 95th percentile of the random-5 damage distribution (20 draws), and ≥ 2× its mean. Falsification criterion : Top-5 damage inside the random-5 distribution (< 95th percentile) — the correlational map fails causal verification; this survival datum (0 or 1) is the first entry of expC's published correlational→causal survival rate, either way. Method : Behavioral metric: margin = mean over minimal pairs of [logit(correct is/are) − logit(incorrect)] at the verb position, teacher-forced prefix "The ". Held-out pairs: fresh seed, deduped against every v2 probe text. Conditions: baseline (no skip); skip top-5 differential layers; skip 5 random layers (excluding top-5), 20 draws; skip bottom-5 differential layers (second control). Same model (Qwen3-0.6B-4bit), deterministic forwards. Baseline / null : baseline margin (sanity: must be > 0, i.e. the model actually prefers grammatical agreement — else the whole question is moot at this size); random-5 and bottom-5 skip distributions. Result : (pending) Interpretation : (pending — with explicit confidence level) Next experiment : (pending) ``` **Declared caveats.** Layer-skip is a coarse intervention (removes ALL of a block's computation, not the agreement direction specifically) — a confirmed result licenses "these layers causally support agreement behavior", NOT "agreement is localized to these layers" (hydra/backup effects, notes §4.4, can mask redundancy). Direction-level interventions (LEACE-style erasure in the forward pass) are the registered follow-up. --- # Hypothesis — expC run #2 (specificity control + direction-level surgery) Registered 2026-08-12 **before** the run, after run #1's non-confirmation. Two declared confounds get instruments: general-damage normalization for layer-skip, and a surgical direction-level intervention the probe map can legitimately pass or fail. ```text Hypothesis (P1) : run #1's bottom-5 damage is GENERAL, not agreement-specific: normalizing margin damage by general damage (mean NLL increase on neutral prose), the bottom-5 specificity ratio falls at or below the random-5 mean ratio, and top-5 does not exceed the random p95 either (the run-1 verdict stands, now with the confound measured). Hypothesis (P2) : direction-level erasure tracks the probe map: erasing the layer-ℓ agreement direction (difference-in-means, estimated on agreement_A mean-pooled reps) at all positions of ℓ's output, minus the damage from erasing a random direction at the same layer, yields a per-layer specific- damage profile that correlates with the differential probe profile: Spearman ρ ≥ 0.4 (permutation p < 0.05, 10k perms). Falsification criterion : (P2) ρ ≤ 0 → the localization claim also dies at direction level: survival ledger 0/2 and the agreement map's layer structure is declared behaviorally void at this granularity. Method : held-out bank enlarged with 6 NEW locations (target ≥ 180 pairs, deduped as before); NLL on 20 neutral prose sentences per condition; direction erasure h' = h − ⟨h−μ, u⟩u applied to every position of layer ℓ's output, u = unit class-mean difference at ℓ, μ = grand mean; random-direction control: 3 seeds per layer, same procedure; all 28 layers scanned. Baseline / null : per-layer random-direction erasure damage; run #1 skip conditions rerun on the enlarged bank. Result : (pending) Interpretation : (pending) Next experiment : (pending) ``` --- # Hypothesis — expC run #3 (replication of the causal direction profile) Registered 2026-08-12 **before** the run. Run #2 found that erasing the diff-of-means agreement direction at any layer 2–15 destroys the behavior. By our own doctrine that profile is unpublishable until it replicates across direction estimates. ```text Hypothesis : The per-layer specific-damage profile is an estimator-stable object: profiles from directions estimated on (i) set A, (ii) set B (disjoint nouns), (iii) 3 bootstrap resamples of set A, agree pairwise at Spearman ρ ≥ 0.7 (mean over the 10 pairs), AND the qualitative claim holds in EVERY source: mean specific damage over layers 2–15 ≥ 3× mean over layers 20–27. Falsification criterion : mean pairwise ρ < 0.7 or any source violating the 3× band claim → the "causal direction profile" is estimator noise; not publishable; expC pivots to steering-based verification. Method : shared random-direction controls (3 dirs/layer, computed once); agreement-direction scan repeated per source (5 sources × 28 layers × 192 pairs); same held-out bank and margins as run #2. Baseline / null : shared random-direction damage per layer. Publication rule : if confirmed → atlas qwen3-0.6b-4bit/ interventions/v1 at Level 2 (probing evidences the direction, erasure confirms causal load, replicated across estimates — NOT Level 3, because the two agreeing methods share the diff-of-means estimator; an independent intervention family (activation addition) is the registered L3 path). Result : (pending) Interpretation : (pending) Next experiment : (pending) ``` --- # Hypothesis — expC run #4 (the narrowed BAND claim) Registered 2026-08-12 **before** the run, after run #3's gate refusal showed the early band replicates (+2.98…+3.34) while the late band is estimator noise. Everything fresh: direction estimates from DISJOINT HALVES of agreement_A and agreement_B (never used whole-set estimates), two new bootstrap seeds, and a NEW held-out behavioral bank (8 locations unseen by any prior run or promptset). ```text Hypothesis : EARLY-BAND claim only — mean specific damage over layers 2–15 is ≥ 2.5 (on the fresh bank's baseline-margin scale) in EVERY one of six fresh direction sources (A-half-1, A-half-2, B-half-1, B-half-2, bootA-fresh, bootB-fresh). The late band (20–27) is REPORTED but carries no claim — declared unstable by run #3. Falsification criterion : any source with early-band mean < 2.5 → the band claim fails replication too; the direction-erasure program is closed for agreement at 0.6B and expC pivots to steering-based verification. Method : full 28-layer scan per source (map format unchanged); shared random-direction controls (3/layer) on the fresh bank; margins as before. Publication rule : pass → atlas interventions/v1 published as a BAND map at Level 2: map.json carries the six profiles, their mean/min, the random-direction damage, and the band statistics; the chart plots mean, worst source, and the raw random-direction reference. Fail → refusal documented, program closed. Result : (pending) Interpretation : (pending) Next experiment : (pending) ``` --- # Hypothesis — expC run #5 (Level-3 path: dose-response steering) Registered 2026-08-12 **before** the run. Erasure (runs #2–#4) established necessity of the agreement direction in the early band. Steering (activation addition) is the INDEPENDENT intervention family required for Level 3: if the direction is the causal handle, pushing along it must move the behavior in the predicted direction, dose-dependently, and random directions must not. ```text Hypothesis : At each of three early-band test layers (4, 8, 12), adding δ = α·σℓ·u to every position of the layer output produces a STRICTLY MONOTONE margin in α ∈ {−2, −1, 0, +1, +2} (σℓ = std of activation projections on u at layer ℓ), with effect size: margin(α=−2) ≤ 0.5 × baseline at every test layer. Specificity: the same ±2σ doses along 3 random directions move the margin by < 25% of baseline (mean absolute change), at every test layer. Falsification criterion : monotonicity broken at any test layer, OR the −2σ dose fails to halve the margin anywhere, OR random directions move the margin ≥ 25% — steering fails, the entry stays Level 2, and the diff-of-means direction is declared necessary-but-not-a-handle. Method : direction u and σℓ estimated from agreement_A (full set — the published v1 object); behavioral bank = run #4's fresh bank (never used for estimation); doses applied at one layer at a time. Publication rule : pass → atlas interventions/v2 at LEVEL 3 (erasure necessity + steering dose-response = two independent intervention families + probing; v1 stays as the Level-2 record). Fail → documented, v1 unchanged. Result : (pending) Interpretation : (pending) Next experiment : (pending) ``` --- # Hypothesis — expC run #6 (minimal L3 claim: layer 12 as a single-layer handle) Registered 2026-08-12 **before** the run, after run #5's conjunction failed while L12 passed every gate. The claim is narrowed to the object that showed textbook behavior — scope discipline again, now applied to layers. ```text Hypothesis : At LAYER 12 ONLY, the diff-of-means agreement direction is a dose-controlled causal handle, and this is estimator-stable: for EVERY one of four fresh direction sources (disjoint halves of agreement_A and agreement_B, new permutation seed), on a THIRD fresh behavioral bank (8 never-used locations): (a) margins strictly monotone in α ∈ {−2,−1,0,+1,+2}·σ; (b) margin(−2σ) ≤ 0.5 × baseline; (c) shared specificity control at L12: 3 random directions at ±2σ move the margin by < 25% of baseline on average. Falsification criterion : any source failing (a) or (b), or (c) failing — the single-layer L3 claim dies and Level 3 is abandoned for this object (steering program closed; candidate_02/expG proceed from the Level-2 band entry). Method : same margin metric; direction + σ re-estimated per source at L12 from half-set mean-pooled reps; doses applied at L12 output, all positions. Publication rule : pass → interventions/v2 at LEVEL 3 with scope declared as {band necessity (v1) + single-layer L12 handle}; fail → documented refusal, v1 unchanged. Result : (pending) Interpretation : (pending) Next experiment : (pending) ```