--- project: modelmap document: qwen3-0.6b-4bit/interventions/v1 — confidence author: Simon-Pierre Boucher contact: contact@spboucher.ai website: https://modelmap.io created: 2026-08-12 status: reviewed --- # Confidence — qwen3-0.6b-4bit / interventions / v1 ```text Level : 2 Seeds : six fresh direction sources (disjoint halves of two promptsets + two fresh bootstraps); fresh behavioral bank Prompt sets: 3 (two estimation sets + held-out behavioral bank) Methods in agreement : 2 (diff-of-means probing; direction erasure) — shared estimator, hence Level 2 and not 3 Causal verification : YES — rank-1 erasure with random-direction nulls ``` Pre-registered band claim (bar 2.5 on a +4.45 baseline margin): - Ahalf1: early-band mean +3.347 - Ahalf2: early-band mean +3.228 - bootA: early-band mean +3.302 - Bhalf1: early-band mean +3.263 - Bhalf2: early-band mean +3.326 - bootB: early-band mean +3.270 Minimum early-band mean across sources: +3.228 — claim PASSES. Granularity discipline: run #3's per-layer profile FAILED replication and was refused by this very gate; the published object is the BAND (layers 2–15). The late band (20–27) is displayed but carries no claim. This map is the causal counterpart of probes/v2, whose decodability ranking it contradicts (survival ledger 0/2) — both stay published, labeled by what they measure.