expC run #1: agreement map FAILS causal verification — survival ledger 0/1
Layer-skip ablation (Tap.skip): top-5 differential layers damage +2.17, BELOW random-5 mean +3.14 (p95 +4.85); bottom-5 (layers 0-4) largest at +4.99. Decodability != causal support (Hase-class dissociation, in-house). Atlas probes/v2 records the failed check (interventions + confidence.md); survival-rate ledger opened in expC analysis. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Showing 8 changed files with +416 and −14
modified
atlas/qwen3-0.6b-4bit/probes/v2/confidence.md
+2 −1
@@ -15,7 +15,8 @@ Level : 1 | ||
| 15 | 15 | Seeds : 5 |
| 16 | 16 | Prompt sets: 6 (token-balanced, structure-borne; overlap certificates in manifest) |
| 17 | 17 | Methods in agreement : 1 (linear probes only — Level 2 requires a second method) |
| 18 | −Causal verification : none (observational; Level 3 requires intervention) | |
| 18 | +Causal verification : ATTEMPTED AND FAILED (expC run #1 layer-skip: top-5 | |
| 19 | + differential layers not confirmed; survival 0/1) | |
| 19 | 20 | ``` |
| 20 | 21 | |
| 21 | 22 | Per-property verdicts (differential real−twin, mean pooling): |
modified
atlas/qwen3-0.6b-4bit/probes/v2/mapcard.json
+4 −2
@@ -59,14 +59,16 @@ | ||
| 59 | 59 | "class token-overlap certificates in promptset manifest" |
| 60 | 60 | ], |
| 61 | 61 | "methods_in_agreement": [], |
| 62 | − "interventions": [], | |
| 62 | + "interventions": [ | |
| 63 | + "layer-skip ablation (expC run #1, 2026-08-12): top-5 differential layers NOT confirmed — damage below random-5 mean; see experiments/micro/expC_causal_verification/analysis.md" | |
| 64 | + ], | |
| 63 | 65 | "replication_rate": 0.5366, |
| 64 | 66 | "per_dataset_agreement": null, |
| 65 | 67 | "ablation_schemes": [], |
| 66 | 68 | "featurizer_class": "natural-basis (mean-pooled + last-token residual)", |
| 67 | 69 | "intervention_protocol": "none (observational map — Level 1 by design)", |
| 68 | 70 | "negative_result": false, |
| 69 | − "notes": "DIFFERENTIAL map (real minus random-init twin), per the doctrine adopted after v1. Mixed outcome by property: agreement and arith_valid carry trained-model signal above the architecture prior; word_order is null-dominated and flagged as such. Strict twin gate (<0.05) still fails on word_order/agreement — only differential claims are published.", | |
| 71 | + "notes": "DIFFERENTIAL map (real minus random-init twin), per the doctrine adopted after v1. Mixed outcome by property: agreement and arith_valid carry trained-model signal above the architecture prior; word_order is null-dominated and flagged as such. Strict twin gate (<0.05) still fails on word_order/agreement — only differential claims are published. CAUSAL CHECK: expC run #1 layer-skip ablation did NOT confirm the top differential layers (survival 0/1); map remains Level 1 and its layer ranking must not be read as causal.", | |
| 70 | 72 | "author": "Simon-Pierre Boucher", |
| 71 | 73 | "contact": "contact@spboucher.ai", |
| 72 | 74 | "website": "https://modelmap.io", |
added
experiments/micro/expC_causal_verification/analysis.md
+68 −0
@@ -0,0 +1,68 @@ | ||
| 1 | +--- | |
| 2 | +project: modelmap | |
| 3 | +document: expC_causal_verification — analysis (run #1) | |
| 4 | +author: Simon-Pierre Boucher | |
| 5 | +contact: contact@spboucher.ai | |
| 6 | +website: https://modelmap.io | |
| 7 | +created: 2026-08-12 | |
| 8 | +modified: 2026-08-12 | |
| 9 | +status: reviewed | |
| 10 | +--- | |
| 11 | + | |
| 12 | +# Analysis — expC run #1: the agreement probe map FAILS causal verification | |
| 13 | + | |
| 14 | +Run: `results/expC_causal_verification/20260812T064534Z/results.json` (5.7 s). | |
| 15 | +Model: Qwen3-0.6B-4bit. Input map: agreement differential (probes/v2). | |
| 16 | +Intervention: layer-skip ablation via the Tap layer. Hypothesis and the | |
| 17 | +binary verdict criterion registered before the run. | |
| 18 | + | |
| 19 | +```text | |
| 20 | +Hypothesis : the top-5 differential layers (by real−twin probe | |
| 21 | + selectivity) causally support agreement behavior — | |
| 22 | + skip damage ≥ p95 of random-5 draws AND ≥ 2× their | |
| 23 | + mean. | |
| 24 | +Falsification criterion : top-5 damage inside the random distribution. | |
| 25 | +Result : FALSIFIED — survival 0/1. Baseline grammatical | |
| 26 | + margin +4.63 (sanity holds: the model robustly | |
| 27 | + prefers correct agreement). Skip damage: | |
| 28 | + top-5 (layers 17,18,19,21,22) = +2.17 — | |
| 29 | + BELOW the random-5 mean (+3.14, p95 +4.85, 20 | |
| 30 | + draws); bottom-5 differential (layers 0–4) = +4.99, | |
| 31 | + the LARGEST of all conditions. | |
| 32 | +Interpretation : Level 1 for the negative claim. Where agreement | |
| 33 | + information is most linearly decodable above the | |
| 34 | + architecture null (late-mid layers) is NOT where | |
| 35 | + the computation is causally load-bearing for the | |
| 36 | + behavior. This is the Hase-class dissociation | |
| 37 | + (localization ≠ causal support — notes §4.6) | |
| 38 | + measured end-to-end in our own pipeline, on a | |
| 39 | + pre-registered binary verdict. The | |
| 40 | + correlational→causal survival ledger opens at 0/1. | |
| 41 | + Declared caveats bite exactly as registered: | |
| 42 | + (a) layer-skip is coarse — early-layer skips | |
| 43 | + plausibly cause GENERAL degradation, not | |
| 44 | + agreement-specific damage (bottom-5 +4.99 reads as | |
| 45 | + "the model breaks", not "agreement lives at layers | |
| 46 | + 0–4"); a specificity control (margin damage | |
| 47 | + normalized by general perplexity damage) is the | |
| 48 | + registered follow-up; (b) held-out pairs = 89 | |
| 49 | + after dedup against every v2 probe text (below the | |
| 50 | + planned 200 — combo pool exhausted; enlarge banks | |
| 51 | + next run). | |
| 52 | +Next experiment : run #2 with (1) perplexity-normalized specificity | |
| 53 | + scores per skip condition, (2) direction-level | |
| 54 | + intervention (LEACE erasure of the agreement | |
| 55 | + direction in the forward pass at layer ℓ) — a | |
| 56 | + surgical test the probe map CAN legitimately pass | |
| 57 | + or fail, (3) larger held-out bank. Then the same | |
| 58 | + protocol on arith_valid. | |
| 59 | +``` | |
| 60 | + | |
| 61 | +## Ledger | |
| 62 | + | |
| 63 | +| # | Correlational claim | Intervention | Survives? | | |
| 64 | +|---|---|---|---| | |
| 65 | +| 1 | agreement top-5 differential layers (probes/v2) | layer-skip ablation | **No** (0/1) | | |
| 66 | + | |
| 67 | +The published survival rate is the running fraction of this table — the | |
| 68 | +charter §8.3 metric, now live. | |
modified
experiments/micro/expC_causal_verification/hypothesis.md
+43 −8
@@ -1,23 +1,58 @@ | ||
| 1 | 1 | --- |
| 2 | 2 | project: modelmap |
| 3 | −document: expC_causal_verification — hypothesis | |
| 3 | +document: expC_causal_verification — hypothesis (run #1, registered before run) | |
| 4 | 4 | author: Simon-Pierre Boucher |
| 5 | 5 | contact: contact@spboucher.ai |
| 6 | 6 | website: https://modelmap.io |
| 7 | 7 | created: 2026-08-12 |
| 8 | −status: draft | |
| 8 | +modified: 2026-08-12 | |
| 9 | +status: reviewed | |
| 9 | 10 | --- |
| 10 | 11 | |
| 11 | −# Hypothesis — expC_causal_verification | |
| 12 | +# Hypothesis — expC run #1 (does the agreement probe map survive ablation?) | |
| 12 | 13 | |
| 13 | −> Causal verification pipeline — correlational-to-causal survival rate | |
| 14 | +> Causal verification pipeline — correlational→causal survival rate. | |
| 15 | +> Registered 2026-08-12 **before** the run. Input: the differential | |
| 16 | +> (real−twin) agreement probe map from expA run #2 | |
| 17 | +> (atlas/qwen3-0.6b-4bit/probes/v2). Intervention: layer-skip ablation | |
| 18 | +> (block contribution removed, residual passes through) via the Tap layer. | |
| 14 | 19 | |
| 15 | 20 | ```text |
| 16 | −Hypothesis : (to be registered before the first run) | |
| 17 | −Falsification criterion : (explicit kill-number, registered in advance) | |
| 18 | −Method : (including controls) | |
| 19 | −Baseline / null : (shuffled labels / random directions / random init) | |
| 21 | +Hypothesis : The top-5 layers of the agreement DIFFERENTIAL map | |
| 22 | + causally support agreement behavior: skipping them | |
| 23 | + damages the model's grammatical preference | |
| 24 | + (mean logit margin of the correct verb form over | |
| 25 | + the incorrect one, on held-out minimal pairs) more | |
| 26 | + than skipping 5 random non-top layers — top-5 | |
| 27 | + damage ≥ the 95th percentile of the random-5 | |
| 28 | + damage distribution (20 draws), and ≥ 2× its mean. | |
| 29 | +Falsification criterion : Top-5 damage inside the random-5 distribution | |
| 30 | + (< 95th percentile) — the correlational map fails | |
| 31 | + causal verification; this survival datum (0 or 1) | |
| 32 | + is the first entry of expC's published | |
| 33 | + correlational→causal survival rate, either way. | |
| 34 | +Method : Behavioral metric: margin = mean over minimal | |
| 35 | + pairs of [logit(correct is/are) − logit(incorrect)] | |
| 36 | + at the verb position, teacher-forced prefix | |
| 37 | + "The <noun(s)> <location>". Held-out pairs: fresh | |
| 38 | + seed, deduped against every v2 probe text. | |
| 39 | + Conditions: baseline (no skip); skip top-5 | |
| 40 | + differential layers; skip 5 random layers | |
| 41 | + (excluding top-5), 20 draws; skip bottom-5 | |
| 42 | + differential layers (second control). Same model | |
| 43 | + (Qwen3-0.6B-4bit), deterministic forwards. | |
| 44 | +Baseline / null : baseline margin (sanity: must be > 0, i.e. the | |
| 45 | + model actually prefers grammatical agreement — | |
| 46 | + else the whole question is moot at this size); | |
| 47 | + random-5 and bottom-5 skip distributions. | |
| 20 | 48 | Result : (pending) |
| 21 | 49 | Interpretation : (pending — with explicit confidence level) |
| 22 | 50 | Next experiment : (pending) |
| 23 | 51 | ``` |
| 52 | + | |
| 53 | +**Declared caveats.** Layer-skip is a coarse intervention (removes ALL of a | |
| 54 | +block's computation, not the agreement direction specifically) — a confirmed | |
| 55 | +result licenses "these layers causally support agreement behavior", NOT | |
| 56 | +"agreement is localized to these layers" (hydra/backup effects, notes §4.4, | |
| 57 | +can mask redundancy). Direction-level interventions (LEACE-style erasure in | |
| 58 | +the forward pass) are the registered follow-up. | |
added
experiments/micro/expC_causal_verification/implementation/benchmark.py
+165 −0
@@ -0,0 +1,165 @@ | ||
| 1 | +#!/usr/bin/env python3 | |
| 2 | +# ============================================================================= | |
| 3 | +# Project : modelmap | |
| 4 | +# File : experiments/micro/expC_causal_verification/implementation/benchmark.py | |
| 5 | +# Purpose : Run #1 — layer-skip ablation test of the agreement probe map | |
| 6 | +# Author : Simon-Pierre Boucher | |
| 7 | +# Contact : contact@spboucher.ai | |
| 8 | +# Website : https://modelmap.io | |
| 9 | +# Created : 2026-08-12 | |
| 10 | +# Modified : 2026-08-12 | |
| 11 | +# Platform : macOS / Apple Silicon (arm64) — MLX / Metal | |
| 12 | +# License : All rights reserved (research code) | |
| 13 | +# ============================================================================= | |
| 14 | +"""expC run #1 (hypothesis registered before this run). | |
| 15 | + | |
| 16 | +Does the agreement DIFFERENTIAL probe map (expA run #2) survive causal | |
| 17 | +testing? Skip the map's top-5 layers vs 20 random-5 draws vs bottom-5, | |
| 18 | +measure the drop in grammatical-agreement logit margin on held-out | |
| 19 | +minimal pairs. | |
| 20 | +""" | |
| 21 | + | |
| 22 | +from __future__ import annotations | |
| 23 | + | |
| 24 | +import json | |
| 25 | +import random | |
| 26 | +import subprocess | |
| 27 | +import sys | |
| 28 | +import time | |
| 29 | +from pathlib import Path | |
| 30 | + | |
| 31 | +import numpy as np | |
| 32 | + | |
| 33 | +ROOT = Path(__file__).resolve().parents[4] | |
| 34 | +sys.path.insert(0, str(ROOT / "src")) | |
| 35 | +sys.path.insert(0, str(ROOT / "benchmarks")) | |
| 36 | +from hardware_manifest import manifest | |
| 37 | + | |
| 38 | +from modelmap.capture.mlx_capture import install_taps | |
| 39 | + | |
| 40 | +MODEL = "mlx-community/Qwen3-0.6B-4bit" | |
| 41 | +MAP_JSON = ROOT / "atlas" / "qwen3-0.6b-4bit" / "probes" / "v2" / "map.json" | |
| 42 | +K = 5 | |
| 43 | +N_RANDOM_DRAWS = 20 | |
| 44 | +N_PAIRS = 200 | |
| 45 | +SEED = 777 | |
| 46 | + | |
| 47 | +NOUN_PAIRS = [("key", "keys"), ("crate", "crates"), ("report", "reports"), ("valve", "valves"), | |
| 48 | + ("ticket", "tickets"), ("ladder", "ladders"), ("sample", "samples"), ("cable", "cables"), | |
| 49 | + ("permit", "permits"), ("beacon", "beacons"), ("filter", "filters"), ("stamp", "stamps")] | |
| 50 | +NEAR = ["near the entrance", "beside the counter", "under the shelf", "behind the gate", | |
| 51 | + "next to the archive", "along the corridor", "opposite the office", "inside the vault", | |
| 52 | + "across the yard", "above the workbench"] | |
| 53 | + | |
| 54 | + | |
| 55 | +def held_out_pairs(rng: random.Random, forbidden: set[str]): | |
| 56 | + """Minimal pairs (prefix, singular?) deduped against every v2 probe text.""" | |
| 57 | + pairs, seen = [], set() | |
| 58 | + combos = [(n, loc) for n in NOUN_PAIRS for loc in NEAR] | |
| 59 | + rng.shuffle(combos) | |
| 60 | + for (sg, pl), loc in combos * 4: | |
| 61 | + for noun, singular in ((sg, True), (pl, False)): | |
| 62 | + prefix = f"The {noun} {loc}" | |
| 63 | + probe_like = f"{prefix} is" if singular else f"{prefix} are" | |
| 64 | + if prefix in seen or any(probe_like in f for f in forbidden): | |
| 65 | + continue | |
| 66 | + seen.add(prefix) | |
| 67 | + pairs.append({"prefix": prefix, "singular": singular}) | |
| 68 | + if len(pairs) >= N_PAIRS: | |
| 69 | + return pairs | |
| 70 | + return pairs | |
| 71 | + | |
| 72 | + | |
| 73 | +def main() -> int: | |
| 74 | + import mlx.core as mx | |
| 75 | + from mlx_lm import load | |
| 76 | + | |
| 77 | + t0 = time.time() | |
| 78 | + model, tokenizer = load(MODEL) | |
| 79 | + taps = install_taps(model) | |
| 80 | + n_layers = len(taps) | |
| 81 | + | |
| 82 | + # top/bottom differential layers from the published map (mean pooling) | |
| 83 | + mdoc = json.loads(MAP_JSON.read_text()) | |
| 84 | + agree = mdoc["properties"]["agreement"] | |
| 85 | + diff = [(a["layer"], a["selectivity_mean"] - t["selectivity_mean"]) | |
| 86 | + for a, t in zip(agree["per_layer"]["A"], agree["twin_null_per_layer_A"])] | |
| 87 | + ranked = [l for l, _ in sorted(diff, key=lambda x: -x[1])] | |
| 88 | + top_k, bottom_k = ranked[:K], ranked[-K:] | |
| 89 | + print(f"top-{K} differential layers: {sorted(top_k)} | bottom-{K}: {sorted(bottom_k)}") | |
| 90 | + | |
| 91 | + # held-out minimal pairs | |
| 92 | + forbidden = set() | |
| 93 | + for f in (ROOT / "benchmarks" / "promptsets").glob("agreement_*.jsonl"): | |
| 94 | + forbidden.update(json.loads(l)["text"] for l in f.read_text().splitlines()) | |
| 95 | + rng = random.Random(SEED) | |
| 96 | + pairs = held_out_pairs(rng, forbidden) | |
| 97 | + print(f"held-out minimal pairs: {len(pairs)}") | |
| 98 | + | |
| 99 | + tok_is = tokenizer.encode(" is") | |
| 100 | + tok_are = tokenizer.encode(" are") | |
| 101 | + assert len(tok_is) == 1 and len(tok_are) == 1, "verb forms must be single tokens" | |
| 102 | + id_is, id_are = tok_is[0], tok_are[0] | |
| 103 | + prefix_ids = [tokenizer.encode(p["prefix"]) for p in pairs] | |
| 104 | + | |
| 105 | + def margin(skip_layers: set[int]) -> float: | |
| 106 | + for i, t in enumerate(taps): | |
| 107 | + t.skip = i in skip_layers | |
| 108 | + margins = [] | |
| 109 | + for ids, p in zip(prefix_ids, pairs): | |
| 110 | + logits = model(mx.array([ids]))[0, -1, :] | |
| 111 | + mx.eval(logits) | |
| 112 | + m = float(logits[id_is] - logits[id_are]) | |
| 113 | + margins.append(m if p["singular"] else -m) | |
| 114 | + for t in taps: | |
| 115 | + t.skip = False | |
| 116 | + return float(np.mean(margins)) | |
| 117 | + | |
| 118 | + baseline = margin(set()) | |
| 119 | + print(f"baseline margin: {baseline:+.4f}") | |
| 120 | + top_m = margin(set(top_k)) | |
| 121 | + bottom_m = margin(set(bottom_k)) | |
| 122 | + candidates = [i for i in range(n_layers) if i not in set(top_k)] | |
| 123 | + random_ms = [] | |
| 124 | + for d in range(N_RANDOM_DRAWS): | |
| 125 | + draw = set(random.Random(SEED + 1 + d).sample(candidates, K)) | |
| 126 | + random_ms.append(margin(draw)) | |
| 127 | + dmg = lambda m: baseline - m | |
| 128 | + rd = np.array([dmg(m) for m in random_ms]) | |
| 129 | + p95 = float(np.percentile(rd, 95)) | |
| 130 | + verdict = bool(dmg(top_m) >= p95 and dmg(top_m) >= 2 * rd.mean()) | |
| 131 | + print(f"damage: top-{K}={dmg(top_m):+.4f} bottom-{K}={dmg(bottom_m):+.4f} " | |
| 132 | + f"random mean={rd.mean():+.4f} p95={p95:+.4f} -> survives={verdict}") | |
| 133 | + | |
| 134 | + commit = subprocess.run(["git", "rev-parse", "HEAD"], cwd=ROOT, | |
| 135 | + capture_output=True, text=True, check=False).stdout.strip() | |
| 136 | + ts = time.strftime("%Y%m%dT%H%M%SZ", time.gmtime()) | |
| 137 | + outdir = ROOT / "results" / "expC_causal_verification" / ts | |
| 138 | + outdir.mkdir(parents=True) | |
| 139 | + doc = { | |
| 140 | + "experiment": "expC_causal_verification", "run": 1, | |
| 141 | + "scope": "layer-skip ablation of the agreement differential probe map", | |
| 142 | + "commit": commit, | |
| 143 | + "config": {"model": MODEL, "k": K, "n_random_draws": N_RANDOM_DRAWS, | |
| 144 | + "n_pairs": len(pairs), "seed": SEED, | |
| 145 | + "map_source": str(MAP_JSON.relative_to(ROOT)), | |
| 146 | + "top_layers": sorted(top_k), "bottom_layers": sorted(bottom_k)}, | |
| 147 | + "manifest": manifest(), | |
| 148 | + "results": { | |
| 149 | + "baseline_margin": baseline, | |
| 150 | + "top_k_margin": top_m, "top_k_damage": dmg(top_m), | |
| 151 | + "bottom_k_margin": bottom_m, "bottom_k_damage": dmg(bottom_m), | |
| 152 | + "random_margins": random_ms, | |
| 153 | + "random_damage_mean": float(rd.mean()), | |
| 154 | + "random_damage_p95": p95, | |
| 155 | + "survives_causal_test": verdict, | |
| 156 | + }, | |
| 157 | + "wall_seconds": round(time.time() - t0, 1), | |
| 158 | + } | |
| 159 | + (outdir / "results.json").write_text(json.dumps(doc, indent=2) + "\n") | |
| 160 | + print(f"results -> {outdir / 'results.json'} ({doc['wall_seconds']} s)") | |
| 161 | + return 0 | |
| 162 | + | |
| 163 | + | |
| 164 | +if __name__ == "__main__": | |
| 165 | + sys.exit(main()) | |
modified
research/LOG.md
+35 −0
@@ -356,3 +356,38 @@ arith_valid: LEACE erasure as second method, then ablation on top layers | ||
| 356 | 356 | (expC entry). (3) The ~0.54 replication number seeds expD's design (more |
| 357 | 357 | seeds, tighter CIs). (4) The agreement differential map is candidate_02's |
| 358 | 358 | quantization-drift target. |
| 359 | + | |
| 360 | +--- | |
| 361 | + | |
| 362 | +## 2026-08-12 07:45 EDT — expC run #1: the agreement map FAILS causal verification (survival ledger opens 0/1) | |
| 363 | + | |
| 364 | +**Question.** Do the top-5 layers of the agreement differential probe map | |
| 365 | +causally support agreement behavior under layer-skip ablation? | |
| 366 | + | |
| 367 | +**Result (5.7 s; pre-registered binary verdict).** FALSIFIED. | |
| 368 | +Baseline grammatical margin +4.63 (the 0.6B model robustly prefers correct | |
| 369 | +agreement — sanity holds). Skip damage: top-5 differential layers | |
| 370 | +(17,18,19,21,22) = +2.17, BELOW the random-5 mean (+3.14; p95 +4.85, 20 | |
| 371 | +draws); bottom-5 differential (layers 0–4) = +4.99, the largest of all | |
| 372 | +conditions. **Where agreement information is most decodable above the | |
| 373 | +architecture null is not where the computation is causally load-bearing.** | |
| 374 | +The Hase-class dissociation (localization ≠ causal support), measured | |
| 375 | +end-to-end in our own pipeline within one day of standing it up. | |
| 376 | + | |
| 377 | +**Ledger.** The correlational→causal survival rate — the charter §8.3 | |
| 378 | +metric — is now live: 0/1. The atlas entry probes/v2 records the failed | |
| 379 | +check in mapcard.interventions and confidence.md ("causal verification: | |
| 380 | +attempted and failed"); the map stays Level 1 and its layer ranking is | |
| 381 | +explicitly flagged as non-causal. | |
| 382 | + | |
| 383 | +**Caveats (registered in advance, both bit).** Layer-skip is coarse: | |
| 384 | +bottom-5 damage plausibly reflects GENERAL degradation (early layers break | |
| 385 | +everything), not agreement-specific structure. Held-out pairs 89 < planned | |
| 386 | +200 (dedup exhausted the combo pool). | |
| 387 | + | |
| 388 | +**Decisions / next.** expC run #2: (1) perplexity-normalized specificity | |
| 389 | +per skip condition; (2) direction-level intervention — LEACE-erase the | |
| 390 | +agreement direction in the forward pass at layer ℓ (a surgical test the | |
| 391 | +probe map can legitimately pass); (3) larger held-out bank; then the same | |
| 392 | +protocol on arith_valid. The pipeline now demonstrably runs the full | |
| 393 | +charter loop: register → measure → verify causally → publish either way. | |
added
results/expC_causal_verification/20260812T064534Z/results.json
+88 −0
@@ -0,0 +1,88 @@ | ||
| 1 | +{ | |
| 2 | + "experiment": "expC_causal_verification", | |
| 3 | + "run": 1, | |
| 4 | + "scope": "layer-skip ablation of the agreement differential probe map", | |
| 5 | + "commit": "9ed234bf4f8e787a8d6fbd01107c25567a403115", | |
| 6 | + "config": { | |
| 7 | + "model": "mlx-community/Qwen3-0.6B-4bit", | |
| 8 | + "k": 5, | |
| 9 | + "n_random_draws": 20, | |
| 10 | + "n_pairs": 89, | |
| 11 | + "seed": 777, | |
| 12 | + "map_source": "atlas/qwen3-0.6b-4bit/probes/v2/map.json", | |
| 13 | + "top_layers": [ | |
| 14 | + 17, | |
| 15 | + 18, | |
| 16 | + 19, | |
| 17 | + 21, | |
| 18 | + 22 | |
| 19 | + ], | |
| 20 | + "bottom_layers": [ | |
| 21 | + 0, | |
| 22 | + 1, | |
| 23 | + 2, | |
| 24 | + 3, | |
| 25 | + 4 | |
| 26 | + ] | |
| 27 | + }, | |
| 28 | + "manifest": { | |
| 29 | + "author": "Simon-Pierre Boucher", | |
| 30 | + "contact": "contact@spboucher.ai", | |
| 31 | + "website": "https://modelmap.io", | |
| 32 | + "chip": { | |
| 33 | + "brand": "Apple M5 Max", | |
| 34 | + "cores_total": 18, | |
| 35 | + "cores_performance": 6, | |
| 36 | + "cores_efficiency": 12 | |
| 37 | + }, | |
| 38 | + "memory": { | |
| 39 | + "unified_gb": 48.0, | |
| 40 | + "pagesize": 16384 | |
| 41 | + }, | |
| 42 | + "os": { | |
| 43 | + "system": "Darwin", | |
| 44 | + "version": "27.0", | |
| 45 | + "arch": "arm64" | |
| 46 | + }, | |
| 47 | + "software": { | |
| 48 | + "python": "3.14.4", | |
| 49 | + "numpy": "2.5.2", | |
| 50 | + "mlx": "0.32.0", | |
| 51 | + "torch": "2.13.0", | |
| 52 | + "safetensors": "0.8.0" | |
| 53 | + } | |
| 54 | + }, | |
| 55 | + "results": { | |
| 56 | + "baseline_margin": 4.631320224719101, | |
| 57 | + "top_k_margin": 2.464185393258427, | |
| 58 | + "top_k_damage": 2.167134831460674, | |
| 59 | + "bottom_k_margin": -0.3570926966292135, | |
| 60 | + "bottom_k_damage": 4.988412921348314, | |
| 61 | + "random_margins": [ | |
| 62 | + 0.17854634831460675, | |
| 63 | + 3.4838483146067416, | |
| 64 | + 3.6688904494382024, | |
| 65 | + 2.7650983146067416, | |
| 66 | + 2.133426966292135, | |
| 67 | + 1.8202247191011236, | |
| 68 | + 0.6693293539325843, | |
| 69 | + 3.425561797752809, | |
| 70 | + 0.5043012640449438, | |
| 71 | + 2.018960674157303, | |
| 72 | + -0.20628511235955055, | |
| 73 | + 0.3945751404494382, | |
| 74 | + 0.6560086025280899, | |
| 75 | + -0.21166169241573032, | |
| 76 | + 4.489115168539326, | |
| 77 | + -0.34467169943820225, | |
| 78 | + 0.15458216292134833, | |
| 79 | + 2.4719101123595504, | |
| 80 | + 1.559691011235955, | |
| 81 | + 0.14381803019662923 | |
| 82 | + ], | |
| 83 | + "random_damage_mean": 3.142556728405899, | |
| 84 | + "random_damage_p95": 4.849632417485955, | |
| 85 | + "survives_causal_test": false | |
| 86 | + }, | |
| 87 | + "wall_seconds": 5.7 | |
| 88 | +} | |
modified
src/modelmap/capture/mlx_capture.py
+11 −3
@@ -26,15 +26,23 @@ import numpy as np | ||
| 26 | 26 | |
| 27 | 27 | |
| 28 | 28 | class Tap: |
| 29 | − """Wraps a decoder layer; optionally retains its output.""" | |
| 29 | + """Wraps a decoder layer; optionally retains its output or skips the block. | |
| 30 | + | |
| 31 | + skip=True implements layer ablation (the block's contribution is removed; | |
| 32 | + the residual stream passes through unchanged) — the Gromov/ShortGPT-style | |
| 33 | + intervention used by expC for causal verification. | |
| 34 | + """ | |
| 30 | 35 | |
| 31 | 36 | def __init__(self, layer): |
| 32 | 37 | self.layer = layer |
| 33 | 38 | self.retained = None |
| 34 | 39 | self.enabled = False |
| 40 | + self.skip = False | |
| 35 | 41 | |
| 36 | − def __call__(self, *args, **kwargs): | |
| 37 | − out = self.layer(*args, **kwargs) | |
| 42 | + def __call__(self, x, *args, **kwargs): | |
| 43 | + if self.skip: | |
| 44 | + return x | |
| 45 | + out = self.layer(x, *args, **kwargs) | |
| 38 | 46 | if self.enabled: |
| 39 | 47 | self.retained = out |
| 40 | 48 | return out |
| 41 | 49 | |