SPB Git

spb/modelmap Public License

Internal cartography of local LLMs on Apple Silicon — registered, gated, negative-first. Public atlas at modelmap.io.

Python 66.3% JavaScript 24.5% CSS 8.1% Shell 0.7%

expC run #1: agreement map FAILS causal verification — survival ledger 0/1

Layer-skip ablation (Tap.skip): top-5 differential layers damage +2.17,
BELOW random-5 mean +3.14 (p95 +4.85); bottom-5 (layers 0-4) largest at
+4.99. Decodability != causal support (Hase-class dissociation, in-house).
Atlas probes/v2 records the failed check (interventions + confidence.md);
survival-rate ledger opened in expC analysis.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Simon-Pierre Boucher committed 3 h ago (Aug 12, 2026) parent 9ed234b

Showing 8 changed files with +416 and −14

modified atlas/qwen3-0.6b-4bit/probes/v2/confidence.md +2 −1
@@ -15,7 +15,8 @@ Level : 1
15 15 Seeds : 5
16 16 Prompt sets: 6 (token-balanced, structure-borne; overlap certificates in manifest)
17 17 Methods in agreement : 1 (linear probes only — Level 2 requires a second method)
18 Causal verification : none (observational; Level 3 requires intervention)
18 +Causal verification : ATTEMPTED AND FAILED (expC run #1 layer-skip: top-5
19 + differential layers not confirmed; survival 0/1)
19 20 ```
20 21
21 22 Per-property verdicts (differential real−twin, mean pooling):
modified atlas/qwen3-0.6b-4bit/probes/v2/mapcard.json +4 −2
@@ -59,14 +59,16 @@
59 59 "class token-overlap certificates in promptset manifest"
60 60 ],
61 61 "methods_in_agreement": [],
62 "interventions": [],
62 + "interventions": [
63 + "layer-skip ablation (expC run #1, 2026-08-12): top-5 differential layers NOT confirmed — damage below random-5 mean; see experiments/micro/expC_causal_verification/analysis.md"
64 + ],
63 65 "replication_rate": 0.5366,
64 66 "per_dataset_agreement": null,
65 67 "ablation_schemes": [],
66 68 "featurizer_class": "natural-basis (mean-pooled + last-token residual)",
67 69 "intervention_protocol": "none (observational map — Level 1 by design)",
68 70 "negative_result": false,
69 "notes": "DIFFERENTIAL map (real minus random-init twin), per the doctrine adopted after v1. Mixed outcome by property: agreement and arith_valid carry trained-model signal above the architecture prior; word_order is null-dominated and flagged as such. Strict twin gate (<0.05) still fails on word_order/agreement — only differential claims are published.",
71 + "notes": "DIFFERENTIAL map (real minus random-init twin), per the doctrine adopted after v1. Mixed outcome by property: agreement and arith_valid carry trained-model signal above the architecture prior; word_order is null-dominated and flagged as such. Strict twin gate (<0.05) still fails on word_order/agreement — only differential claims are published. CAUSAL CHECK: expC run #1 layer-skip ablation did NOT confirm the top differential layers (survival 0/1); map remains Level 1 and its layer ranking must not be read as causal.",
70 72 "author": "Simon-Pierre Boucher",
71 73 "contact": "contact@spboucher.ai",
72 74 "website": "https://modelmap.io",
added experiments/micro/expC_causal_verification/analysis.md +68 −0
@@ -0,0 +1,68 @@
1 +---
2 +project: modelmap
3 +document: expC_causal_verification — analysis (run #1)
4 +author: Simon-Pierre Boucher
5 +contact: contact@spboucher.ai
6 +website: https://modelmap.io
7 +created: 2026-08-12
8 +modified: 2026-08-12
9 +status: reviewed
10 +---
11 +
12 +# Analysis — expC run #1: the agreement probe map FAILS causal verification
13 +
14 +Run: `results/expC_causal_verification/20260812T064534Z/results.json` (5.7 s).
15 +Model: Qwen3-0.6B-4bit. Input map: agreement differential (probes/v2).
16 +Intervention: layer-skip ablation via the Tap layer. Hypothesis and the
17 +binary verdict criterion registered before the run.
18 +
19 +```text
20 +Hypothesis : the top-5 differential layers (by real−twin probe
21 + selectivity) causally support agreement behavior —
22 + skip damage ≥ p95 of random-5 draws AND ≥ 2× their
23 + mean.
24 +Falsification criterion : top-5 damage inside the random distribution.
25 +Result : FALSIFIED — survival 0/1. Baseline grammatical
26 + margin +4.63 (sanity holds: the model robustly
27 + prefers correct agreement). Skip damage:
28 + top-5 (layers 17,18,19,21,22) = +2.17 —
29 + BELOW the random-5 mean (+3.14, p95 +4.85, 20
30 + draws); bottom-5 differential (layers 0–4) = +4.99,
31 + the LARGEST of all conditions.
32 +Interpretation : Level 1 for the negative claim. Where agreement
33 + information is most linearly decodable above the
34 + architecture null (late-mid layers) is NOT where
35 + the computation is causally load-bearing for the
36 + behavior. This is the Hase-class dissociation
37 + (localization ≠ causal support — notes §4.6)
38 + measured end-to-end in our own pipeline, on a
39 + pre-registered binary verdict. The
40 + correlational→causal survival ledger opens at 0/1.
41 + Declared caveats bite exactly as registered:
42 + (a) layer-skip is coarse — early-layer skips
43 + plausibly cause GENERAL degradation, not
44 + agreement-specific damage (bottom-5 +4.99 reads as
45 + "the model breaks", not "agreement lives at layers
46 + 0–4"); a specificity control (margin damage
47 + normalized by general perplexity damage) is the
48 + registered follow-up; (b) held-out pairs = 89
49 + after dedup against every v2 probe text (below the
50 + planned 200 — combo pool exhausted; enlarge banks
51 + next run).
52 +Next experiment : run #2 with (1) perplexity-normalized specificity
53 + scores per skip condition, (2) direction-level
54 + intervention (LEACE erasure of the agreement
55 + direction in the forward pass at layer ℓ) — a
56 + surgical test the probe map CAN legitimately pass
57 + or fail, (3) larger held-out bank. Then the same
58 + protocol on arith_valid.
59 +```
60 +
61 +## Ledger
62 +
63 +| # | Correlational claim | Intervention | Survives? |
64 +|---|---|---|---|
65 +| 1 | agreement top-5 differential layers (probes/v2) | layer-skip ablation | **No** (0/1) |
66 +
67 +The published survival rate is the running fraction of this table — the
68 +charter §8.3 metric, now live.
modified experiments/micro/expC_causal_verification/hypothesis.md +43 −8
@@ -1,23 +1,58 @@
1 1 ---
2 2 project: modelmap
3 document: expC_causal_verification — hypothesis
3 +document: expC_causal_verification — hypothesis (run #1, registered before run)
4 4 author: Simon-Pierre Boucher
5 5 contact: contact@spboucher.ai
6 6 website: https://modelmap.io
7 7 created: 2026-08-12
8 status: draft
8 +modified: 2026-08-12
9 +status: reviewed
9 10 ---
10 11
11 # Hypothesis — expC_causal_verification
12 +# Hypothesis — expC run #1 (does the agreement probe map survive ablation?)
12 13
13 > Causal verification pipeline — correlational-to-causal survival rate
14 +> Causal verification pipeline — correlational→causal survival rate.
15 +> Registered 2026-08-12 **before** the run. Input: the differential
16 +> (real−twin) agreement probe map from expA run #2
17 +> (atlas/qwen3-0.6b-4bit/probes/v2). Intervention: layer-skip ablation
18 +> (block contribution removed, residual passes through) via the Tap layer.
14 19
15 20 ```text
16 Hypothesis : (to be registered before the first run)
17 Falsification criterion : (explicit kill-number, registered in advance)
18 Method : (including controls)
19 Baseline / null : (shuffled labels / random directions / random init)
21 +Hypothesis : The top-5 layers of the agreement DIFFERENTIAL map
22 + causally support agreement behavior: skipping them
23 + damages the model's grammatical preference
24 + (mean logit margin of the correct verb form over
25 + the incorrect one, on held-out minimal pairs) more
26 + than skipping 5 random non-top layers — top-5
27 + damage ≥ the 95th percentile of the random-5
28 + damage distribution (20 draws), and ≥ 2× its mean.
29 +Falsification criterion : Top-5 damage inside the random-5 distribution
30 + (< 95th percentile) — the correlational map fails
31 + causal verification; this survival datum (0 or 1)
32 + is the first entry of expC's published
33 + correlational→causal survival rate, either way.
34 +Method : Behavioral metric: margin = mean over minimal
35 + pairs of [logit(correct is/are) − logit(incorrect)]
36 + at the verb position, teacher-forced prefix
37 + "The <noun(s)> <location>". Held-out pairs: fresh
38 + seed, deduped against every v2 probe text.
39 + Conditions: baseline (no skip); skip top-5
40 + differential layers; skip 5 random layers
41 + (excluding top-5), 20 draws; skip bottom-5
42 + differential layers (second control). Same model
43 + (Qwen3-0.6B-4bit), deterministic forwards.
44 +Baseline / null : baseline margin (sanity: must be > 0, i.e. the
45 + model actually prefers grammatical agreement —
46 + else the whole question is moot at this size);
47 + random-5 and bottom-5 skip distributions.
20 48 Result : (pending)
21 49 Interpretation : (pending — with explicit confidence level)
22 50 Next experiment : (pending)
23 51 ```
52 +
53 +**Declared caveats.** Layer-skip is a coarse intervention (removes ALL of a
54 +block's computation, not the agreement direction specifically) — a confirmed
55 +result licenses "these layers causally support agreement behavior", NOT
56 +"agreement is localized to these layers" (hydra/backup effects, notes §4.4,
57 +can mask redundancy). Direction-level interventions (LEACE-style erasure in
58 +the forward pass) are the registered follow-up.
added experiments/micro/expC_causal_verification/implementation/benchmark.py +165 −0
@@ -0,0 +1,165 @@
1 +#!/usr/bin/env python3
2 +# =============================================================================
3 +# Project : modelmap
4 +# File : experiments/micro/expC_causal_verification/implementation/benchmark.py
5 +# Purpose : Run #1 — layer-skip ablation test of the agreement probe map
6 +# Author : Simon-Pierre Boucher
7 +# Contact : contact@spboucher.ai
8 +# Website : https://modelmap.io
9 +# Created : 2026-08-12
10 +# Modified : 2026-08-12
11 +# Platform : macOS / Apple Silicon (arm64) — MLX / Metal
12 +# License : All rights reserved (research code)
13 +# =============================================================================
14 +"""expC run #1 (hypothesis registered before this run).
15 +
16 +Does the agreement DIFFERENTIAL probe map (expA run #2) survive causal
17 +testing? Skip the map's top-5 layers vs 20 random-5 draws vs bottom-5,
18 +measure the drop in grammatical-agreement logit margin on held-out
19 +minimal pairs.
20 +"""
21 +
22 +from __future__ import annotations
23 +
24 +import json
25 +import random
26 +import subprocess
27 +import sys
28 +import time
29 +from pathlib import Path
30 +
31 +import numpy as np
32 +
33 +ROOT = Path(__file__).resolve().parents[4]
34 +sys.path.insert(0, str(ROOT / "src"))
35 +sys.path.insert(0, str(ROOT / "benchmarks"))
36 +from hardware_manifest import manifest
37 +
38 +from modelmap.capture.mlx_capture import install_taps
39 +
40 +MODEL = "mlx-community/Qwen3-0.6B-4bit"
41 +MAP_JSON = ROOT / "atlas" / "qwen3-0.6b-4bit" / "probes" / "v2" / "map.json"
42 +K = 5
43 +N_RANDOM_DRAWS = 20
44 +N_PAIRS = 200
45 +SEED = 777
46 +
47 +NOUN_PAIRS = [("key", "keys"), ("crate", "crates"), ("report", "reports"), ("valve", "valves"),
48 + ("ticket", "tickets"), ("ladder", "ladders"), ("sample", "samples"), ("cable", "cables"),
49 + ("permit", "permits"), ("beacon", "beacons"), ("filter", "filters"), ("stamp", "stamps")]
50 +NEAR = ["near the entrance", "beside the counter", "under the shelf", "behind the gate",
51 + "next to the archive", "along the corridor", "opposite the office", "inside the vault",
52 + "across the yard", "above the workbench"]
53 +
54 +
55 +def held_out_pairs(rng: random.Random, forbidden: set[str]):
56 + """Minimal pairs (prefix, singular?) deduped against every v2 probe text."""
57 + pairs, seen = [], set()
58 + combos = [(n, loc) for n in NOUN_PAIRS for loc in NEAR]
59 + rng.shuffle(combos)
60 + for (sg, pl), loc in combos * 4:
61 + for noun, singular in ((sg, True), (pl, False)):
62 + prefix = f"The {noun} {loc}"
63 + probe_like = f"{prefix} is" if singular else f"{prefix} are"
64 + if prefix in seen or any(probe_like in f for f in forbidden):
65 + continue
66 + seen.add(prefix)
67 + pairs.append({"prefix": prefix, "singular": singular})
68 + if len(pairs) >= N_PAIRS:
69 + return pairs
70 + return pairs
71 +
72 +
73 +def main() -> int:
74 + import mlx.core as mx
75 + from mlx_lm import load
76 +
77 + t0 = time.time()
78 + model, tokenizer = load(MODEL)
79 + taps = install_taps(model)
80 + n_layers = len(taps)
81 +
82 + # top/bottom differential layers from the published map (mean pooling)
83 + mdoc = json.loads(MAP_JSON.read_text())
84 + agree = mdoc["properties"]["agreement"]
85 + diff = [(a["layer"], a["selectivity_mean"] - t["selectivity_mean"])
86 + for a, t in zip(agree["per_layer"]["A"], agree["twin_null_per_layer_A"])]
87 + ranked = [l for l, _ in sorted(diff, key=lambda x: -x[1])]
88 + top_k, bottom_k = ranked[:K], ranked[-K:]
89 + print(f"top-{K} differential layers: {sorted(top_k)} | bottom-{K}: {sorted(bottom_k)}")
90 +
91 + # held-out minimal pairs
92 + forbidden = set()
93 + for f in (ROOT / "benchmarks" / "promptsets").glob("agreement_*.jsonl"):
94 + forbidden.update(json.loads(l)["text"] for l in f.read_text().splitlines())
95 + rng = random.Random(SEED)
96 + pairs = held_out_pairs(rng, forbidden)
97 + print(f"held-out minimal pairs: {len(pairs)}")
98 +
99 + tok_is = tokenizer.encode(" is")
100 + tok_are = tokenizer.encode(" are")
101 + assert len(tok_is) == 1 and len(tok_are) == 1, "verb forms must be single tokens"
102 + id_is, id_are = tok_is[0], tok_are[0]
103 + prefix_ids = [tokenizer.encode(p["prefix"]) for p in pairs]
104 +
105 + def margin(skip_layers: set[int]) -> float:
106 + for i, t in enumerate(taps):
107 + t.skip = i in skip_layers
108 + margins = []
109 + for ids, p in zip(prefix_ids, pairs):
110 + logits = model(mx.array([ids]))[0, -1, :]
111 + mx.eval(logits)
112 + m = float(logits[id_is] - logits[id_are])
113 + margins.append(m if p["singular"] else -m)
114 + for t in taps:
115 + t.skip = False
116 + return float(np.mean(margins))
117 +
118 + baseline = margin(set())
119 + print(f"baseline margin: {baseline:+.4f}")
120 + top_m = margin(set(top_k))
121 + bottom_m = margin(set(bottom_k))
122 + candidates = [i for i in range(n_layers) if i not in set(top_k)]
123 + random_ms = []
124 + for d in range(N_RANDOM_DRAWS):
125 + draw = set(random.Random(SEED + 1 + d).sample(candidates, K))
126 + random_ms.append(margin(draw))
127 + dmg = lambda m: baseline - m
128 + rd = np.array([dmg(m) for m in random_ms])
129 + p95 = float(np.percentile(rd, 95))
130 + verdict = bool(dmg(top_m) >= p95 and dmg(top_m) >= 2 * rd.mean())
131 + print(f"damage: top-{K}={dmg(top_m):+.4f} bottom-{K}={dmg(bottom_m):+.4f} "
132 + f"random mean={rd.mean():+.4f} p95={p95:+.4f} -> survives={verdict}")
133 +
134 + commit = subprocess.run(["git", "rev-parse", "HEAD"], cwd=ROOT,
135 + capture_output=True, text=True, check=False).stdout.strip()
136 + ts = time.strftime("%Y%m%dT%H%M%SZ", time.gmtime())
137 + outdir = ROOT / "results" / "expC_causal_verification" / ts
138 + outdir.mkdir(parents=True)
139 + doc = {
140 + "experiment": "expC_causal_verification", "run": 1,
141 + "scope": "layer-skip ablation of the agreement differential probe map",
142 + "commit": commit,
143 + "config": {"model": MODEL, "k": K, "n_random_draws": N_RANDOM_DRAWS,
144 + "n_pairs": len(pairs), "seed": SEED,
145 + "map_source": str(MAP_JSON.relative_to(ROOT)),
146 + "top_layers": sorted(top_k), "bottom_layers": sorted(bottom_k)},
147 + "manifest": manifest(),
148 + "results": {
149 + "baseline_margin": baseline,
150 + "top_k_margin": top_m, "top_k_damage": dmg(top_m),
151 + "bottom_k_margin": bottom_m, "bottom_k_damage": dmg(bottom_m),
152 + "random_margins": random_ms,
153 + "random_damage_mean": float(rd.mean()),
154 + "random_damage_p95": p95,
155 + "survives_causal_test": verdict,
156 + },
157 + "wall_seconds": round(time.time() - t0, 1),
158 + }
159 + (outdir / "results.json").write_text(json.dumps(doc, indent=2) + "\n")
160 + print(f"results -> {outdir / 'results.json'} ({doc['wall_seconds']} s)")
161 + return 0
162 +
163 +
164 +if __name__ == "__main__":
165 + sys.exit(main())
modified research/LOG.md +35 −0
@@ -356,3 +356,38 @@ arith_valid: LEACE erasure as second method, then ablation on top layers
356 356 (expC entry). (3) The ~0.54 replication number seeds expD's design (more
357 357 seeds, tighter CIs). (4) The agreement differential map is candidate_02's
358 358 quantization-drift target.
359 +
360 +---
361 +
362 +## 2026-08-12 07:45 EDT — expC run #1: the agreement map FAILS causal verification (survival ledger opens 0/1)
363 +
364 +**Question.** Do the top-5 layers of the agreement differential probe map
365 +causally support agreement behavior under layer-skip ablation?
366 +
367 +**Result (5.7 s; pre-registered binary verdict).** FALSIFIED.
368 +Baseline grammatical margin +4.63 (the 0.6B model robustly prefers correct
369 +agreement — sanity holds). Skip damage: top-5 differential layers
370 +(17,18,19,21,22) = +2.17, BELOW the random-5 mean (+3.14; p95 +4.85, 20
371 +draws); bottom-5 differential (layers 0–4) = +4.99, the largest of all
372 +conditions. **Where agreement information is most decodable above the
373 +architecture null is not where the computation is causally load-bearing.**
374 +The Hase-class dissociation (localization ≠ causal support), measured
375 +end-to-end in our own pipeline within one day of standing it up.
376 +
377 +**Ledger.** The correlational→causal survival rate — the charter §8.3
378 +metric — is now live: 0/1. The atlas entry probes/v2 records the failed
379 +check in mapcard.interventions and confidence.md ("causal verification:
380 +attempted and failed"); the map stays Level 1 and its layer ranking is
381 +explicitly flagged as non-causal.
382 +
383 +**Caveats (registered in advance, both bit).** Layer-skip is coarse:
384 +bottom-5 damage plausibly reflects GENERAL degradation (early layers break
385 +everything), not agreement-specific structure. Held-out pairs 89 < planned
386 +200 (dedup exhausted the combo pool).
387 +
388 +**Decisions / next.** expC run #2: (1) perplexity-normalized specificity
389 +per skip condition; (2) direction-level intervention — LEACE-erase the
390 +agreement direction in the forward pass at layer ℓ (a surgical test the
391 +probe map can legitimately pass); (3) larger held-out bank; then the same
392 +protocol on arith_valid. The pipeline now demonstrably runs the full
393 +charter loop: register → measure → verify causally → publish either way.
added results/expC_causal_verification/20260812T064534Z/results.json +88 −0
@@ -0,0 +1,88 @@
1 +{
2 + "experiment": "expC_causal_verification",
3 + "run": 1,
4 + "scope": "layer-skip ablation of the agreement differential probe map",
5 + "commit": "9ed234bf4f8e787a8d6fbd01107c25567a403115",
6 + "config": {
7 + "model": "mlx-community/Qwen3-0.6B-4bit",
8 + "k": 5,
9 + "n_random_draws": 20,
10 + "n_pairs": 89,
11 + "seed": 777,
12 + "map_source": "atlas/qwen3-0.6b-4bit/probes/v2/map.json",
13 + "top_layers": [
14 + 17,
15 + 18,
16 + 19,
17 + 21,
18 + 22
19 + ],
20 + "bottom_layers": [
21 + 0,
22 + 1,
23 + 2,
24 + 3,
25 + 4
26 + ]
27 + },
28 + "manifest": {
29 + "author": "Simon-Pierre Boucher",
30 + "contact": "contact@spboucher.ai",
31 + "website": "https://modelmap.io",
32 + "chip": {
33 + "brand": "Apple M5 Max",
34 + "cores_total": 18,
35 + "cores_performance": 6,
36 + "cores_efficiency": 12
37 + },
38 + "memory": {
39 + "unified_gb": 48.0,
40 + "pagesize": 16384
41 + },
42 + "os": {
43 + "system": "Darwin",
44 + "version": "27.0",
45 + "arch": "arm64"
46 + },
47 + "software": {
48 + "python": "3.14.4",
49 + "numpy": "2.5.2",
50 + "mlx": "0.32.0",
51 + "torch": "2.13.0",
52 + "safetensors": "0.8.0"
53 + }
54 + },
55 + "results": {
56 + "baseline_margin": 4.631320224719101,
57 + "top_k_margin": 2.464185393258427,
58 + "top_k_damage": 2.167134831460674,
59 + "bottom_k_margin": -0.3570926966292135,
60 + "bottom_k_damage": 4.988412921348314,
61 + "random_margins": [
62 + 0.17854634831460675,
63 + 3.4838483146067416,
64 + 3.6688904494382024,
65 + 2.7650983146067416,
66 + 2.133426966292135,
67 + 1.8202247191011236,
68 + 0.6693293539325843,
69 + 3.425561797752809,
70 + 0.5043012640449438,
71 + 2.018960674157303,
72 + -0.20628511235955055,
73 + 0.3945751404494382,
74 + 0.6560086025280899,
75 + -0.21166169241573032,
76 + 4.489115168539326,
77 + -0.34467169943820225,
78 + 0.15458216292134833,
79 + 2.4719101123595504,
80 + 1.559691011235955,
81 + 0.14381803019662923
82 + ],
83 + "random_damage_mean": 3.142556728405899,
84 + "random_damage_p95": 4.849632417485955,
85 + "survives_causal_test": false
86 + },
87 + "wall_seconds": 5.7
88 +}
modified src/modelmap/capture/mlx_capture.py +11 −3
@@ -26,15 +26,23 @@ import numpy as np
26 26
27 27
28 28 class Tap:
29 """Wraps a decoder layer; optionally retains its output."""
29 + """Wraps a decoder layer; optionally retains its output or skips the block.
30 +
31 + skip=True implements layer ablation (the block's contribution is removed;
32 + the residual stream passes through unchanged) — the Gromov/ShortGPT-style
33 + intervention used by expC for causal verification.
34 + """
30 35
31 36 def __init__(self, layer):
32 37 self.layer = layer
33 38 self.retained = None
34 39 self.enabled = False
40 + self.skip = False
35 41
36 def __call__(self, *args, **kwargs):
37 out = self.layer(*args, **kwargs)
42 + def __call__(self, x, *args, **kwargs):
43 + if self.skip:
44 + return x
45 + out = self.layer(x, *args, **kwargs)
38 46 if self.enabled:
39 47 self.retained = out
40 48 return out
41 49