SPB Git

spb/modelmap Public License

Internal cartography of local LLMs on Apple Silicon — registered, gated, negative-first. Public atlas at modelmap.io.

Python 66.3% JavaScript 24.5% CSS 8.1% Shell 0.7%

expC run #4: BAND claim passes — first Level-2 atlas entry (causal map)

Early-band (2-15) mean specific damage +3.23..+3.35 across six fresh
direction sources (min +3.228 vs pre-registered bar 2.5; spread 0.12) on
a fresh behavioral bank (baseline +4.45). Late band unstable, no claim.
Published atlas/qwen3-0.6b-4bit/interventions/v1 (Level 2): a single
diff-of-means agreement direction, erased at any early-band layer,
removes ~73-75% of grammatical preference. The gate that refused run #3's
per-layer map passed the band version. Band map leads the home page.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Simon-Pierre Boucher committed 1 h ago (Aug 12, 2026) parent f7c298d

Showing 12 changed files with +1,236 and −67

added atlas/qwen3-0.6b-4bit/interventions/v1/confidence.md +36 −0
@@ -0,0 +1,36 @@
1 +---
2 +project: modelmap
3 +document: qwen3-0.6b-4bit/interventions/v1 — confidence
4 +author: Simon-Pierre Boucher
5 +contact: contact@spboucher.ai
6 +website: https://modelmap.io
7 +created: 2026-08-12
8 +status: reviewed
9 +---
10 +
11 +# Confidence — qwen3-0.6b-4bit / interventions / v1
12 +
13 +```text
14 +Level : 2
15 +Seeds : six fresh direction sources (disjoint halves of two promptsets
16 + + two fresh bootstraps); fresh behavioral bank
17 +Prompt sets: 3 (two estimation sets + held-out behavioral bank)
18 +Methods in agreement : 2 (diff-of-means probing; direction erasure) — shared
19 + estimator, hence Level 2 and not 3
20 +Causal verification : YES — rank-1 erasure with random-direction nulls
21 +```
22 +
23 +Pre-registered band claim (bar 2.5 on a +4.45 baseline margin):
24 +- Ahalf1: early-band mean +3.347
25 +- Ahalf2: early-band mean +3.228
26 +- bootA: early-band mean +3.302
27 +- Bhalf1: early-band mean +3.263
28 +- Bhalf2: early-band mean +3.326
29 +- bootB: early-band mean +3.270
30 +Minimum early-band mean across sources: +3.228 — claim PASSES.
31 +
32 +Granularity discipline: run #3's per-layer profile FAILED replication and was
33 +refused by this very gate; the published object is the BAND (layers 2–15).
34 +The late band (20–27) is displayed but carries no claim. This map is the
35 +causal counterpart of probes/v2, whose decodability ranking it contradicts
36 +(survival ledger 0/2) — both stay published, labeled by what they measure.
added atlas/qwen3-0.6b-4bit/interventions/v1/map.json +305 −0
@@ -0,0 +1,305 @@
1 +{
2 + "author": "Simon-Pierre Boucher",
3 + "contact": "contact@spboucher.ai",
4 + "website": "https://modelmap.io",
5 + "map_type": "interventions",
6 + "model_id": "mlx-community/Qwen3-0.6B-4bit",
7 + "claim": "BAND claim: erasing the diff-of-means agreement direction at any single layer in the early band (2-15) destroys most of the grammatical-agreement margin, replicated across six fresh direction estimates on a fresh behavioral bank. The late band (20-27) is reported but carries NO claim (declared estimator-unstable by run #3).",
8 + "behavior_metric": "logit margin correct-vs-incorrect verb, held-out minimal pairs",
9 + "baseline_margin": 4.450520833333333,
10 + "band": {
11 + "early_layers": [
12 + 2,
13 + 15
14 + ],
15 + "late_layers": [
16 + 20,
17 + 27
18 + ],
19 + "bar": 2.5,
20 + "early_means_per_source": [
21 + 3.3470672123015865,
22 + 3.2282172309027777,
23 + 3.3024708581349214,
24 + 3.263111901661706,
25 + 3.3256831093439985,
26 + 3.270344567677331
27 + ],
28 + "min_early_mean": 3.2282172309027777
29 + },
30 + "per_layer": {
31 + "mean": [
32 + 0.0042317708333339255,
33 + 0.02083333333333363,
34 + 2.0526258680555554,
35 + 3.2511393229166665,
36 + 0.89776611328125,
37 + 4.292575412326388,
38 + 3.8483106825086817,
39 + 3.572519938151041,
40 + 3.896267361111111,
41 + 2.8672688802083335,
42 + 4.019599066840278,
43 + 2.9330512152777786,
44 + 3.843017578125,
45 + 3.8774685329861107,
46 + 2.924452039930556,
47 + 3.7766927083333326,
48 + 4.187445746527779,
49 + 3.8103027343749996,
50 + 3.910481770833334,
51 + 3.8522135416666674,
52 + 3.426052517361111,
53 + 3.187174479166668,
54 + 2.733425564236112,
55 + 2.2225070529513893,
56 + 1.4873589409722214,
57 + 1.3905164930555547,
58 + 1.6435818142361114,
59 + 0.8997938368055557
60 + ],
61 + "min": [
62 + -0.029405381944443754,
63 + -0.005316840277777679,
64 + 2.032199435763889,
65 + 3.235975477430556,
66 + 0.884114583333333,
67 + 4.256130642361111,
68 + 3.8278537326388897,
69 + 3.5364040798611103,
70 + 3.861246744791667,
71 + 2.7926161024305554,
72 + 3.9384223090277777,
73 + 2.870795355902778,
74 + 3.7218967013888884,
75 + 3.7080620659722223,
76 + 2.6974283854166665,
77 + 3.6133897569444438,
78 + 4.05029296875,
79 + 3.554307725694444,
80 + 3.7395833333333335,
81 + 3.700846354166667,
82 + 2.9704318576388893,
83 + 2.4511176215277786,
84 + 1.5763888888888893,
85 + -0.569742838541667,
86 + -2.4296332465277786,
87 + -1.4393988715277786,
88 + -0.99951171875,
89 + 0.31629774305555536
90 + ],
91 + "random_direction_damage": [
92 + -0.0008680555555562464,
93 + -0.0028211805555562464,
94 + 2.340847439236111,
95 + 1.1391059027777772,
96 + 3.474690755208333,
97 + 0.05262586805555536,
98 + 0.4880642361111103,
99 + 0.7284071180555558,
100 + 0.4264322916666661,
101 + 1.4096137152777777,
102 + 0.24273003472222232,
103 + 1.309950086805555,
104 + 0.3861762152777777,
105 + 0.3366970486111107,
106 + 1.3014322916666665,
107 + 0.4308810763888893,
108 + 0.0016276041666660745,
109 + 0.15190972222222232,
110 + 0.04720052083333304,
111 + -0.040364583333333925,
112 + 0.09483506944444375,
113 + -0.05674913194444553,
114 + -0.11707899305555625,
115 + -0.05338541666666696,
116 + 0.028103298611111605,
117 + 0.0866970486111116,
118 + 0.06803385416666607,
119 + -0.0021701388888892836
120 + ]
121 + },
122 + "profiles_per_source": {
123 + "Ahalf1": [
124 + -0.029405381944443754,
125 + -0.005316840277777679,
126 + 2.068657769097222,
127 + 3.288953993055556,
128 + 0.920817057291667,
129 + 4.322129991319445,
130 + 3.8747287326388897,
131 + 3.5920681423611103,
132 + 3.9365234375,
133 + 2.9437391493055554,
134 + 4.149278428819444,
135 + 3.007921006944445,
136 + 3.9253472222222223,
137 + 3.9666883680555554,
138 + 3.0131835937499996,
139 + 3.8489040798611107,
140 + 4.263671875,
141 + 3.9997829861111107,
142 + 3.8670247395833335,
143 + 3.7986653645833335,
144 + 2.9704318576388893,
145 + 2.7823350694444455,
146 + 1.7347547743055558,
147 + 1.9698893229166665,
148 + 1.3794487847222214,
149 + 1.2564019097222214,
150 + 1.6738281250000004,
151 + 0.7954644097222223
152 + ],
153 + "Ahalf2": [
154 + 0.052951388888889284,
155 + 0.040256076388889284,
156 + 2.068901909722222,
157 + 3.2481011284722228,
158 + 0.892659505208333,
159 + 4.256130642361111,
160 + 3.8350151909722228,
161 + 3.5364040798611103,
162 + 3.919596354166667,
163 + 2.7926161024305554,
164 + 3.9950629340277777,
165 + 2.909776475694445,
166 + 3.7218967013888884,
167 + 3.7080620659722223,
168 + 2.6974283854166665,
169 + 3.6133897569444438,
170 + 4.05029296875,
171 + 3.554307725694444,
172 + 3.7395833333333335,
173 + 3.9003906250000004,
174 + 3.2191297743055562,
175 + 2.4511176215277786,
176 + 1.5763888888888893,
177 + -0.569742838541667,
178 + -2.4296332465277786,
179 + -1.4393988715277786,
180 + -0.99951171875,
181 + 0.31629774305555536
182 + ],
183 + "bootA": [
184 + -0.008572048611110716,
185 + 0.018446180555556246,
186 + 2.036593967013889,
187 + 3.259657118055556,
188 + 0.912272135416667,
189 + 4.290757921006945,
190 + 3.8606906467013897,
191 + 3.584581163194444,
192 + 3.877766927083334,
193 + 2.8962131076388884,
194 + 4.020941840277778,
195 + 2.930284288194445,
196 + 3.8477105034722223,
197 + 3.9113498263888893,
198 + 2.9856770833333335,
199 + 3.8200954861111107,
200 + 4.232096354166667,
201 + 3.8371853298611107,
202 + 4.0078125,
203 + 3.9334309895833335,
204 + 3.6519097222222228,
205 + 3.2728949652777786,
206 + 2.6348198784722223,
207 + 2.4124348958333335,
208 + 2.005099826388888,
209 + 1.606336805555555,
210 + 1.8066406250000004,
211 + 0.8556857638888888
212 + ],
213 + "Bhalf1": [
214 + -0.005967881944443754,
215 + 0.015190972222222321,
216 + 2.060845269097222,
217 + 3.237765842013889,
218 + 0.887776692708333,
219 + 4.281765407986111,
220 + 3.8278537326388897,
221 + 3.542344835069444,
222 + 3.875325520833334,
223 + 2.8387586805555554,
224 + 3.9384223090277777,
225 + 2.870795355902778,
226 + 3.8202853732638884,
227 + 3.8525933159722223,
228 + 2.9090169270833335,
229 + 3.7400173611111107,
230 + 4.126627604166667,
231 + 3.7250434027777772,
232 + 3.8818359375,
233 + 3.7548828125000004,
234 + 3.491102430555556,
235 + 3.3700629340277786,
236 + 3.2643771701388893,
237 + 3.141764322916667,
238 + 2.7155490451388884,
239 + 2.4268120659722214,
240 + 2.977213541666667,
241 + 1.2258029513888888
242 + ],
243 + "Bhalf2": [
244 + 0.007703993055556246,
245 + 0.029513888888889284,
246 + 2.032199435763889,
247 + 3.235975477430556,
248 + 0.888956705729167,
249 + 4.313869900173611,
250 + 3.8546481662326397,
251 + 3.623155381944444,
252 + 3.907145182291667,
253 + 2.8999565972222223,
254 + 4.044542100694444,
255 + 2.986273871527778,
256 + 3.9251844618055554,
257 + 3.9627821180555554,
258 + 3.0100911458333335,
259 + 3.8747829861111107,
260 + 4.273600260416667,
261 + 3.9958767361111107,
262 + 4.093912760416667,
263 + 4.025065104166667,
264 + 3.724500868055556,
265 + 3.752712673611112,
266 + 3.6870659722222228,
267 + 3.225260416666667,
268 + 2.6506076388888884,
269 + 2.2886284722222214,
270 + 2.0901692708333335,
271 + 1.0425347222222223
272 + ],
273 + "bootB": [
274 + 0.008680555555556246,
275 + 0.02690972222222232,
276 + 2.048556857638889,
277 + 3.2363823784722228,
278 + 0.884114583333333,
279 + 4.290798611111111,
280 + 3.836927625868056,
281 + 3.556566026475694,
282 + 3.861246744791667,
283 + 2.8323296440972223,
284 + 3.9693467881944438,
285 + 2.893256293402778,
286 + 3.8176812065972223,
287 + 3.8633355034722223,
288 + 2.9313151041666665,
289 + 3.7629665798611107,
290 + 4.178385416666667,
291 + 3.749620225694444,
292 + 3.8727213541666665,
293 + 3.700846354166667,
294 + 3.4992404513888893,
295 + 3.4939236111111116,
296 + 3.5031467013888893,
297 + 3.155436197916667,
298 + 2.6030815972222214,
299 + 2.204318576388888,
300 + 2.313151041666667,
301 + 1.1629774305555558
302 + ]
303 + },
304 + "source_results": "results/expC_causal_verification/20260812T071954Z/results.json"
305 +}
added atlas/qwen3-0.6b-4bit/interventions/v1/mapcard.json +76 −0
@@ -0,0 +1,76 @@
1 +{
2 + "map_id": "atlas/qwen3-0.6b-4bit/interventions/v1",
3 + "map_type": "interventions",
4 + "model_id": "mlx-community/Qwen3-0.6B-4bit",
5 + "model_hash": "392e8d466d56100ada00eb82031fb854297fc9e389b7d303eba3af114e87bce2",
6 + "quantization": "q4 (mlx)",
7 + "commit": "f7c298d2859ed43654d388618b7274266e7e562b",
8 + "config": "results/expC_causal_verification/20260812T071954Z/results.json",
9 + "created": "2026-08-12",
10 + "hardware_manifest": {
11 + "author": "Simon-Pierre Boucher",
12 + "contact": "contact@spboucher.ai",
13 + "website": "https://modelmap.io",
14 + "chip": {
15 + "brand": "Apple M5 Max",
16 + "cores_total": 18,
17 + "cores_performance": 6,
18 + "cores_efficiency": 12
19 + },
20 + "memory": {
21 + "unified_gb": 48.0,
22 + "pagesize": 16384
23 + },
24 + "os": {
25 + "system": "Darwin",
26 + "version": "27.0",
27 + "arch": "arm64"
28 + },
29 + "software": {
30 + "python": "3.14.4",
31 + "numpy": "2.5.2",
32 + "mlx": "0.32.0",
33 + "torch": "2.13.0",
34 + "safetensors": "0.8.0"
35 + }
36 + },
37 + "confidence_level": 2,
38 + "regenerate_command": ".venv/bin/python experiments/micro/expC_causal_verification/implementation/benchmark_v4.py && .venv/bin/python experiments/micro/expC_causal_verification/implementation/make_interventions_mapcard.py",
39 + "seeds": [
40 + 780,
41 + "Ahalf1",
42 + "Ahalf2",
43 + "bootA",
44 + "Bhalf1",
45 + "Bhalf2",
46 + "bootB"
47 + ],
48 + "prompt_sets": [
49 + "agreement_A halves (direction est.)",
50 + "agreement_B halves (direction est.)",
51 + "fresh held-out minimal-pair bank (8 unseen locations)"
52 + ],
53 + "controls": [
54 + "random-direction erasure per layer (3 dirs, netted out)",
55 + "six-source replication with pre-registered bar (runs #3-#4)",
56 + "late band excluded as estimator-unstable (run #3 refusal)"
57 + ],
58 + "methods_in_agreement": [
59 + "difference-in-means probing (direction exists, decodable)",
60 + "direction erasure (causally load-bearing)"
61 + ],
62 + "interventions": [
63 + "rank-1 direction erasure at each layer's output, all positions"
64 + ],
65 + "replication_rate": 0.7254,
66 + "per_dataset_agreement": null,
67 + "ablation_schemes": [],
68 + "featurizer_class": "linear (difference-in-means direction)",
69 + "intervention_protocol": "erase h' = h − ⟨h−μ,u⟩u at layer ℓ; logit-margin metric; random-direction null netted out; BAND granularity",
70 + "negative_result": false,
71 + "notes": "Level 2, NOT 3: the two agreeing methods share the diff-of-means estimator; activation-addition steering is the registered Level-3 path. This entry exists because run #3's per-layer version was REFUSED by the gate — the band is the granularity that replicates. Anti-correlates with probes/v2 layer ranking (survival ledger 0/2): decodability peaks ≠ causal joints.",
72 + "author": "Simon-Pierre Boucher",
73 + "contact": "contact@spboucher.ai",
74 + "website": "https://modelmap.io",
75 + "schema_version": "0.1"
76 +}
added atlas/qwen3-0.6b-4bit/interventions/v1/provenance.json +63 −0
@@ -0,0 +1,63 @@
1 +{
2 + "author": "Simon-Pierre Boucher",
3 + "contact": "contact@spboucher.ai",
4 + "website": "https://modelmap.io",
5 + "model_id": "mlx-community/Qwen3-0.6B-4bit",
6 + "map_type": "interventions",
7 + "version": "v1",
8 + "commit": "f7c298d2859ed43654d388618b7274266e7e562b",
9 + "model_hash": "392e8d466d56100ada00eb82031fb854297fc9e389b7d303eba3af114e87bce2",
10 + "config": {
11 + "model": "mlx-community/Qwen3-0.6B-4bit",
12 + "sources": [
13 + "Ahalf1",
14 + "Ahalf2",
15 + "bootA",
16 + "Bhalf1",
17 + "Bhalf2",
18 + "bootB"
19 + ],
20 + "bar": 2.5,
21 + "early_layers": [
22 + 2,
23 + 15
24 + ],
25 + "late_layers": [
26 + 20,
27 + 27
28 + ],
29 + "n_random_dirs": 3,
30 + "n_pairs": 192,
31 + "seed": 780
32 + },
33 + "seed": 780,
34 + "hardware_manifest": {
35 + "author": "Simon-Pierre Boucher",
36 + "contact": "contact@spboucher.ai",
37 + "website": "https://modelmap.io",
38 + "chip": {
39 + "brand": "Apple M5 Max",
40 + "cores_total": 18,
41 + "cores_performance": 6,
42 + "cores_efficiency": 12
43 + },
44 + "memory": {
45 + "unified_gb": 48.0,
46 + "pagesize": 16384
47 + },
48 + "os": {
49 + "system": "Darwin",
50 + "version": "27.0",
51 + "arch": "arm64"
52 + },
53 + "software": {
54 + "python": "3.14.4",
55 + "numpy": "2.5.2",
56 + "mlx": "0.32.0",
57 + "torch": "2.13.0",
58 + "safetensors": "0.8.0"
59 + }
60 + },
61 + "created": "2026-08-12",
62 + "source_results": "results/expC_causal_verification/20260812T071954Z/results.json"
63 +}
modified experiments/micro/expC_causal_verification/analysis.md +43 −0
@@ -177,3 +177,46 @@ Next experiment : run #4 with the NARROWER pre-registered claim:
177 177 map as an early-band BAND claim (not a per-layer
178 178 ranking), Level 2.
179 179 ```
180 +
181 +---
182 +
183 +# Analysis — expC run #4: the BAND claim passes — first Level-2 atlas entry
184 +
185 +Run: `results/expC_causal_verification/20260812T071954Z/results.json`.
186 +Everything fresh: six direction sources (disjoint halves of A and B + two
187 +new bootstraps), new behavioral bank (8 unseen locations, 192 pairs,
188 +baseline margin +4.45). Claim, bar, and publication rule registered before
189 +the run.
190 +
191 +```text
192 +Hypothesis : early-band (2–15) mean specific damage ≥ 2.5 in
193 + EVERY fresh source; late band reported, no claim.
194 +Result : CONFIRMED — early-band means: Ahalf1 +3.35,
195 + Ahalf2 +3.23, bootA +3.30, Bhalf1 +3.26,
196 + Bhalf2 +3.33, bootB +3.27 (min +3.228 vs bar
197 + 2.5). Spread across six estimators: 0.12 — the
198 + band is a tight, estimator-stable causal object.
199 + Late band again unstable (+0.27…+2.83), as
200 + declared; it carries no claim.
201 +Interpretation : Level 2, published: atlas/qwen3-0.6b-4bit/
202 + interventions/v1 — the project's first Level-2
203 + entry and first causal map. The claim: a single
204 + linear direction (diff-of-means over correct vs
205 + violated agreement), erased at ANY one layer in
206 + 2–15, removes ~73–75% of the model's grammatical
207 + preference on held-out pairs, replicated across
208 + six independent direction estimates and netted
209 + against random-direction damage. NOT Level 3:
210 + both agreeing methods share the diff-of-means
211 + estimator — activation-addition steering is the
212 + registered L3 path. Granularity discipline paid
213 + off: the per-layer version of this map was
214 + refused (run #3); the band version replicates.
215 +Next experiment : (1) L3 path: steering (add the direction) should
216 + INCREASE the margin on violated-preference pairs;
217 + (2) same band protocol on arith_valid; (3)
218 + quantization drift of the band map (candidate_02:
219 + does the early band move under Q8/FP16?);
220 + (4) cross-model: does the band replicate on
221 + Qwen3-1.7B (expG entry)?
222 +```
modified experiments/micro/expC_causal_verification/hypothesis.md +37 −0
@@ -140,3 +140,40 @@ Result : (pending)
140 140 Interpretation : (pending)
141 141 Next experiment : (pending)
142 142 ```
143 +
144 +---
145 +
146 +# Hypothesis — expC run #4 (the narrowed BAND claim)
147 +
148 +Registered 2026-08-12 **before** the run, after run #3's gate refusal
149 +showed the early band replicates (+2.98…+3.34) while the late band is
150 +estimator noise. Everything fresh: direction estimates from DISJOINT HALVES
151 +of agreement_A and agreement_B (never used whole-set estimates), two new
152 +bootstrap seeds, and a NEW held-out behavioral bank (8 locations unseen by
153 +any prior run or promptset).
154 +
155 +```text
156 +Hypothesis : EARLY-BAND claim only — mean specific damage over
157 + layers 2–15 is ≥ 2.5 (on the fresh bank's
158 + baseline-margin scale) in EVERY one of six fresh
159 + direction sources (A-half-1, A-half-2, B-half-1,
160 + B-half-2, bootA-fresh, bootB-fresh). The late
161 + band (20–27) is REPORTED but carries no claim —
162 + declared unstable by run #3.
163 +Falsification criterion : any source with early-band mean < 2.5 → the band
164 + claim fails replication too; the direction-erasure
165 + program is closed for agreement at 0.6B and expC
166 + pivots to steering-based verification.
167 +Method : full 28-layer scan per source (map format
168 + unchanged); shared random-direction controls
169 + (3/layer) on the fresh bank; margins as before.
170 +Publication rule : pass → atlas interventions/v1 published as a BAND
171 + map at Level 2: map.json carries the six profiles,
172 + their mean/min, the random-direction damage, and
173 + the band statistics; the chart plots mean, worst
174 + source, and the raw random-direction reference.
175 + Fail → refusal documented, program closed.
176 +Result : (pending)
177 +Interpretation : (pending)
178 +Next experiment : (pending)
179 +```
added experiments/micro/expC_causal_verification/implementation/benchmark_v4.py +191 −0
@@ -0,0 +1,191 @@
1 +#!/usr/bin/env python3
2 +# =============================================================================
3 +# Project : modelmap
4 +# File : experiments/micro/expC_causal_verification/implementation/benchmark_v4.py
5 +# Purpose : Run #4 — the narrowed early-band claim, everything fresh
6 +# Author : Simon-Pierre Boucher
7 +# Contact : contact@spboucher.ai
8 +# Website : https://modelmap.io
9 +# Created : 2026-08-12
10 +# Modified : 2026-08-12
11 +# Platform : macOS / Apple Silicon (arm64) — MLX / Metal
12 +# License : All rights reserved (research code)
13 +# =============================================================================
14 +"""expC run #4 (hypothesis registered before this run).
15 +
16 +Band claim: early-band (layers 2-15) mean specific damage >= 2.5 in EVERY
17 +of six fresh direction sources (disjoint halves of A and B + two fresh
18 +bootstraps), on a fresh behavioral bank. Late band reported, no claim.
19 +"""
20 +
21 +from __future__ import annotations
22 +
23 +import json
24 +import random
25 +import subprocess
26 +import sys
27 +import time
28 +from pathlib import Path
29 +
30 +import numpy as np
31 +
32 +ROOT = Path(__file__).resolve().parents[4]
33 +sys.path.insert(0, str(ROOT / "src"))
34 +sys.path.insert(0, str(ROOT / "benchmarks"))
35 +from hardware_manifest import manifest
36 +
37 +from modelmap.capture.mlx_capture import capture_pooled, install_taps
38 +
39 +MODEL = "mlx-community/Qwen3-0.6B-4bit"
40 +N_RANDOM_DIRS = 3
41 +SEED = 780
42 +EARLY = slice(2, 16)
43 +LATE = slice(20, 28)
44 +BAR = 2.5
45 +
46 +NOUN_PAIRS = [("key", "keys"), ("crate", "crates"), ("report", "reports"), ("valve", "valves"),
47 + ("ticket", "tickets"), ("ladder", "ladders"), ("sample", "samples"), ("cable", "cables"),
48 + ("permit", "permits"), ("beacon", "beacons"), ("filter", "filters"), ("stamp", "stamps")]
49 +FRESH_NEAR = ["past the boiler room", "near the loading dock", "behind the ticket booth",
50 + "under the mezzanine", "beside the flagpole", "opposite the greenhouse",
51 + "inside the stairwell", "along the towpath"]
52 +
53 +
54 +def directions_from(reps, labels, idx):
55 + dirs, mus = [], []
56 + r, l = reps[idx], labels[idx]
57 + for layer in range(reps.shape[1]):
58 + x = r[:, layer, :]
59 + mu = x.mean(0)
60 + u = x[l == "correct"].mean(0) - x[l == "violated"].mean(0)
61 + u = u / (np.linalg.norm(u) + 1e-8)
62 + dirs.append(u); mus.append(mu)
63 + return dirs, mus
64 +
65 +
66 +def main() -> int:
67 + import mlx.core as mx
68 + from mlx_lm import load
69 +
70 + t0 = time.time()
71 + model, tokenizer = load(MODEL)
72 + taps = install_taps(model)
73 + n_layers = len(taps)
74 +
75 + rng = random.Random(SEED)
76 + combos = [(n, loc) for n in NOUN_PAIRS for loc in FRESH_NEAR]
77 + rng.shuffle(combos)
78 + pairs = []
79 + for (sg, pl), loc in combos:
80 + pairs.append({"prefix": f"The {sg} {loc}", "singular": True})
81 + pairs.append({"prefix": f"The {pl} {loc}", "singular": False})
82 + prefix_ids = [tokenizer.encode(p["prefix"]) for p in pairs]
83 + id_is, id_are = tokenizer.encode(" is")[0], tokenizer.encode(" are")[0]
84 +
85 + def margin() -> float:
86 + out = []
87 + for ids, p in zip(prefix_ids, pairs):
88 + logits = model(mx.array([ids]))[0, -1, :]
89 + mx.eval(logits)
90 + m = float(logits[id_is] - logits[id_are])
91 + out.append(m if p["singular"] else -m)
92 + return float(np.mean(out))
93 +
94 + def erase_fn(u_np, mu_np):
95 + u = mx.array(u_np.astype(np.float32))
96 + mu = mx.array(mu_np.astype(np.float32))
97 + def fn(out):
98 + h = out.astype(mx.float32)
99 + coef = ((h - mu) * u).sum(axis=-1, keepdims=True)
100 + return (h - coef * u).astype(out.dtype)
101 + return fn
102 +
103 + base_m = margin()
104 + print(f"fresh bank: {len(pairs)} pairs | baseline margin {base_m:+.4f}", flush=True)
105 +
106 + # fresh direction sources: disjoint halves + fresh bootstraps
107 + reps_cache = {}
108 + for s in ("A", "B"):
109 + items = [json.loads(l) for l in
110 + (ROOT / "benchmarks" / "promptsets" / f"agreement_{s}.jsonl").read_text().splitlines()]
111 + toks = [tokenizer.encode(it["text"]) for it in items]
112 + labels = np.array([it["label"] for it in items])
113 + print(f"capture agreement_{s}…", flush=True)
114 + reps_cache[s] = (capture_pooled(model, taps, toks)["mean"], labels)
115 +
116 + sources = {}
117 + for s in ("A", "B"):
118 + reps, labels = reps_cache[s]
119 + n = len(labels)
120 + perm = np.random.default_rng(SEED + ord(s)).permutation(n)
121 + sources[f"{s}half1"] = directions_from(reps, labels, perm[: n // 2])
122 + sources[f"{s}half2"] = directions_from(reps, labels, perm[n // 2:])
123 + boot = np.random.default_rng(200 + ord(s)).integers(0, n, n)
124 + sources[f"boot{s}"] = directions_from(reps, labels, boot)
125 + del sources["bootB"] # keep 6 by hypothesis? A/B halves (4) + bootA + bootB = 6 — keep both
126 + sources["bootB"] = directions_from(reps_cache["B"][0], reps_cache["B"][1],
127 + np.random.default_rng(202).integers(0, len(reps_cache["B"][1]),
128 + len(reps_cache["B"][1])))
129 +
130 + rand_damage = []
131 + d_model = reps_cache["A"][0].shape[-1]
132 + for layer in range(n_layers):
133 + ms = []
134 + for s in range(N_RANDOM_DIRS):
135 + ru = np.random.default_rng(2000 * layer + s).standard_normal(d_model)
136 + ru /= np.linalg.norm(ru)
137 + taps[layer].edit = erase_fn(ru.astype(np.float32), sources["Ahalf1"][1][layer])
138 + ms.append(margin())
139 + taps[layer].edit = None
140 + rand_damage.append(base_m - float(np.mean(ms)))
141 + print("random-direction controls done", flush=True)
142 +
143 + profiles, bands = {}, {}
144 + for name, (dirs, mus) in sources.items():
145 + prof = []
146 + for layer in range(n_layers):
147 + taps[layer].edit = erase_fn(dirs[layer], mus[layer])
148 + m = margin()
149 + taps[layer].edit = None
150 + prof.append((base_m - m) - rand_damage[layer])
151 + profiles[name] = prof
152 + bands[name] = {"early_mean": float(np.mean(prof[EARLY])),
153 + "late_mean": float(np.mean(prof[LATE]))}
154 + print(f"{name:8s} early={bands[name]['early_mean']:+.3f} late={bands[name]['late_mean']:+.3f}",
155 + flush=True)
156 +
157 + early_means = [bands[n]["early_mean"] for n in sources]
158 + passes = bool(all(e >= BAR for e in early_means))
159 + print(f"BAND CLAIM (early mean >= {BAR} in every source): "
160 + f"min={min(early_means):+.3f} -> passes={passes}")
161 +
162 + arr = np.array(list(profiles.values()))
163 + commit = subprocess.run(["git", "rev-parse", "HEAD"], cwd=ROOT,
164 + capture_output=True, text=True, check=False).stdout.strip()
165 + ts = time.strftime("%Y%m%dT%H%M%SZ", time.gmtime())
166 + outdir = ROOT / "results" / "expC_causal_verification" / ts
167 + outdir.mkdir(parents=True)
168 + (outdir / "results.json").write_text(json.dumps({
169 + "experiment": "expC_causal_verification", "run": 4,
170 + "scope": "narrowed early-band claim, fresh sources + fresh behavioral bank",
171 + "commit": commit,
172 + "config": {"model": MODEL, "sources": list(sources), "bar": BAR,
173 + "early_layers": [2, 15], "late_layers": [20, 27],
174 + "n_random_dirs": N_RANDOM_DIRS, "n_pairs": len(pairs), "seed": SEED},
175 + "manifest": manifest(),
176 + "baseline_margin": base_m,
177 + "random_direction_damage": rand_damage,
178 + "profiles": profiles,
179 + "bands": bands,
180 + "profile_mean": arr.mean(axis=0).tolist(),
181 + "profile_min": arr.min(axis=0).tolist(),
182 + "band_claim": {"bar": BAR, "early_means": early_means,
183 + "min_early_mean": float(min(early_means)), "passes": passes},
184 + "wall_seconds": round(time.time() - t0, 1),
185 + }, indent=2) + "\n")
186 + print(f"results -> {outdir / 'results.json'}")
187 + return 0
188 +
189 +
190 +if __name__ == "__main__":
191 + sys.exit(main())
modified experiments/micro/expC_causal_verification/implementation/make_interventions_mapcard.py +55 −52
@@ -2,7 +2,7 @@
2 2 # =============================================================================
3 3 # Project : modelmap
4 4 # File : experiments/micro/expC_causal_verification/implementation/make_interventions_mapcard.py
5 # Purpose : Publish the causal direction-erasure profile as an atlas entry
5 +# Purpose : Publish the BAND-claim interventions map from expC run #4
6 6 # Author : Simon-Pierre Boucher
7 7 # Contact : contact@spboucher.ai
8 8 # Website : https://modelmap.io
@@ -11,9 +11,11 @@
11 11 # Platform : macOS / Apple Silicon (arm64)
12 12 # License : All rights reserved (research code)
13 13 # =============================================================================
14 """Builds atlas/qwen3-0.6b-4bit/interventions/v1 from expC run #3 —
15 REFUSES to publish unless the run's own replication verdict is true
16 (mean pairwise rho >= 0.7 and the band claim holds in every source)."""
14 +"""Builds atlas/qwen3-0.6b-4bit/interventions/v1 from expC run #4 —
15 +REFUSES unless the run's pre-registered band claim passed (early-band mean
16 +specific damage >= bar in EVERY fresh source). Run #3's full-profile
17 +version of this gate refused publication (documented in analysis.md);
18 +this version publishes a BAND claim, the granularity that replicates."""
17 19
18 20 from __future__ import annotations
19 21
@@ -23,8 +25,6 @@ import sys
23 25 import time
24 26 from pathlib import Path
25 27
26 import numpy as np
27
28 28 ROOT = Path(__file__).resolve().parents[4]
29 29 sys.path.insert(0, str(ROOT / "src"))
30 30
@@ -34,12 +34,12 @@ ENTRY = ROOT / "atlas" / "qwen3-0.6b-4bit" / "interventions" / "v1"
34 34 MODEL_ID = "mlx-community/Qwen3-0.6B-4bit"
35 35
36 36
37 def newest_run3() -> Path:
37 +def newest_run4() -> Path:
38 38 for d in sorted((ROOT / "results" / "expC_causal_verification").iterdir(), reverse=True):
39 39 doc = json.loads((d / "results.json").read_text())
40 if doc.get("run") == 3:
40 + if doc.get("run") == 4:
41 41 return d / "results.json"
42 raise SystemExit("no run-3 results found")
42 + raise SystemExit("no run-4 results found")
43 43
44 44
45 45 def model_hash() -> str:
@@ -52,35 +52,36 @@ def model_hash() -> str:
52 52
53 53
54 54 def main() -> int:
55 res_path = newest_run3()
55 + res_path = newest_run4()
56 56 doc = json.loads(res_path.read_text())
57 rep = doc["replication"]
58 if not rep["publishable"]:
59 print(f"REFUSED: replication verdict is false "
60 f"(mean rho {rep['mean_rho']:.3f}, band {rep['band_claim_all']}) — "
61 "per the registered publication rule this profile stays unpublished.")
57 + claim = doc["band_claim"]
58 + if not claim["passes"]:
59 + print(f"REFUSED: band claim failed (min early mean {claim['min_early_mean']:+.3f} "
60 + f"< bar {claim['bar']}) — per the registered rule nothing is published.")
62 61 return 1
63 62
64 63 ENTRY.mkdir(parents=True, exist_ok=True)
65 profiles = doc["profiles"]
66 boots = [k for k in profiles if k.startswith("bootA")]
67 boot_mean = np.mean([profiles[b] for b in boots], axis=0).tolist()
68 64 map_doc = {
69 65 "author": "Simon-Pierre Boucher", "contact": "contact@spboucher.ai",
70 66 "website": "https://modelmap.io",
71 67 "map_type": "interventions", "model_id": MODEL_ID,
72 "claim": "A single diff-of-means agreement direction, erased at any layer in the "
73 "early-mid band, destroys most of the model's grammatical-agreement "
74 "preference; late layers carry little. Per-layer specific damage = "
75 "(agreement-direction erasure damage) minus (random-direction damage).",
76 "behavior_metric": "logit margin correct-vs-incorrect verb on held-out minimal pairs",
68 + "claim": "BAND claim: erasing the diff-of-means agreement direction at any single "
69 + "layer in the early band (2-15) destroys most of the grammatical-agreement "
70 + "margin, replicated across six fresh direction estimates on a fresh "
71 + "behavioral bank. The late band (20-27) is reported but carries NO claim "
72 + "(declared estimator-unstable by run #3).",
73 + "behavior_metric": "logit margin correct-vs-incorrect verb, held-out minimal pairs",
77 74 "baseline_margin": doc["baseline_margin"],
75 + "band": {"early_layers": doc["config"]["early_layers"],
76 + "late_layers": doc["config"]["late_layers"],
77 + "bar": claim["bar"], "early_means_per_source": claim["early_means"],
78 + "min_early_mean": claim["min_early_mean"]},
78 79 "per_layer": {
79 "setA": profiles["setA"], "setB": profiles["setB"],
80 "boot_mean": boot_mean,
80 + "mean": doc["profile_mean"],
81 + "min": doc["profile_min"],
81 82 "random_direction_damage": doc["random_direction_damage"],
82 83 },
83 "replication": rep,
84 + "profiles_per_source": doc["profiles"],
84 85 "source_results": str(res_path.relative_to(ROOT)),
85 86 }
86 87 (ENTRY / "map.json").write_text(json.dumps(map_doc, indent=2) + "\n")
@@ -103,31 +104,31 @@ def main() -> int:
103 104 config=str(res_path.relative_to(ROOT)), created=created,
104 105 hardware_manifest=doc["manifest"], confidence_level=2,
105 106 regenerate_command=(
106 ".venv/bin/python experiments/micro/expC_causal_verification/implementation/benchmark_v3.py && "
107 + ".venv/bin/python experiments/micro/expC_causal_verification/implementation/benchmark_v4.py && "
107 108 ".venv/bin/python experiments/micro/expC_causal_verification/implementation/make_interventions_mapcard.py"),
108 seeds=[doc["config"]["seed"], "bootA0-2 (direction bootstrap)"],
109 prompt_sets=["agreement_A (direction est.)", "agreement_B (direction est.)",
110 "held-out minimal pairs (behavior, disjoint locations)"],
109 + seeds=[doc["config"]["seed"], *doc["config"]["sources"]],
110 + prompt_sets=["agreement_A halves (direction est.)", "agreement_B halves (direction est.)",
111 + "fresh held-out minimal-pair bank (8 unseen locations)"],
111 112 controls=["random-direction erasure per layer (3 dirs, netted out)",
112 "cross-promptset direction replication (A vs B)",
113 "bootstrap direction replication (3 resamples)"],
113 + "six-source replication with pre-registered bar (runs #3-#4)",
114 + "late band excluded as estimator-unstable (run #3 refusal)"],
114 115 methods_in_agreement=["difference-in-means probing (direction exists, decodable)",
115 116 "direction erasure (causally load-bearing)"],
116 117 interventions=["rank-1 direction erasure at each layer's output, all positions"],
117 replication_rate=round(rep["mean_rho"], 4),
118 + replication_rate=round(claim["min_early_mean"] / doc["baseline_margin"], 4),
118 119 featurizer_class="linear (difference-in-means direction)",
119 intervention_protocol="erase h' = h − ⟨h−μ,u⟩u at layer ℓ output; logit-margin metric; "
120 "random-direction null netted out",
120 + intervention_protocol="erase h' = h − ⟨h−μ,u⟩u at layer ℓ; logit-margin metric; "
121 + "random-direction null netted out; BAND granularity",
121 122 notes="Level 2, NOT 3: the two agreeing methods share the diff-of-means estimator; "
122 "an independent intervention family (activation addition/steering) is the "
123 "registered Level-3 path. Layer profile ANTI-correlates with the probes/v2 "
124 "decodability map (expC ledger 0/2) — decodability peaks and causal joints are "
125 "different map types.",
123 + "activation-addition steering is the registered Level-3 path. This entry "
124 + "exists because run #3's per-layer version was REFUSED by the gate — the "
125 + "band is the granularity that replicates. Anti-correlates with probes/v2 "
126 + "layer ranking (survival ledger 0/2): decodability peaks ≠ causal joints.",
126 127 )
127 128 (ENTRY / "mapcard.json").write_text(card.to_json())
128 129
129 early = float(np.mean(profiles["setA"][2:16]))
130 late = float(np.mean(profiles["setA"][20:28]))
130 + per_src = "\n".join(f"- {n}: early-band mean {e:+.3f}"
131 + for n, e in zip(doc["config"]["sources"], claim["early_means"]))
131 132 (ENTRY / "confidence.md").write_text(f"""---
132 133 project: modelmap
133 134 document: qwen3-0.6b-4bit/interventions/v1 — confidence
@@ -142,28 +143,30 @@ status: reviewed
142 143
143 144 ```text
144 145 Level : 2
145 Seeds : direction bootstrap x3 + cross-promptset (A/B)
146 Prompt sets: 3 (two direction-estimation sets + held-out behavioral bank)
146 +Seeds : six fresh direction sources (disjoint halves of two promptsets
147 + + two fresh bootstraps); fresh behavioral bank
148 +Prompt sets: 3 (two estimation sets + held-out behavioral bank)
147 149 Methods in agreement : 2 (diff-of-means probing; direction erasure) — shared
148 150 estimator, hence Level 2 and not 3
149 151 Causal verification : YES — rank-1 erasure with random-direction nulls
150 152 ```
151 153
152 Replication: mean pairwise Spearman ρ = {rep['mean_rho']:.3f} (min {rep['min_rho']:.3f})
153 across 5 direction sources; early-band (layers 215) mean specific damage
154 {early:+.2f} vs late-band (2027) {late:+.2f} on a {doc['baseline_margin']:+.2f} baseline margin.
154 +Pre-registered band claim (bar {claim['bar']} on a {doc['baseline_margin']:+.2f} baseline margin):
155 +{per_src}
156 +Minimum early-band mean across sources: {claim['min_early_mean']:+.3f} — claim PASSES.
155 157
156 The claim is about a DIRECTION, not a place: the same linear feature is
157 causally load-bearing wherever it is erased in the early-mid band. This map
158 is the causal counterpart of probes/v2, whose decodability ranking it
159 contradicts (survival ledger 0/2) — both are published, labeled by what they
160 actually measure.
158 +Granularity discipline: run #3's per-layer profile FAILED replication and was
159 +refused by this very gate; the published object is the BAND (layers 215).
160 +The late band (2027) is displayed but carries no claim. This map is the
161 +causal counterpart of probes/v2, whose decodability ranking it contradicts
162 +(survival ledger 0/2) — both stay published, labeled by what they measure.
161 163 """)
162 164 errs = card.validate()
163 165 if errs:
164 166 print("CARD INVALID:", errs)
165 167 return 1
166 print(f"atlas entry written: {ENTRY.relative_to(ROOT)} (Level 2, rho {rep['mean_rho']:.2f})")
168 + print(f"atlas entry written: {ENTRY.relative_to(ROOT)} "
169 + f"(Level 2, band min {claim['min_early_mean']:+.3f} / bar {claim['bar']})")
167 170 return 0
168 171
169 172
modified research/LOG.md +31 −0
@@ -448,3 +448,34 @@ as unstable). If it passes, the interventions map is published as a BAND
448 448 claim, not a per-layer ranking, at Level 2. Meta-lesson for methodology.md:
449 449 map artifacts must declare their stable granularity — per-layer rankings
450 450 were too fine for this object; bands are the honest resolution.
451 +
452 +---
453 +
454 +## 2026-08-12 09:15 EDT — expC run #4: BAND claim passes — FIRST LEVEL-2 ATLAS ENTRY
455 +
456 +**Result (pre-registered; everything fresh).** CONFIRMED — early-band
457 +(layers 2–15) mean specific damage per source: +3.35 / +3.23 / +3.30 /
458 ++3.26 / +3.33 / +3.27 (min +3.228 vs bar 2.5; spread 0.12 across six
459 +independent direction estimates; fresh bank baseline +4.45). Late band
460 +again unstable (+0.27…+2.83) and carries no claim, as declared.
461 +
462 +**Published.** atlas/qwen3-0.6b-4bit/interventions/v1 — the project's
463 +first causal, Level-2 map: *a single diff-of-means agreement direction,
464 +erased at any one layer in the early band, removes ~73–75% of the model's
465 +grammatical preference, replicated across six fresh estimators, netted
466 +against random-direction damage.* The gate that refused run #3's
467 +per-layer version passed run #4's band version — granularity discipline
468 +enforced by tooling, start to finish. Site: the band map now leads the
469 +home page and the atlas entry renders mean / worst-source / random-
470 +reference curves.
471 +
472 +**The day's arc, as the atlas now shows it:** probes/v1 (negative —
473 +architecture null wins), probes/v2 (Level 1 — differential signal,
474 +per-property verdicts), interventions/v1 (Level 2 — causal band claim),
475 +survival ledger 0/2 for correlational layer rankings, one publication-gate
476 +refusal on record. Every claim at its measured level.
477 +
478 +**Next.** (1) Level-3 path: activation-addition steering (the direction
479 +should *raise* the margin where it is weak); (2) same band protocol on
480 +arith_valid; (3) candidate_02 entry: quantization drift of the band map
481 +(FP16 vs Q8 vs Q4); (4) expG entry: does the band replicate on Qwen3-1.7B?
added results/expC_causal_verification/20260812T071954Z/results.json +369 −0
@@ -0,0 +1,369 @@
1 +{
2 + "experiment": "expC_causal_verification",
3 + "run": 4,
4 + "scope": "narrowed early-band claim, fresh sources + fresh behavioral bank",
5 + "commit": "f7c298d2859ed43654d388618b7274266e7e562b",
6 + "config": {
7 + "model": "mlx-community/Qwen3-0.6B-4bit",
8 + "sources": [
9 + "Ahalf1",
10 + "Ahalf2",
11 + "bootA",
12 + "Bhalf1",
13 + "Bhalf2",
14 + "bootB"
15 + ],
16 + "bar": 2.5,
17 + "early_layers": [
18 + 2,
19 + 15
20 + ],
21 + "late_layers": [
22 + 20,
23 + 27
24 + ],
25 + "n_random_dirs": 3,
26 + "n_pairs": 192,
27 + "seed": 780
28 + },
29 + "manifest": {
30 + "author": "Simon-Pierre Boucher",
31 + "contact": "contact@spboucher.ai",
32 + "website": "https://modelmap.io",
33 + "chip": {
34 + "brand": "Apple M5 Max",
35 + "cores_total": 18,
36 + "cores_performance": 6,
37 + "cores_efficiency": 12
38 + },
39 + "memory": {
40 + "unified_gb": 48.0,
41 + "pagesize": 16384
42 + },
43 + "os": {
44 + "system": "Darwin",
45 + "version": "27.0",
46 + "arch": "arm64"
47 + },
48 + "software": {
49 + "python": "3.14.4",
50 + "numpy": "2.5.2",
51 + "mlx": "0.32.0",
52 + "torch": "2.13.0",
53 + "safetensors": "0.8.0"
54 + }
55 + },
56 + "baseline_margin": 4.450520833333333,
57 + "random_direction_damage": [
58 + -0.0008680555555562464,
59 + -0.0028211805555562464,
60 + 2.340847439236111,
61 + 1.1391059027777772,
62 + 3.474690755208333,
63 + 0.05262586805555536,
64 + 0.4880642361111103,
65 + 0.7284071180555558,
66 + 0.4264322916666661,
67 + 1.4096137152777777,
68 + 0.24273003472222232,
69 + 1.309950086805555,
70 + 0.3861762152777777,
71 + 0.3366970486111107,
72 + 1.3014322916666665,
73 + 0.4308810763888893,
74 + 0.0016276041666660745,
75 + 0.15190972222222232,
76 + 0.04720052083333304,
77 + -0.040364583333333925,
78 + 0.09483506944444375,
79 + -0.05674913194444553,
80 + -0.11707899305555625,
81 + -0.05338541666666696,
82 + 0.028103298611111605,
83 + 0.0866970486111116,
84 + 0.06803385416666607,
85 + -0.0021701388888892836
86 + ],
87 + "profiles": {
88 + "Ahalf1": [
89 + -0.029405381944443754,
90 + -0.005316840277777679,
91 + 2.068657769097222,
92 + 3.288953993055556,
93 + 0.920817057291667,
94 + 4.322129991319445,
95 + 3.8747287326388897,
96 + 3.5920681423611103,
97 + 3.9365234375,
98 + 2.9437391493055554,
99 + 4.149278428819444,
100 + 3.007921006944445,
101 + 3.9253472222222223,
102 + 3.9666883680555554,
103 + 3.0131835937499996,
104 + 3.8489040798611107,
105 + 4.263671875,
106 + 3.9997829861111107,
107 + 3.8670247395833335,
108 + 3.7986653645833335,
109 + 2.9704318576388893,
110 + 2.7823350694444455,
111 + 1.7347547743055558,
112 + 1.9698893229166665,
113 + 1.3794487847222214,
114 + 1.2564019097222214,
115 + 1.6738281250000004,
116 + 0.7954644097222223
117 + ],
118 + "Ahalf2": [
119 + 0.052951388888889284,
120 + 0.040256076388889284,
121 + 2.068901909722222,
122 + 3.2481011284722228,
123 + 0.892659505208333,
124 + 4.256130642361111,
125 + 3.8350151909722228,
126 + 3.5364040798611103,
127 + 3.919596354166667,
128 + 2.7926161024305554,
129 + 3.9950629340277777,
130 + 2.909776475694445,
131 + 3.7218967013888884,
132 + 3.7080620659722223,
133 + 2.6974283854166665,
134 + 3.6133897569444438,
135 + 4.05029296875,
136 + 3.554307725694444,
137 + 3.7395833333333335,
138 + 3.9003906250000004,
139 + 3.2191297743055562,
140 + 2.4511176215277786,
141 + 1.5763888888888893,
142 + -0.569742838541667,
143 + -2.4296332465277786,
144 + -1.4393988715277786,
145 + -0.99951171875,
146 + 0.31629774305555536
147 + ],
148 + "bootA": [
149 + -0.008572048611110716,
150 + 0.018446180555556246,
151 + 2.036593967013889,
152 + 3.259657118055556,
153 + 0.912272135416667,
154 + 4.290757921006945,
155 + 3.8606906467013897,
156 + 3.584581163194444,
157 + 3.877766927083334,
158 + 2.8962131076388884,
159 + 4.020941840277778,
160 + 2.930284288194445,
161 + 3.8477105034722223,
162 + 3.9113498263888893,
163 + 2.9856770833333335,
164 + 3.8200954861111107,
165 + 4.232096354166667,
166 + 3.8371853298611107,
167 + 4.0078125,
168 + 3.9334309895833335,
169 + 3.6519097222222228,
170 + 3.2728949652777786,
171 + 2.6348198784722223,
172 + 2.4124348958333335,
173 + 2.005099826388888,
174 + 1.606336805555555,
175 + 1.8066406250000004,
176 + 0.8556857638888888
177 + ],
178 + "Bhalf1": [
179 + -0.005967881944443754,
180 + 0.015190972222222321,
181 + 2.060845269097222,
182 + 3.237765842013889,
183 + 0.887776692708333,
184 + 4.281765407986111,
185 + 3.8278537326388897,
186 + 3.542344835069444,
187 + 3.875325520833334,
188 + 2.8387586805555554,
189 + 3.9384223090277777,
190 + 2.870795355902778,
191 + 3.8202853732638884,
192 + 3.8525933159722223,
193 + 2.9090169270833335,
194 + 3.7400173611111107,
195 + 4.126627604166667,
196 + 3.7250434027777772,
197 + 3.8818359375,
198 + 3.7548828125000004,
199 + 3.491102430555556,
200 + 3.3700629340277786,
201 + 3.2643771701388893,
202 + 3.141764322916667,
203 + 2.7155490451388884,
204 + 2.4268120659722214,
205 + 2.977213541666667,
206 + 1.2258029513888888
207 + ],
208 + "Bhalf2": [
209 + 0.007703993055556246,
210 + 0.029513888888889284,
211 + 2.032199435763889,
212 + 3.235975477430556,
213 + 0.888956705729167,
214 + 4.313869900173611,
215 + 3.8546481662326397,
216 + 3.623155381944444,
217 + 3.907145182291667,
218 + 2.8999565972222223,
219 + 4.044542100694444,
220 + 2.986273871527778,
221 + 3.9251844618055554,
222 + 3.9627821180555554,
223 + 3.0100911458333335,
224 + 3.8747829861111107,
225 + 4.273600260416667,
226 + 3.9958767361111107,
227 + 4.093912760416667,
228 + 4.025065104166667,
229 + 3.724500868055556,
230 + 3.752712673611112,
231 + 3.6870659722222228,
232 + 3.225260416666667,
233 + 2.6506076388888884,
234 + 2.2886284722222214,
235 + 2.0901692708333335,
236 + 1.0425347222222223
237 + ],
238 + "bootB": [
239 + 0.008680555555556246,
240 + 0.02690972222222232,
241 + 2.048556857638889,
242 + 3.2363823784722228,
243 + 0.884114583333333,
244 + 4.290798611111111,
245 + 3.836927625868056,
246 + 3.556566026475694,
247 + 3.861246744791667,
248 + 2.8323296440972223,
249 + 3.9693467881944438,
250 + 2.893256293402778,
251 + 3.8176812065972223,
252 + 3.8633355034722223,
253 + 2.9313151041666665,
254 + 3.7629665798611107,
255 + 4.178385416666667,
256 + 3.749620225694444,
257 + 3.8727213541666665,
258 + 3.700846354166667,
259 + 3.4992404513888893,
260 + 3.4939236111111116,
261 + 3.5031467013888893,
262 + 3.155436197916667,
263 + 2.6030815972222214,
264 + 2.204318576388888,
265 + 2.313151041666667,
266 + 1.1629774305555558
267 + ]
268 + },
269 + "bands": {
270 + "Ahalf1": {
271 + "early_mean": 3.3470672123015865,
272 + "late_mean": 1.820319281684028
273 + },
274 + "Ahalf2": {
275 + "early_mean": 3.2282172309027777,
276 + "late_mean": 0.2655809190538194
277 + },
278 + "bootA": {
279 + "early_mean": 3.3024708581349214,
280 + "late_mean": 2.280727810329861
281 + },
282 + "Bhalf1": {
283 + "early_mean": 3.263111901661706,
284 + "late_mean": 2.826585557725694
285 + },
286 + "Bhalf2": {
287 + "early_mean": 3.3256831093439985,
288 + "late_mean": 2.8076850043402777
289 + },
290 + "bootB": {
291 + "early_mean": 3.270344567677331,
292 + "late_mean": 2.741909450954861
293 + }
294 + },
295 + "profile_mean": [
296 + 0.0042317708333339255,
297 + 0.02083333333333363,
298 + 2.0526258680555554,
299 + 3.2511393229166665,
300 + 0.89776611328125,
301 + 4.292575412326388,
302 + 3.8483106825086817,
303 + 3.572519938151041,
304 + 3.896267361111111,
305 + 2.8672688802083335,
306 + 4.019599066840278,
307 + 2.9330512152777786,
308 + 3.843017578125,
309 + 3.8774685329861107,
310 + 2.924452039930556,
311 + 3.7766927083333326,
312 + 4.187445746527779,
313 + 3.8103027343749996,
314 + 3.910481770833334,
315 + 3.8522135416666674,
316 + 3.426052517361111,
317 + 3.187174479166668,
318 + 2.733425564236112,
319 + 2.2225070529513893,
320 + 1.4873589409722214,
321 + 1.3905164930555547,
322 + 1.6435818142361114,
323 + 0.8997938368055557
324 + ],
325 + "profile_min": [
326 + -0.029405381944443754,
327 + -0.005316840277777679,
328 + 2.032199435763889,
329 + 3.235975477430556,
330 + 0.884114583333333,
331 + 4.256130642361111,
332 + 3.8278537326388897,
333 + 3.5364040798611103,
334 + 3.861246744791667,
335 + 2.7926161024305554,
336 + 3.9384223090277777,
337 + 2.870795355902778,
338 + 3.7218967013888884,
339 + 3.7080620659722223,
340 + 2.6974283854166665,
341 + 3.6133897569444438,
342 + 4.05029296875,
343 + 3.554307725694444,
344 + 3.7395833333333335,
345 + 3.700846354166667,
346 + 2.9704318576388893,
347 + 2.4511176215277786,
348 + 1.5763888888888893,
349 + -0.569742838541667,
350 + -2.4296332465277786,
351 + -1.4393988715277786,
352 + -0.99951171875,
353 + 0.31629774305555536
354 + ],
355 + "band_claim": {
356 + "bar": 2.5,
357 + "early_means": [
358 + 3.3470672123015865,
359 + 3.2282172309027777,
360 + 3.3024708581349214,
361 + 3.263111901661706,
362 + 3.3256831093439985,
363 + 3.270344567677331
364 + ],
365 + "min_early_mean": 3.2282172309027777,
366 + "passes": true
367 + },
368 + "wall_seconds": 149.9
369 +}
modified site/data/atlas-index.json +11 −0
@@ -3,6 +3,17 @@
3 3 "contact": "contact@spboucher.ai",
4 4 "website": "https://modelmap.io",
5 5 "entries": [
6 + {
7 + "path": "atlas/qwen3-0.6b-4bit/interventions/v1",
8 + "map_id": "atlas/qwen3-0.6b-4bit/interventions/v1",
9 + "map_type": "interventions",
10 + "model_id": "mlx-community/Qwen3-0.6B-4bit",
11 + "quantization": "q4 (mlx)",
12 + "confidence_level": 2,
13 + "replication_rate": 0.7254,
14 + "negative_result": false,
15 + "created": "2026-08-12"
16 + },
6 17 {
7 18 "path": "atlas/qwen3-0.6b-4bit/probes/v1",
8 19 "map_id": "atlas/qwen3-0.6b-4bit/probes/v1",
modified site/server.js +19 −15
@@ -53,28 +53,31 @@ function probeMapCharts(onlyProp, mapRel) {
53 53 const MAP_V2 = "atlas/qwen3-0.6b-4bit/probes/v2/map.json";
54 54 const MAP_INT = "atlas/qwen3-0.6b-4bit/interventions/v1/map.json";
55 55
56 /** Causal direction-erasure profile chart (interventions map). */
56 +/** Causal direction-erasure BAND map chart (interventions map, run #4 format). */
57 57 function interventionsMapCharts(mapRel) {
58 58 const m = C.readJson(mapRel || MAP_INT);
59 if (!m || !m.per_layer) return "";
60 const all = [...m.per_layer.setA, ...m.per_layer.setB, ...m.per_layer.boot_mean];
59 + if (!m || !m.per_layer || !m.per_layer.mean) return "";
60 + const all = [...m.per_layer.mean, ...m.per_layer.min, ...m.per_layer.random_direction_damage];
61 61 const yMin = Math.floor(Math.min(...all, 0) * 2) / 2 - 0.5;
62 62 const yMax = Math.ceil(Math.max(...all) * 2) / 2 + 0.5;
63 + const band = m.band || {};
63 64 return Charts.layerLineSvg({
64 title: "agreement — causal direction-erasure profile (specific damage by layer)",
65 + title: "agreement — causal direction-erasure BAND map (specific damage by layer)",
65 66 yLabel: "specific margin damage",
66 67 yMin, yMax,
67 68 series: [
68 { label: "direction from set A", color: Charts.S1, values: m.per_layer.setA },
69 { label: "direction from set B", color: Charts.S2, values: m.per_layer.setB },
70 { label: "bootstrap mean (×3)", color: Charts.REF, dash: true, ref: true,
71 values: m.per_layer.boot_mean },
69 + { label: "mean of 6 fresh sources", color: Charts.S1, values: m.per_layer.mean },
70 + { label: "worst source (min)", color: Charts.S2, values: m.per_layer.min },
71 + { label: "random-direction damage", color: Charts.REF, dash: true, ref: true,
72 + values: m.per_layer.random_direction_damage },
72 73 ],
73 caption: "Erasing the diff-of-means agreement direction at any single early-mid layer " +
74 "destroys most of the grammatical-agreement margin (baseline " +
75 (m.baseline_margin ? m.baseline_margin.toFixed(2) : "—") + "); random-direction damage is " +
76 "netted out. Level 2 — causally verified, replicated across direction estimates. " +
77 "This profile ANTI-correlates with the probes/v2 decodability ranking (survival ledger 0/2).",
74 + caption: `The published claim is the BAND (layers ${band.early_layers ? band.early_layers.join("–") : "2–15"}): ` +
75 + "erasing the diff-of-means agreement direction at any single early-band layer destroys most of " +
76 + "the grammatical margin (baseline " + (m.baseline_margin ? m.baseline_margin.toFixed(2) : "—") +
77 + "), replicated across six fresh direction estimates on a fresh behavioral bank (Level 2). " +
78 + "The late band is displayed but carries no claim — run #3's per-layer version was refused by " +
79 + "the publication gate as estimator-unstable. Anti-correlates with the probes/v2 decodability " +
80 + "ranking (survival ledger 0/2): decodability peaks ≠ causal joints.",
78 81 });
79 82 }
80 83
@@ -212,15 +215,16 @@ app.get("/", (req, res) => {
212 215 </section>
213 216
214 217 ${(() => {
218 + const causal = interventionsMapCharts();
215 219 const probe2 = probeMapCharts("agreement", MAP_V2);
216 220 const probe1 = probeMapCharts("lang_id");
217 221 const exph = expHCharts();
218 return probe1 || probe2 || exph ? `<section class="card">
222 + return causal || probe1 || probe2 || exph ? `<section class="card">
219 223 <h2>First measured maps</h2>
220 224 <p class="lede-small">Every figure below is regenerated from committed code + versioned results —
221 225 raw JSON with hardware manifests under <a href="/results">Results</a>. Full per-property maps live
222 226 in the <a href="/atlas">Atlas</a>.</p>
223 ${probe2}${probe1}${exph}
227 + ${causal}${probe2}${probe1}${exph}
224 228 </section>` : "";
225 229 })()}
226 230
227 231