SPB Git

spb/anomaly-atlas Public License

Systematic discovery & rigorous validation of statistical anomalies in open HF market data (hfmarketdata.io) — pre-registered, artifact-null-driven, fully reproducible. Live atlas: www.anomaly-atlas.io

Python 61.4% JavaScript 28.7% CSS 8.6% Shell 0.7% Makefile 0.5%

README: English rewrite with metric badges, headline results, author info

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Simon-Pierre Boucher committed 3 h ago (Aug 12, 2026) parent 409ef07

Showing 1 changed file with +82 and −38

modified README.md +82 −38
@@ -5,61 +5,105 @@ author: Simon-Pierre Boucher
5 5 contact: contact@spboucher.ai
6 6 data_source: hfmarketdata.io
7 7 created: 2026-08-12
8 status: draft
8 +modified: 2026-08-12
9 +status: reviewed
9 10 ---
10 11
11 12 # anomaly-atlas
12 13
14 +<p>
15 + <a href="https://www.anomaly-atlas.io"><img alt="Live atlas" src="https://img.shields.io/badge/atlas-anomaly--atlas.io-2456c4"></a>
16 + <img alt="Experiments" src="https://img.shields.io/badge/micro--experiments-8%2F8%20complete-1e7a4d">
17 + <img alt="Synthetic gate" src="https://img.shields.io/badge/synthetic%20gate-29%20tests%20passing-1e7a4d">
18 + <img alt="Findings" src="https://img.shields.io/badge/atlas%20findings-2%20%C3%97%20Level%202-2456c4">
19 + <img alt="Hypothesis budget" src="https://img.shields.io/badge/hypothesis%20budget-22%20(declared)-8a5a1e">
20 + <img alt="Data" src="https://img.shields.io/badge/data-hfmarketdata.io%20(sole%20source)-4a3aa7">
21 + <img alt="Rows analyzed" src="https://img.shields.io/badge/bars%20analyzed-~68M-5d6167">
22 + <img alt="Platform" src="https://img.shields.io/badge/platform-macOS%20%2F%20Apple%20Silicon-1a1c20">
23 + <img alt="License" src="https://img.shields.io/badge/license-all%20rights%20reserved-8b1e1e">
24 +</p>
25 +
13 26 **Which statistical regularities in open high-frequency market data are real —
14 surviving out-of-sample testing, multiple-comparison correction, transaction
15 costs and robustness checks — and which are artifacts (bid-ask bounce, stale
16 timestamps, survivorship, look-ahead, microstructure noise)?**
27 +and which are artifacts?** A systematic, pre-registered, fully reproducible
28 +research project that scans 1-minute-to-daily bars (equities, ETFs, futures,
29 +indices, FX, crypto, options chains) for mean-reversion, lead-lag, and
30 +calendar anomalies, then pushes every candidate through a three-layer
31 +validation ladder: **measured artifact nulls → multiple-testing correction →
32 +transaction costs → out-of-sample confirmation**.
17 33
18 This repository is a living research project (charter: [CLAUDE.md](CLAUDE.md)).
19 Data source: **[hfmarketdata.io](https://www.hfmarketdata.io) — the sole permitted
20 source**; every result is reproducible from it alone.
34 +> **Honesty doctrine.** Every candidate anomaly is an artifact until proven
35 +> otherwise. In-sample results are never findings. Negative results are
36 +> first-class. Nothing here is investment advice or a trading system.
21 37
22 The distinction kept sharp everywhere:
38 +## Headline results (train 2000–2016 → validation 2016–2021)
23 39
24 ```text
25 statistically detectable in-sample
26 ≠ reproducible out-of-sample
27 ≠ robust to artifacts and specification choices
28 ≠ economically meaningful after costs
29 ```
40 +| Validation layer | Survivors |
41 +|---|---|
42 +| Searched rule universe (2 signs × every scanned cell/pair/class) | 372 |
43 +| Naive \|t\| > 1.96 | 232 (62 %) |
44 +| Benjamini–Hochberg FDR 5 % | 226 (61 %) |
45 +| Hansen SPA (data-snooping correction) | 68 (18 %) — *gross, artifact-laden* |
46 +| EDGE cost model, full half-spread per trade | **0** |
47 +| Out-of-sample (validation split, opened once) | **0** — the negative replicates |
48 +
49 +The SPA survivors carried paper Sharpes of 10–31 — bounce harvesting, not
50 +economics — and a deliberately included known artifact (the SPX→SPY "lead")
51 +passed statistical correction unharmed: **statistical correction corrects
52 +for search, not for mechanism.** Median breakeven cost: the surviving rules
53 +capture ~1 % of one half-spread per trade. Meanwhile the artifacts
54 +themselves replicate out-of-sample perfectly.
30 55
31 - **Author:** Simon-Pierre Boucher — <contact@spboucher.ai>
32 - **Primary platform:** Apple Silicon Mac (DuckDB + parquet out-of-core), macOS 14+
33 - **Status:** bootstrap. Next: Phase 0.5 (Experiment A — data reality check). See `research/LOG.md`.
56 +**First atlas entries (Level 2 — corrected, OOS-confirmed, robust):**
34 57
35 ## Honesty doctrine (charter §2.1)
58 +- **F001***Nothing in the searched universe survives the full ladder* (negative finding, the project's headline).
59 +- **F002***The SPX→SPY minute-scale "lead" is index content staleness* (Fisher 1966, measured live; survives print synchronization AND SPA).
36 60
37 Every candidate anomaly is an **artifact until proven otherwise**. In-sample
38 results are never findings. Nothing here is trading advice, a trading system,
39 or a claim of arbitrage; past statistical regularity does not imply future
40 returns. Negative results are first-class published findings.
61 +Full write-up: [**P001 — The Artifact Frontier, Part I**](https://www.anomaly-atlas.io/publications/P001_artifact_frontier_part1)
62 +(every figure regenerates live from committed `results.json`).
41 63
42 ## Layout
64 +## What's in the box
43 65
44 66 | Path | Contents |
45 67 |---|---|
46 | `research/` | Paper trail: log, data-source profile, state of the art, artifact taxonomy, gaps, ranking, methodology, bibliography |
47 | `src/anomaly_atlas/` | Library: `data/` (single `hf_client`, cache, universe, calendars, cleaning), `stats/`, `validation/`, `atlas/`, `viz/`, `instrumentation/` |
48 | `experiments/` | Micro-experiments A–H and candidate prototypes (hypothesis → falsification criterion → analysis) |
49 | `benchmarks/` | Hardware manifest + synthetic series with known properties ("test the tests") |
50 | `atlas/` | Confidence-labeled findings (Level 0–3) with provenance, via `tools/new_finding.py` |
51 | `results/` | Raw runs, reproducible from commit + config + data-manifest + seed + manifest |
52 | `web/` | The public atlas platform (the charter's `site/`, implemented as in localvm-research) — deployed at www.anomaly-atlas.io |
53 | `tools/` | `check_headers.py`, `new_experiment.py`, `new_finding.py` |
68 +| `research/` | Pre-registered charter artifacts: [data-source profile](research/data_source_profile.md), [artifact taxonomy T1–T7 with measured magnitudes](research/artifact_taxonomy.md), [52-source verified bibliography](research/bibliography.md), [22-hypothesis budget](research/research_gaps.md), append-only [LOG](research/LOG.md), [publications](research/publications/) |
69 +| `src/anomaly_atlas/` | The library: single cache-first API client (never silently refetches; committed data manifest), gated statistics (VR, AC1, lead-lag, block/stationary bootstrap, BH-FDR, White RC, Hansen SPA, DSR), artifact detectors (Roll, EDGE, staleness, LOCF) |
70 +| `benchmarks/synthetic/` | **The §8.1 gate** — 29 tests on series with known properties; no detector touches real data before passing (it caught 2 real bugs) |
71 +| `experiments/micro/` | expA–expH, each with pre-registered `hypothesis.md` (falsification criterion + artifact nulls) and `analysis.md` |
72 +| `atlas/` | Confidence-labeled findings (Level 0–3) with full provenance, created only via `tools/new_finding.py` |
73 +| `web/` | The public platform (Node/Express, server-rendered SVG figures from results JSON, mobile-first, comments) |
54 74
55 ## Confidence taxonomy
75 +## Methodology in one paragraph
76 +
77 +Universes, time splits (train / validation / **sealed holdout 2022→**), and
78 +the 22-hypothesis budget were frozen in writing before any scan. Every
79 +detector passes a synthetic gate first (random walk → nothing; planted
80 +effects → recovered; pure bounce → flagged artifact). Scans report effects
81 +*net of measured artifact nulls* (variance-consistent bounce null, both-fresh
82 +synchronization, permuted calendar). Survivors face White RC / Hansen SPA
83 +over the full searched universe, then an EDGE-spread cost sweep, then the
84 +validation split — opened exactly once. All data flows through one frozen
85 +cache indexed by a committed manifest; two experiments ran with **zero
86 +network requests**.
87 +
88 +## Reproduce
89 +
90 +```bash
91 +make setup # venv + deps (macOS / Apple Silicon)
92 +make test # 29 synthetic-gate + unit tests
93 +make headers # author-header compliance
94 +python experiments/micro/expA_data_reality/benchmark.py # then B..H in order
95 +```
56 96
57 `Level 0` in-sample only (never published) → `Level 1` corrected & OOS →
58 `Level 2` robust → `Level 3` cost-real & confirmed on the once-touched holdout.
97 +Every result JSON embeds the hardware manifest, client instrumentation, and
98 +attribution; every finding cites its commits and the SHA-256 of the data
99 +manifest.
59 100
60 ## Reproducibility
101 +## Author
61 102
62 Every finding: commit hash + config + data-manifest index + seed + hardware
63 manifest (`benchmarks/hardware_manifest.py`). No result from an uncommitted tree.
103 +**Simon-Pierre Boucher** — <contact@spboucher.ai>
104 +Data source: [hfmarketdata.io](https://www.hfmarketdata.io) (sole source) ·
105 +Live atlas: [www.anomaly-atlas.io](https://www.anomaly-atlas.io)
64 106
65 *Research on statistical properties of market data — not investment advice.*
107 +*Research on statistical properties of market data. Not investment advice,
108 +not a trading system; past statistical regularity does not imply future
109 +returns. All rights reserved.*
66 110