SPB Git

spb/anomaly-atlas Public License

Systematic discovery & rigorous validation of statistical anomalies in open HF market data (hfmarketdata.io) — pre-registered, artifact-null-driven, fully reproducible. Live atlas: www.anomaly-atlas.io

Python 61.4% JavaScript 28.7% CSS 8.6% Shell 0.7% Makefile 0.5%
5.5 KB

# project: anomaly-atlas document: Candidate ranking author: Simon-Pierre Boucher contact: contact@spboucher.ai data_source: hfmarketdata.io created: 2026-08-12 modified: 2026-08-12 status: reviewed

# Candidate ranking (Phase 4)

All 22 hypotheses from research_gaps.md, scored 1–10 per charter-§7 axis (10 = best). Axes: T testability on hfmarketdata.io · N cleanliness of the artifact null · O OOS feasibility · M multiple-testing discipline · S survival odds after costs (5 = n/a, structural claim) · V novelty of the honest test · R reproducibility · P public-atlas value · I implementation simplicity · D safety against self-deception.

H Hypothesis (short) T N O M S V R P I D Σ
H15 Intraday momentum post-2018 decay 9 7 9 9 6 9 9 10 7 8 83
H11 SPX→SPY lead = staleness artifact 9 9 8 9 5 8 9 9 8 9 83
H13 Turn-of-month survived 2006–2026 9 8 9 8 6 8 9 9 8 8 82
H03 Index VR residual = Fisher share 9 9 8 9 5 7 9 8 8 9 81
H16 Overnight premium convention share 9 9 8 9 5 7 9 8 8 9 81
H20 Intraday artifact U-profile (input) 9 9 8 9 5 6 9 8 8 10 81
H14 Monday effect stays dead (control) 9 9 9 9 5 4 9 7 9 10 80
H09 ES→SPY ordering at the 1min floor 8 8 8 9 5 8 9 8 7 8 78
H10 Crypto weekend → Monday equity gap 8 7 7 9 5 9 8 9 7 7 76
H21 Vol persistence (positive control) 9 9 9 9 5 3 9 6 7 10 76
H01 Liquid 1–60min reversion = 0 net 9 9 8 8 2 4 9 7 8 9 73
H08 SPY→sector ETFs zero on fresh pairs 9 8 8 8 3 5 9 6 8 9 73
H07 Large→small decay + artifact share 8 8 8 7 3 6 8 8 6 8 70
H05 Crypto 1min reversion by maturity 8 7 7 7 4 7 8 7 7 7 69
H17 Crypto hour-of-week calendar 8 8 8 5 4 7 8 7 7 5 67
H02 Daily reversal decay curve 8 7 8 7 2 5 8 7 6 7 65
H18 Holiday effects are dead 7 8 7 7 4 4 8 5 8 7 65
H12 Options activity → next-day vol 7 6 8 7 5 7 7 7 4 6 64
H22 OpEx-week volume/vol patterns 7 6 7 6 4 7 8 7 6 6 64
H19 DST-transition distortions 7 7 7 7 4 6 8 5 6 5 62
H04 Post-jump 1min overreaction 7 5 7 6 3 7 7 7 5 5 59
H06 FX session-boundary reversion 7 6 7 5 3 6 8 5 6 6 59

# Selected for prototyping (charter: 3–5)

Candidate 01 — H15, intraday momentum post-publication decay (Σ83). The single highest-information test: a famous published effect (Gao et al. 2018) with a freezable specification, an untouched 2018–2026 out-of-sample window that only time could create, a mechanism proxy available in our options chains (Baltussen et al.), moderate cost exposure (2 trades/day in SPY), and near-zero researcher degrees of freedom. Every outcome — persistence, decay, or reversal — is a publishable McLean–Pontiff-style data point.

Candidate 02 — H13, turn-of-month 2006–2026 (Σ82). The last calendar survivor per Marquering et al. (2006), untested since on open data; low-frequency (12×/yr → real cost-survival chance), one pre-registered window, permuted-calendar null + SPA. Highest chance in the whole list of a genuine Level-3 finding — and a clean negative if dead.

Candidate 03 — H11 + H03, the index-artifact decomposition (Σ83/81). One machinery, two published deliverables: prove the SPX→SPY 1min "lead" is pure print-staleness (result-type C/D: artifact demonstration), and quantify the Fisher share of index-level variance-ratio momentum. Cleanest nulls of the entire list; already seeded by expB (+0.065 measured).

Candidate 04 — H10, crypto-weekend → Monday equity gap (Σ76, novelty pick). The one genuinely NEW question this dataset is structurally positioned to ask (24/7 crypto × equity session opens). Pre-registered exposed set {COIN, MSTR, RIOT, MARA, HUT} + placebo set + SPY-gap control + permuted-weekend null. Higher self-deception risk (selection, short joint history ~400 weekends) — ranked last of the selected four accordingly.

Reserve — H09 (ES→SPY at the 1min floor): promoted if a selected candidate dies early at the data layer.

# Not selected (why, in one line each)

Controls/inputs H14, H20, H21 run inside the micro-experiment pipeline (expE, expB/expG, framework) — they are mandatory, not candidates. H16 (strong score) folds into candidate 03's convention machinery as a shared deliverable. H01/H08 are expected-negatives that expC/expD produce as scan output without candidate-level investment. H07/H02 are decay re-measurements scheduled after the scans (reuse candidate-03 machinery). H05/H17 (crypto) wait for volume-semantics verification (expA follow-up noted in the data profile). H12/H22 need the options ingestion layer (heavier implementation). H04/H06/H19 have the weakest nulls or lowest power — revisit only if the scans surface something.

Multiple-testing accounting: selection made on priors and design cleanliness, not on any real-data effect sizes — no data peeking occurred beyond expA/expB artifact levels. The 22-hypothesis budget stands for expF.