spb/anomaly-atlas Public License
Systematic discovery & rigorous validation of statistical anomalies in open HF market data (hfmarketdata.io) — pre-registered, artifact-null-driven, fully reproducible. Live atlas: www.anomaly-atlas.io
Python 61.4%
JavaScript 28.7%
CSS 8.6%
Shell 0.7%
Makefile 0.5%
1---2project: anomaly-atlas3document: P001 — The Artifact Frontier, Part I4author: Simon-Pierre Boucher5contact: contact@spboucher.ai6data_source: hfmarketdata.io7created: 2026-08-128status: published9pub_id: P00110version: 111title: "The Artifact Frontier: What Survives Honest Testing in Open High-Frequency Market Data?"12subtitle: "Part I — Methods, artifact taxonomy, and train-period results (2000–2016)"13abstract: "We systematically scan open 1-minute-to-daily market data (hfmarketdata.io, sole source) for short-horizon mean-reversion, lead-lag, and calendar anomalies, under a pre-registered protocol: every detector must first pass a synthetic-data gate; every scan runs against measured artifact nulls; hypothesis budgets are declared before testing; and validation layers (multiple-testing correction, transaction costs) are applied in sequence. On the 2000–2016 train split, 62% of the 372 searched rules are naively 'significant' and 18% survive Hansen's SPA — yet the survivors carry physically implausible paper Sharpes, and a deliberately included known artifact (the SPX→SPY 'lead') survives statistical correction unharmed. The cost layer then eliminates essentially everything: the median surviving rule breaks even at 1.1% of one half-spread per trade, and zero intraday rules survive paying the full half-spread. The calendar family — including turn-of-month, the last published survivor — dies against a permuted-calendar null with an 8-test budget. Our principal positive contributions are a measured artifact taxonomy for this dataset and a demonstrated three-layer validation doctrine: artifact nulls, search correction, and costs are independent filters, and no one of them substitutes for another. Out-of-sample confirmation on the untouched validation split is reported in Part II."14---1516# The Artifact Frontier — Part I1718*Every number in this document regenerates from committed `results.json`19files, and every figure below is rendered live from them by the site. The20[research log](/doc/research/LOG.md) is the audit trail; the21[charter](/doc/CLAUDE.md) is the pre-registered protocol.*2223## 1. Question and doctrine2425Given only open high-frequency market data — 1-minute to daily bars, no26quotes — which statistical regularities are *real* (reproducible27out-of-sample, robust to artifacts, economically nonzero after costs), and28which are plumbing? The project's doctrine, stated before any data was29touched: **every candidate anomaly is an artifact until proven otherwise;30in-sample results are never findings; negative results are first-class.**3132Three design rules operationalize this:33341. **Test the tests.** No detector touches real data before passing a35 synthetic gate: it must find *nothing* in a pure random walk, recover36 planted effects, and flag a pure bid-ask-bounce series as artifact.37 The gate is not ceremonial — it caught two real bugs before they could38 contaminate results (a variance-ratio estimator reading ~1/q on random39 walks, and an end-freshness-only synchronization mask that silently40 attenuated correlations via the Epps mechanism).412. **Pre-registration.** Universes, time splits (train 2000–2016 /42 validation 2016–2021 / sealed holdout 2022→), hypothesis budgets (22,43 declared in [research_gaps](/doc/research/research_gaps.md)), and every44 experiment's falsification criterion were frozen in writing before the45 corresponding run.463. **A frozen cache as the reproducibility anchor.** All data flows through47 one client that never silently refetches; every response is indexed in a48 committed data manifest. Two of the seven experiments below ran with49 **zero network requests**, entirely from the frozen cache.5051## 2. The data, measured5253Experiment A established empirically what the source actually provides:54true 1-minute bars for ~7,700 stocks and ~5,200 ETFs from January 200055(futures and indices from 2008, FX from 2010, crypto from 2013), plus daily56options chains with quotes and Greeks over 67 quarters. Load-bearing facts57that the documentation does not state: timestamps are US-Eastern wall-clock,58bar-start labeled; **bars exist only where trades occurred** (no zero-volume59placeholders — an illiquid name printed 38 bars in a full session); daily60bars carry the official auction close and consolidated volume, both absent61from the 1-minute series; and dividend-adjusted prices are re-based to the62vendor's build date, so adjusted series are not point-in-time stable.63Responses are hard-capped at 50,000 rows. Full profile:64[data_source_profile](/doc/research/data_source_profile.md).6566## 3. The artifact taxonomy, with magnitudes6768The project's first deliverable is a catalogue of the mechanisms in *this69dataset* that manufacture fake anomalies — each with detection code and a70measured magnitude ([artifact_taxonomy](/doc/research/artifact_taxonomy.md)).71Experiment B measured the key ones on a pre-specified 42-ticker universe72(Q1 2024, RTH 1-minute):7374{{figure:expB_artifact_baselines}}7576Highlights: bid-ask bounce alone produces AC1 of −0.23 and VR(30) of 0.5577in the stalest liquidity tercile with zero planted economics; LOCF joins78make SPY spuriously "lead" mid-staleness names (+0.047 at +1 min, Spearman79vs staleness +0.43) while *diluting* the lead of ultra-stale names — the80artifact is non-monotone; and the SPX index print lags SPY by one minute81(+0.065 at 0.965 contemporaneous correlation) — a Fisher (1966) effect,82measured live. The intraday profile (Experiment E) adds the time-of-day83dimension: volatility is U-shaped (6.6 bp at the open, 2.4 midday, 2.9 at84the close) while the effective spread declines monotonically (2.8 → 1.2 bp)85— any "first-30-minutes" return claim fights 2–3× the midday artifact86level.8788## 4. The scans (train split only, Level 0 by construction)8990**Mean-reversion (Experiment C).** 127 cells (ticker × timeframe ×91sub-period), each tested against a bounce null and FDR-corrected. The92scan's most valuable output was about the null itself: a daily effective93spread combined with pure Roll alternation predicts *impossible* intraday94autocorrelations (−3 to −27), because consecutive intraday closes rarely95flip sides. The corrected, variance-consistent triage — an MA(1) null that96absorbs *all* lag-1 effects — leaves 14 cells of genuine multi-lag97reversion, concentrated in a daily 2008–2015 mega-cap/index family98(XOM excess AC1 −0.13, SPY −0.055, both FDR) and a few 1-minute cells99(JPM −0.24 VR-excess at the 30-minute horizon).100101{{figure:expC_reversion_scan}}102103**Lead-lag (Experiment D).** 49 pairs, two windows, with the raw-LOCF104versus both-fresh comparison built in — so the non-synchronicity artifact105is *measured*, not consumed. In 2006–2007, minute-scale market→component106diffusion was real on synchronized samples (SPY led every sector ETF by107+0.07..+0.15). By 2014–2015 it had collapsed to ±0.05 — the cleanest decay108measurement of the project. Two structural facts survive synchronization:109a splice-invariant ES↔SPY cross-serial effect (−0.032, identical across110all three futures adjustments), and the SPX→SPY "lead" (+0.132) — which111survives *because synchronizing print times cannot fix a computed index*.112113{{figure:expD_leadlag_scan}}114115**Calendar (Experiment E).** Eight pre-declared tests (day-of-week ×5,116turn-of-month, pre/post-holiday) on SPY against a within-year117permuted-calendar null with a family-wise max-statistic. **Nothing118survives** (best marginal p = 0.24; family-wise p ≥ 0.93 everywhere).119Turn-of-month — the last survivor in the published literature as of 2006 —120fails and decays inside the train period (+7.9 bp in 2000–2007 → +1.6 bp121in 2008–2015). Both pipeline controls behaved: the Monday effect stayed122dead, and the volatility U-shape was strongly present.123124{{figure:expE_calendar_scan}}125126## 5. The survival curve: statistical correction is not artifact correction127128Experiment F pushed *everything the scans searched* — 372 signed rules —129through the correction battery: naive t-tests, Benjamini–Hochberg FDR,130White's Reality Check and Hansen's SPA over stationary bootstraps, and the131Deflated Sharpe Ratio.132133{{figure:expF_multiple_testing}}134135The result that matters is not the 18% SPA survival rate — it is *what*136survives: rules with paper Sharpes of 10–31 annualized, physically137implausible, dominated by bounce harvesting (a contrarian rule mechanically138earns −autocov₁ on paper, which is precisely the spread it would pay in139reality). The canary proves the point: the SPX→SPY rule — an artifact we140had already measured twice — **passes SPA comfortably**. Statistical141correction corrects for *search*; it is structurally blind to *mechanism*.142143## 6. The cost frontier closes the loop144145Experiment G re-priced the double-filtered pool (rules that beat both the146artifact nulls and the search correction; 31 rules) under a declared cost147model: net = gross − κ · (EDGE half-spread) · turnover, κ swept from 0 to 2.148149{{figure:expG_cost_frontier}}150151The median rule breaks even at **κ\* = 0.011** — it captures about 1% of152one half-spread per trade. Three rules survive κ = 0.1 (all in sparse153names), one survives κ = 0.25 (CKX, an ultra-sparse name with a wide, noisy154spread estimate — the classic profile of an estimation artifact, forwarded155to the validation split with a skeptical prior rather than discarded by156hand), and **zero rules — none — survive paying the full half-spread.**157The pre-registered falsification clause ("the costs-kill story fails if any158intraday rule survives κ = 1") did not trigger.159160## 7. What Part I establishes1611621. **A three-layer validation doctrine, demonstrated rather than asserted.**163 Artifact nulls, search correction, and transaction costs filter164 *different* failure modes; each layer passed things the next one killed.1652. **A measured artifact taxonomy (T1–T7)** for open bar data, with the166 magnitudes above and neutralization rules, validated on synthetic ground167 truth.1683. **Negative results with teeth**: the calendar family is empty under an169 honest budget; minute-scale lead-lag decayed an order of magnitude170 between 2006 and 2015; and nothing in the searched universe pays for its171 own spread on the train split.1724. **Methodological findings**: spread-based bounce nulls must be173 variance-consistent with the target series; synchronization on print174 times cannot de-artifact computed indices; LOCF distortion of lead-lag175 is non-monotone in staleness.176177## 8. Limitations178179All Part-I numbers are train-split, in-sample by design — their180out-of-sample fate on the untouched 2016–2021 validation split is Part II.181Bar data carries no quotes: costs are estimated (EDGE), not observed.182Holiday-class calendar tests are low-powered (n = 144). The universe is183U.S.-equity-centric; crypto and FX hypotheses await volume-semantics184verification. Capacity is out of scope.185186## 9. Reproducibility187188Everything regenerates from the repository: commit `f59e891` (data profile,189client), `9b6beea` (artifact baselines), `c4977f2` (reversion scan),190`a7f66cb` (lead-lag scan), `b228862` (calendar scan), `7a82cc1` (survival191battery), `881399d` (cost frontier). Each experiment's `results.json`192embeds the hardware manifest and the client's instrumentation; the data193manifest (`data_manifest/index.jsonl`) indexes every API response consumed.194Detector gates: `benchmarks/synthetic/` (29 tests at the time of writing).195196## References197198Key sources (full annotated list with DOIs:199[bibliography](/doc/research/bibliography.md)): Roll (1984); Fisher (1966);200Scholes & Williams (1977); Epps (1979); Lo & MacKinlay (1988, 1990);201Sullivan, Timmermann & White (2001); White (2000); Hansen (2005); Bailey &202López de Prado (2014); Harvey, Liu & Zhu (2016); McLean & Pontiff (2016);203Novy-Marx & Velikov (2016); Chordia, Roll & Subrahmanyam (2005); Ardia,204Guidotti & Kroencke (2024); Chen & Velikov (2022).205