project: anomaly-atlas document: State of the art author: Simon-Pierre Boucher contact: contact@spboucher.ai data_source: hfmarketdata.io created: 2026-08-12 modified: 2026-08-12 status: reviewed
State of the art (Phase 2)
Critical map of each anomaly family and method, in the charter §5 template.
Sources: research/bibliography.md (52 verified references, accessed
2026-08-12); dataset facts: research/data_source_profile.md; measured
artifact levels: research/artifact_taxonomy.md (T1–T7).
Epistemic status legend: robust / decayed / disputed / likely-artifact.
A1 — Short-horizon individual-stock reversal
- What it claims. Individual stock returns revert at daily–monthly horizons (Lehmann 1990 weekly; Jegadeesh 1990 monthly).
- Granularity required. Daily suffices; our 1min adds the ability to separate close-convention effects. ✅ testable.
- Known artifact confounds. Bid-ask bounce (T1 — Blume–Stambaugh showed it halves such effects), stale prices (T2), auction-close mismatch (T4).
- Decay evidence. McLean–Pontiff −58 % post-publication; Chordia et al. 2014 attenuation with liquidity. Largely gone in liquid U.S. names.
- Correct test. Cross-sectional reversal portfolios with bounce-robust prices + VR/AC1 net of the liquidity-bucket bounce null; block-bootstrap CIs (assumes stationarity within blocks).
- Multiple-testing exposure. Moderate: horizon × universe × weighting grid. Pre-specify or FDR-correct.
- Cost sensitivity. Extreme — highest-turnover class; Novy-Marx–Velikov prior: dies net of costs.
- Open-source impl. Our own
stats/reversion.py(gated §8.1); arch'sVarianceRatioas cross-check — both build on macOS arm64. ✅ - Main limitation. Without quote data, bounce correction is estimated, not measured.
- Honest new test here. A 2000–2026 decay curve of daily reversal net of the measured bounce null, by liquidity bucket — a decay re-measurement, not a discovery claim. Status: decayed (gross); likely-artifact (net).
A2 — Index/portfolio variance-ratio momentum
- Claims. Weekly index returns positively autocorrelated, VR(q) > 1 (Lo–MacKinlay 1988).
- Granularity. Daily/weekly from our 1day bars (2000→) and intradaily aggregation. ✅
- Confounds. Fisher stale-constituent effect (T2/T3) inflated early index autocorrelation; largely gone in ETF prices (SPY trades fresh).
- Decay. The classic effect faded post-1990s; on ETFs (traded prices, not stale indices) it was always weaker.
- Correct test. Lo–MacKinlay VR with heteroskedasticity-robust CIs / block bootstrap; on BOTH the index (SPX) and the ETF (SPY) — divergence measures the Fisher artifact directly.
- MT exposure. Low if q-grid pre-specified (q ∈ {2,5,10,30}).
- Cost sensitivity. n/a as stated (it's a statistical property claim).
- Impl. Ours + arch. ✅
- Limitation. Regime breaks (2008, 2020) dominate long windows — Bai–Perron sub-periods mandatory.
- Honest new test. SPX-vs-SPY VR divergence as a quantified Fisher artifact 2008–2026 — methodological contribution. Status: decayed; index-level residual = likely-artifact.
A3 — Lead-lag: large caps → small caps
- Claims. Returns of large stocks lead small stocks (Lo–MacKinlay 1990); the source of "contrarian" profits.
- Granularity. Daily and 1min both usable. ✅
- Confounds. Non-synchronous trading (T3) — THE canonical confound (Scholes–Williams); our expB measured SPY spuriously leading stale names +0.047 at 1min.
- Decay. Chordia–Roll–Subrahmanyam: minute-scale predictability arbitraged within 5–60 min by 2005; expect near-zero today in fresh pairs.
- Correct test. Lagged cross-correlation/Granger ONLY on both-fresh subsamples, against the staleness-matched null (expB machinery); Epps-aware at 1min.
- MT exposure. High (pairs explosion) — pre-specify a small pair set.
- Cost sensitivity. Extreme for any tradable interpretation.
- Impl. Ours (
stats/leadlag.py, gated). ✅ - Limitation. No trade timestamps within the bar; sub-minute lead-lag invisible.
- Honest new test. Decay curve of large→small lead-lag 2000–2026 net of the staleness null — with the artifact share reported alongside the total. Status: decayed (fresh pairs); the textbook effect is largely T3 artifact in modern data.
A4 — Futures/ETF/index lead-lag (price-discovery ordering)
- Claims. Futures (ES) lead cash ETFs (SPY) which lead the index print (SPX) at minute scale.
- Granularity. 1min is coarse for this (the true lead is seconds) but the ordering may still be detectable. ⚠️ marginal.
- Confounds. T3/T7 (session semantics, index staleness — expB measured SPX lagging SPY +0.065); futures splice choice (3 variants — testable).
- Decay. At seconds-scale this is permanent structure; at 1min it may be fully arbitraged/invisible.
- Correct test. Both-fresh 1min xcorr ES↔SPY with staleness null; robustness across the three futures adjustment variants.
- MT exposure. Low (one pre-specified triple).
- Cost sensitivity. n/a (structural claim, not a strategy).
- Impl. Ours. ✅
- Limitation. 1min floor; ES data starts 2008.
- Honest new test. Is ANY ES→SPY lead detectable at 1min after the staleness null, and is SPX→anything pure artifact? Status: robust at sub-second (literature); unknown at 1min on open data — genuine gap.
A5 — Cross-asset information flow: crypto ↔ crypto-exposed equities
- Claims. (Thin literature.) 24/7 crypto prices embed information that equity prices can only reflect at the next open.
- Granularity. 1min crypto (24/7) + equity opens. ✅ — this is a structural granularity advantage of our dataset.
- Confounds. Overnight-gap conventions (T4), selection of "exposed" equities (must be pre-specified), regime dependence (crypto-equity beta varies).
- Decay. Unknown — modern, underexplored on open data.
- Correct test. Does BTC's Friday-close→Monday-preopen return predict the Monday opening gap of pre-specified crypto-exposed equities, vs a placebo set and a permuted-weekend null?
- MT exposure. Low if the equity set and horizon are pre-registered.
- Cost sensitivity. Open-auction execution is costly; report the frontier.
- Impl. Ours. ✅
- Limitation. Short joint history (crypto-exposed equities mostly 2018→); few independent weekends (~400).
- Honest new test. Exactly the above — one of the few places our data can ask something not already answered. Status: unknown/genuine gap.
A6 — Weekend / Monday effect
- Claims. Negative Monday returns (French 1980).
- Granularity. Daily. ✅
- Confounds. Close conventions (T4); DST weeks (T7).
- Decay. The cleanest corpse: gone post-publication (Schwert 2003; Marquering et al. 2006).
- Correct test. Day-of-week means with permuted-calendar null + SPA against the full day-of-week universe (STW 2001 protocol).
- MT exposure. High by construction — the calendar space.
- Cost sensitivity. Any exploitation is high-turnover.
- Impl. Ours. ✅
- Honest new test. Re-confirmation of absence on 2000–2026 open data, published as a negative control for the calendar pipeline. Status: decayed.
A7 — Turn-of-month
- Claims. Returns concentrate around month boundaries (Ariel 1987; Lakonishok–Smidt 1988).
- Granularity. Daily. ✅
- Confounds. Month-boundary volume/flows are real mechanics (pension/401k flows) — a mechanism, not an artifact; but overlap with OpEx week and quarter-ends must be disentangled.
- Decay. The last survivor as of Marquering et al. 2006. Post-2006 behavior on open data = open question.
- Correct test. Pre-specified window (−1..+3 trading days), permuted- calendar null, SPA vs the full window universe, sub-period stability.
- MT exposure. Moderate — window choice is the researcher degree of freedom; pre-register ONE window.
- Cost sensitivity. Low-frequency (12×/year) — the rare calendar effect that could survive costs if real.
- Impl. Ours. ✅
- Honest new test. Did the last survivor survive 2006–2026? Status: disputed — the most interesting calendar re-test.
A8 — Intraday momentum (first → last half-hour)
- Claims. First half-hour market return predicts last half-hour (Gao et al. 2018); mechanism: gamma hedging (Baltussen et al. 2021).
- Granularity. 1min/30min SPY. ✅ perfect fit.
- Confounds. Overnight-gap inclusion choice; T4 close convention; spread seasonality (U-shape) at both ends of the day.
- Decay. Published 2018 — post-publication window (2018–2026) is exactly what open data can measure now.
- Correct test. Pre-registered replication (their exact spec) + OOS post-2018 sample + gamma-state split using our options chains; DSR for the spec search.
- MT exposure. Low if the published spec is frozen.
- Cost sensitivity. Two trades/day at the most liquid instrument's most liquid hours — survivable in principle; measure.
- Impl. Ours. ✅
- Honest new test. The cleanest possible decay measurement: published effect, published spec, untouched post-publication data. Status: disputed (post-2018 fate unknown).
A9 — Intraday U-shape (open/close vol & spread concentration)
- Claims. Volatility, volume, spreads peak at open and close (Wood et al. 1985).
- Granularity. 1min. ✅
- Confounds. None — this one is real microstructure, and it is itself a confounder for other intraday claims.
- Decay. Robust across decades.
- Correct test. Descriptive profile with bootstrap bands.
- Cost sensitivity. n/a (input to the cost model, not a strategy).
- Honest contribution. Measure the intraday profile of OUR bounce null (taxonomy open item) so expE can subtract it. Status: robust — use as positive control + cost-model input.
A10 — Overnight vs intraday return split
- Claims. Equity returns accrue disproportionately overnight.
- Granularity. Daily open/close (+1min for convention checks). ✅
- Confounds. T4 is central: auction close vs last bar changes overnight returns mechanically; stale opens for illiquid names.
- Decay. Persistent in the literature but convention-sensitive — disputed as economics vs plumbing.
- Correct test. Recompute under BOTH close conventions and both open definitions (first 1min bar vs daily open field); effect must survive all four.
- MT exposure. Low.
- Cost sensitivity. High (daily turnover).
- Honest new test. Quantify how much of the overnight premium is convention-dependent on this dataset. Status: disputed.
Methods inventory (with macOS arm64 status)
| Method | Use | Implementation | arm64 |
|---|---|---|---|
| Lo–MacKinlay VR + block bootstrap | Q1 scans | ours (§8.1-gated) + arch.unitroot.VarianceRatio cross-check |
✅ |
| Lo (1991) modified R/S | long-memory claims | to implement, gate on synthetic long-memory | ✅ |
| Lagged xcorr / Granger | Q2 | ours + statsmodels grangercausalitytests |
✅ |
| Roll / Corwin–Schultz / CHL / EDGE spreads | cost model, T1 null | ours (Roll); bidask package (EDGE, pure numpy) + own CS/CHL |
✅ |
| Moving-block / stationary bootstrap | all CIs | ours (Künsch); Politis–Romano to add for RC/SPA | ✅ |
| White RC / Hansen SPA / Romano–Wolf StepM | expF | no maintained OSS — implement ourselves, gate on synthetic | ✅ (numpy) |
| Benjamini–Hochberg FDR | scan triage | trivial; statsmodels.stats.multitest |
✅ |
| Deflated Sharpe / PBO-CSCV | expF/expH | implement from Bailey–López de Prado formulas | ✅ |
| Bai–Perron breaks | regime robustness | statsmodels (partial) / own dynamic-programming impl |
✅ |
Cross-cutting conclusion. On modern liquid U.S. equities, the honest prior for every gross short-horizon effect is decay toward zero, and for every net effect, death by costs. The genuine opportunities on THIS data are: (i) decay re-measurements with artifact shares reported (A1–A3, A8), (ii) the structural questions our dataset uniquely reaches (A4 at 1min, A5 cross-asset 24/7), (iii) the last-survivor calendar re-test (A7), and (iv) methodological artifact quantifications (A2 Fisher share, A10 convention share) — all publishable regardless of sign.