--- project: anomaly-atlas document: P001 — The Artifact Frontier, Part I author: Simon-Pierre Boucher contact: contact@spboucher.ai data_source: hfmarketdata.io created: 2026-08-12 status: published pub_id: P001 version: 1 title: "The Artifact Frontier: What Survives Honest Testing in Open High-Frequency Market Data?" subtitle: "Part I — Methods, artifact taxonomy, and train-period results (2000–2016)" abstract: "We systematically scan open 1-minute-to-daily market data (hfmarketdata.io, sole source) for short-horizon mean-reversion, lead-lag, and calendar anomalies, under a pre-registered protocol: every detector must first pass a synthetic-data gate; every scan runs against measured artifact nulls; hypothesis budgets are declared before testing; and validation layers (multiple-testing correction, transaction costs) are applied in sequence. On the 2000–2016 train split, 62% of the 372 searched rules are naively 'significant' and 18% survive Hansen's SPA — yet the survivors carry physically implausible paper Sharpes, and a deliberately included known artifact (the SPX→SPY 'lead') survives statistical correction unharmed. The cost layer then eliminates essentially everything: the median surviving rule breaks even at 1.1% of one half-spread per trade, and zero intraday rules survive paying the full half-spread. The calendar family — including turn-of-month, the last published survivor — dies against a permuted-calendar null with an 8-test budget. Our principal positive contributions are a measured artifact taxonomy for this dataset and a demonstrated three-layer validation doctrine: artifact nulls, search correction, and costs are independent filters, and no one of them substitutes for another. Out-of-sample confirmation on the untouched validation split is reported in Part II." --- # The Artifact Frontier — Part I *Every number in this document regenerates from committed `results.json` files, and every figure below is rendered live from them by the site. The [research log](/doc/research/LOG.md) is the audit trail; the [charter](/doc/CLAUDE.md) is the pre-registered protocol.* ## 1. Question and doctrine Given only open high-frequency market data — 1-minute to daily bars, no quotes — which statistical regularities are *real* (reproducible out-of-sample, robust to artifacts, economically nonzero after costs), and which are plumbing? The project's doctrine, stated before any data was touched: **every candidate anomaly is an artifact until proven otherwise; in-sample results are never findings; negative results are first-class.** Three design rules operationalize this: 1. **Test the tests.** No detector touches real data before passing a synthetic gate: it must find *nothing* in a pure random walk, recover planted effects, and flag a pure bid-ask-bounce series as artifact. The gate is not ceremonial — it caught two real bugs before they could contaminate results (a variance-ratio estimator reading ~1/q on random walks, and an end-freshness-only synchronization mask that silently attenuated correlations via the Epps mechanism). 2. **Pre-registration.** Universes, time splits (train 2000–2016 / validation 2016–2021 / sealed holdout 2022→), hypothesis budgets (22, declared in [research_gaps](/doc/research/research_gaps.md)), and every experiment's falsification criterion were frozen in writing before the corresponding run. 3. **A frozen cache as the reproducibility anchor.** All data flows through one client that never silently refetches; every response is indexed in a committed data manifest. Two of the seven experiments below ran with **zero network requests**, entirely from the frozen cache. ## 2. The data, measured Experiment A established empirically what the source actually provides: true 1-minute bars for ~7,700 stocks and ~5,200 ETFs from January 2000 (futures and indices from 2008, FX from 2010, crypto from 2013), plus daily options chains with quotes and Greeks over 67 quarters. Load-bearing facts that the documentation does not state: timestamps are US-Eastern wall-clock, bar-start labeled; **bars exist only where trades occurred** (no zero-volume placeholders — an illiquid name printed 38 bars in a full session); daily bars carry the official auction close and consolidated volume, both absent from the 1-minute series; and dividend-adjusted prices are re-based to the vendor's build date, so adjusted series are not point-in-time stable. Responses are hard-capped at 50,000 rows. Full profile: [data_source_profile](/doc/research/data_source_profile.md). ## 3. The artifact taxonomy, with magnitudes The project's first deliverable is a catalogue of the mechanisms in *this dataset* that manufacture fake anomalies — each with detection code and a measured magnitude ([artifact_taxonomy](/doc/research/artifact_taxonomy.md)). Experiment B measured the key ones on a pre-specified 42-ticker universe (Q1 2024, RTH 1-minute): {{figure:expB_artifact_baselines}} Highlights: bid-ask bounce alone produces AC1 of −0.23 and VR(30) of 0.55 in the stalest liquidity tercile with zero planted economics; LOCF joins make SPY spuriously "lead" mid-staleness names (+0.047 at +1 min, Spearman vs staleness +0.43) while *diluting* the lead of ultra-stale names — the artifact is non-monotone; and the SPX index print lags SPY by one minute (+0.065 at 0.965 contemporaneous correlation) — a Fisher (1966) effect, measured live. The intraday profile (Experiment E) adds the time-of-day dimension: volatility is U-shaped (6.6 bp at the open, 2.4 midday, 2.9 at the close) while the effective spread declines monotonically (2.8 → 1.2 bp) — any "first-30-minutes" return claim fights 2–3× the midday artifact level. ## 4. The scans (train split only, Level 0 by construction) **Mean-reversion (Experiment C).** 127 cells (ticker × timeframe × sub-period), each tested against a bounce null and FDR-corrected. The scan's most valuable output was about the null itself: a daily effective spread combined with pure Roll alternation predicts *impossible* intraday autocorrelations (−3 to −27), because consecutive intraday closes rarely flip sides. The corrected, variance-consistent triage — an MA(1) null that absorbs *all* lag-1 effects — leaves 14 cells of genuine multi-lag reversion, concentrated in a daily 2008–2015 mega-cap/index family (XOM excess AC1 −0.13, SPY −0.055, both FDR) and a few 1-minute cells (JPM −0.24 VR-excess at the 30-minute horizon). {{figure:expC_reversion_scan}} **Lead-lag (Experiment D).** 49 pairs, two windows, with the raw-LOCF versus both-fresh comparison built in — so the non-synchronicity artifact is *measured*, not consumed. In 2006–2007, minute-scale market→component diffusion was real on synchronized samples (SPY led every sector ETF by +0.07..+0.15). By 2014–2015 it had collapsed to ±0.05 — the cleanest decay measurement of the project. Two structural facts survive synchronization: a splice-invariant ES↔SPY cross-serial effect (−0.032, identical across all three futures adjustments), and the SPX→SPY "lead" (+0.132) — which survives *because synchronizing print times cannot fix a computed index*. {{figure:expD_leadlag_scan}} **Calendar (Experiment E).** Eight pre-declared tests (day-of-week ×5, turn-of-month, pre/post-holiday) on SPY against a within-year permuted-calendar null with a family-wise max-statistic. **Nothing survives** (best marginal p = 0.24; family-wise p ≥ 0.93 everywhere). Turn-of-month — the last survivor in the published literature as of 2006 — fails and decays inside the train period (+7.9 bp in 2000–2007 → +1.6 bp in 2008–2015). Both pipeline controls behaved: the Monday effect stayed dead, and the volatility U-shape was strongly present. {{figure:expE_calendar_scan}} ## 5. The survival curve: statistical correction is not artifact correction Experiment F pushed *everything the scans searched* — 372 signed rules — through the correction battery: naive t-tests, Benjamini–Hochberg FDR, White's Reality Check and Hansen's SPA over stationary bootstraps, and the Deflated Sharpe Ratio. {{figure:expF_multiple_testing}} The result that matters is not the 18% SPA survival rate — it is *what* survives: rules with paper Sharpes of 10–31 annualized, physically implausible, dominated by bounce harvesting (a contrarian rule mechanically earns −autocov₁ on paper, which is precisely the spread it would pay in reality). The canary proves the point: the SPX→SPY rule — an artifact we had already measured twice — **passes SPA comfortably**. Statistical correction corrects for *search*; it is structurally blind to *mechanism*. ## 6. The cost frontier closes the loop Experiment G re-priced the double-filtered pool (rules that beat both the artifact nulls and the search correction; 31 rules) under a declared cost model: net = gross − κ · (EDGE half-spread) · turnover, κ swept from 0 to 2. {{figure:expG_cost_frontier}} The median rule breaks even at **κ\* = 0.011** — it captures about 1% of one half-spread per trade. Three rules survive κ = 0.1 (all in sparse names), one survives κ = 0.25 (CKX, an ultra-sparse name with a wide, noisy spread estimate — the classic profile of an estimation artifact, forwarded to the validation split with a skeptical prior rather than discarded by hand), and **zero rules — none — survive paying the full half-spread.** The pre-registered falsification clause ("the costs-kill story fails if any intraday rule survives κ = 1") did not trigger. ## 7. What Part I establishes 1. **A three-layer validation doctrine, demonstrated rather than asserted.** Artifact nulls, search correction, and transaction costs filter *different* failure modes; each layer passed things the next one killed. 2. **A measured artifact taxonomy (T1–T7)** for open bar data, with the magnitudes above and neutralization rules, validated on synthetic ground truth. 3. **Negative results with teeth**: the calendar family is empty under an honest budget; minute-scale lead-lag decayed an order of magnitude between 2006 and 2015; and nothing in the searched universe pays for its own spread on the train split. 4. **Methodological findings**: spread-based bounce nulls must be variance-consistent with the target series; synchronization on print times cannot de-artifact computed indices; LOCF distortion of lead-lag is non-monotone in staleness. ## 8. Limitations All Part-I numbers are train-split, in-sample by design — their out-of-sample fate on the untouched 2016–2021 validation split is Part II. Bar data carries no quotes: costs are estimated (EDGE), not observed. Holiday-class calendar tests are low-powered (n = 144). The universe is U.S.-equity-centric; crypto and FX hypotheses await volume-semantics verification. Capacity is out of scope. ## 9. Reproducibility Everything regenerates from the repository: commit `f59e891` (data profile, client), `9b6beea` (artifact baselines), `c4977f2` (reversion scan), `a7f66cb` (lead-lag scan), `b228862` (calendar scan), `7a82cc1` (survival battery), `881399d` (cost frontier). Each experiment's `results.json` embeds the hardware manifest and the client's instrumentation; the data manifest (`data_manifest/index.jsonl`) indexes every API response consumed. Detector gates: `benchmarks/synthetic/` (29 tests at the time of writing). ## References Key sources (full annotated list with DOIs: [bibliography](/doc/research/bibliography.md)): Roll (1984); Fisher (1966); Scholes & Williams (1977); Epps (1979); Lo & MacKinlay (1988, 1990); Sullivan, Timmermann & White (2001); White (2000); Hansen (2005); Bailey & López de Prado (2014); Harvey, Liu & Zhu (2016); McLean & Pontiff (2016); Novy-Marx & Velikov (2016); Chordia, Roll & Subrahmanyam (2005); Ardia, Guidotti & Kroencke (2024); Chen & Velikov (2022).