SPB Git

spb/anomaly-atlas Public License

Systematic discovery & rigorous validation of statistical anomalies in open HF market data (hfmarketdata.io) — pre-registered, artifact-null-driven, fully reproducible. Live atlas: www.anomaly-atlas.io

Python 61.4% JavaScript 28.7% CSS 8.6% Shell 0.7% Makefile 0.5%

# project: anomaly-atlas document: expG_cost_frontier/hypothesis author: Simon-Pierre Boucher contact: contact@spboucher.ai data_source: hfmarketdata.io created: 2026-08-12 modified: 2026-08-12 status: final

# Hypothesis — expG_cost_frontier

Pre-specified 2026-08-12 before the sweep ran.

text
Hypothesis
  Tests H-family question Q5 on the expF double-filtered pool (the only
  rules that beat BOTH the artifact nulls and the search correction,
  gross): 12 reversion cells + ES->SPY + the expD 2014-2015 FDR lead-lag
  set. Prior (Novy-Marx-Velikov; Chen-Velikov): high-turnover short-horizon
  rules die well below realistic costs. Expectation: every intraday rule
  has breakeven cost multiplier kappa* << 0.25 (a quarter of the half
  spread — the patient-execution floor of Frazzini et al.); at most the
  1day cells are ambiguous.

Falsification criterion
  The "costs kill short-horizon anomalies" story is falsified if ANY
  intraday rule from the pool survives kappa = 1 (paying the full
  half-spread per trade) with positive net mean.

Artifact null(s)
  None new — this experiment IS the cost layer. The spread input is the
  EDGE estimate from daily OHLC of the TRADED instrument over the rule's
  own period (independent granularity, as in expC).

Method (pre-declared)
  Pool: mechanical intersection recomputed from committed expC/expF/expD
  results (no hand-picking). Rule return streams rebuilt from the frozen
  cache exactly as in expF, now with per-day TURNOVER = sum |delta
  position| (entry included). Cost model: net_day = gross_day - kappa *
  half_spread * turnover_day, half_spread = EDGE/2 per (traded instrument,
  period). Sweep kappa in {0, 0.1, 0.25, 0.5, 1.0, 2.0}; for each rule
  report gross mean, turnover/day, half-spread (bp), net mean and its
  block-bootstrap t at each kappa, and the analytic breakeven
  kappa* = gross_mean / (half_spread * mean_turnover). Survivor counts at
  each kappa level; ES uses the same machinery with the futures caveat
  declared (cost structure differs; kappa* still reported).

Result
  Run 20260812T072128Z (cache-served). Pool: 31 rules (12 reversion cells,
  16 lead-lag incl. ES x3 splices, sparse-name pairs). kappa* median =
  0.0114, max = 0.28 (all 31 defined after invalid-day filtering): at the MEDIAN the double-filtered survivors capture
  0.7% of one half-spread per trade. Survivors: kappa=0.1 -> 3 rules (all
  sparse-name: AXDX 30min, CKX 1day, HTD 5min); kappa=0.25 -> CKX 1day
  alone (kappa*=0.28); kappa=0.5 -> NONE; kappa=1.0 -> NONE. Zero intraday
  rules survive kappa=1 — the pre-registered falsification did NOT trigger.
  ES->SPY: kappa*=0.0028, identical across splices.

Interpretation
  (Level 0.) The cost frontier does exactly what the literature priors
  said it would (Novy-Marx-Velikov; Chen-Velikov): everything that
  survived the artifact nulls AND the search correction dies at a fraction
  of realistic costs. The gross "profits" were spread capture one cannot
  buy. Q5's answer on this pool: the frontier sits at ~0.01-0.04 of a
  half-spread for intraday rules — an order of magnitude below even the
  most optimistic patient-execution assumptions. The single kappa=0.25
  survivor (CKX 1day, an ultra-sparse name with a wide, noisy EDGE
  estimate) is exactly the profile of a measurement artifact — it goes to
  expH's validation split with a strong skeptical prior rather than being
  discarded by hand.

Next experiment
  expH: validation-split evaluation of whatever survives kappa >= 0.25
  (if anything); otherwise expH validates the negative finding and the
  atlas receives its first confidence-labeled entries.