SPB Git

spb/anomaly-atlas Public License

Systematic discovery & rigorous validation of statistical anomalies in open HF market data (hfmarketdata.io) — pre-registered, artifact-null-driven, fully reproducible. Live atlas: www.anomaly-atlas.io

Python 61.4% JavaScript 28.7% CSS 8.6% Shell 0.7% Makefile 0.5%

Phase 1: literature sweep — 6 theme notes + 52-source verified bibliography

All citations verified against OpenAlex (DOI/venue/year, accessed
2026-08-12). Priors adopted: McLean-Pontiff decay base rate, STW calendar
reality check, t>=3 hurdle, EDGE spread estimator for expG.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Simon-Pierre Boucher committed 5 h ago (Aug 12, 2026) parent 9b6beea

Showing 9 changed files with +440 and −3

modified research/LOG.md +26 −0
@@ -82,3 +82,29 @@ criteria failed to trigger. All Level 0 by construction.
82 82 **Decision.** artifact_taxonomy.md now carries 7 entries with measured
83 83 magnitudes (T1–T7). Next: expC consumes the bucket-matched bounce null;
84 84 Phase 1 literature sweep in parallel.
85 +
86 +## 2026-08-12 02:45 ET — Phase 1 complete: literature sweep (52 verified sources)
87 +
88 +**Question.** What does the literature establish about Q1-Q3 anomalies, the
89 +statistics of not fooling yourself, microstructure artifacts, and costs?
90 +
91 +**Method.** Three parallel verification passes against OpenAlex (every
92 +citation confirmed: title, authors, year, venue, DOI; access date
93 +2026-08-12), plus discovery searches for 2005-2025 decay/replication work.
94 +Six theme notes written in research/notes/; consolidated bibliography.
95 +
96 +**Key priors adopted.** (i) McLean-Pontiff: -26% OOS, -58% post-publication
97 +— the base rate. (ii) Sullivan-Timmermann-White: calendar effects vanish
98 +under the Reality Check. (iii) Chordia-Goyal-Saretto/Harvey-Liu-Zhu:
99 +t-hurdle >= 3.0 floor, 3.4-3.8 band reported. (iv) Novy-Marx-Velikov +
100 +Chen-Velikov: high-turnover (short-horizon) anomalies mostly die at costs
101 +(~4bp/month average net). (v) Chordia-Roll-Subrahmanyam: minute-scale
102 +inefficiencies arbitraged within 5-60 min already by 2005. (vi) Marquering
103 +et al.: turn-of-month was the last calendar survivor (2006) — the single
104 +most interesting re-test. (vii) EDGE (Ardia et al. 2024) chosen as primary
105 +OHLC spread estimator for expG.
106 +
107 +**Decision.** Error-rate ladder fixed for Phase 9: FDR for scans -> SPA/StepM
108 +vs artifact nulls for Level 1 -> DSR with logged trial counts. Next: Phase 2
109 +state-of-the-art map, then research_gaps (Phase 3). In parallel (user
110 +request): major web platform upgrade (mobile + comments).
modified research/bibliography.md +88 −2
@@ -5,9 +5,95 @@ author: Simon-Pierre Boucher
5 5 contact: contact@spboucher.ai
6 6 data_source: hfmarketdata.io
7 7 created: 2026-08-12
8 status: draft
8 +modified: 2026-08-12
9 +status: reviewed
9 10 ---
10 11
11 12 # Bibliography
12 13
13 Every consulted source with URL and access date. Phase 1 populates this.
14 +Every consulted source, with URL and access date. All entries verified
15 +against OpenAlex on **2026-08-12** (metadata: title, authors, year, venue,
16 +DOI). Where online/print years differ, the print-issue year is used with the
17 +online year noted in the theme notes. Theme notes in `research/notes/`.
18 +
19 +## §4.1 — Short-horizon anomalies & lead-lag
20 +
21 +- Lo, A. W. & MacKinlay, A. C. (1988). Stock Market Prices Do Not Follow Random Walks: Evidence from a Simple Specification Test. *Review of Financial Studies* 1(1), 41–66. https://doi.org/10.1093/rfs/1.1.41 (accessed 2026-08-12)
22 +- Lehmann, B. N. (1990). Fads, Martingales, and Market Efficiency. *Quarterly Journal of Economics* 105(1), 1–28. https://doi.org/10.2307/2937816 (accessed 2026-08-12)
23 +- Jegadeesh, N. (1990). Evidence of Predictable Behavior of Security Returns. *Journal of Finance* 45(3), 881–898. https://doi.org/10.1111/j.1540-6261.1990.tb05110.x (accessed 2026-08-12)
24 +- Lo, A. W. & MacKinlay, A. C. (1990). When Are Contrarian Profits Due to Stock Market Overreaction? *Review of Financial Studies* 3(2), 175–205. https://doi.org/10.1093/rfs/3.2.175 (accessed 2026-08-12)
25 +- Epps, T. W. (1979). Comovements in Stock Prices in the Very Short Run. *Journal of the American Statistical Association* 74(366a), 291–298. https://doi.org/10.1080/01621459.1979.10482508 (accessed 2026-08-12)
26 +- Scholes, M. & Williams, J. (1977). Estimating betas from nonsynchronous data. *Journal of Financial Economics* 5(3), 309–327. https://doi.org/10.1016/0304-405X(77)90041-1 (accessed 2026-08-12)
27 +- Chordia, T., Roll, R. & Subrahmanyam, A. (2005). Evidence on the speed of convergence to market efficiency. *Journal of Financial Economics* 76(2), 271–292. https://doi.org/10.1016/j.jfineco.2004.06.004 (accessed 2026-08-12)
28 +
29 +## §4.1 — Calendar & intraday effects
30 +
31 +- French, K. R. (1980). Stock returns and the weekend effect. *Journal of Financial Economics* 8(1), 55–69. https://doi.org/10.1016/0304-405X(80)90021-5 (accessed 2026-08-12)
32 +- Ariel, R. A. (1987). A monthly effect in stock returns. *Journal of Financial Economics* 18(1), 161–174. https://doi.org/10.1016/0304-405X(87)90066-3 (accessed 2026-08-12)
33 +- Lakonishok, J. & Smidt, S. (1988). Are Seasonal Anomalies Real? A Ninety-Year Perspective. *Review of Financial Studies* 1(4), 403–425. https://doi.org/10.1093/rfs/1.4.403 (accessed 2026-08-12)
34 +- Wood, R. A., McInish, T. H. & Ord, J. K. (1985). An Investigation of Transactions Data for NYSE Stocks. *Journal of Finance* 40(3), 723–739. https://doi.org/10.2307/2327796 (accessed 2026-08-12)
35 +- Gao, L., Han, Y., Li, S. Z. & Zhou, G. (2018). Market intraday momentum. *Journal of Financial Economics* 129(2), 394–414. https://doi.org/10.1016/j.jfineco.2018.05.009 (accessed 2026-08-12)
36 +- Baltussen, G., Da, Z., Lammers, S. & Martens, M. (2021). Hedging demand and market intraday momentum. *Journal of Financial Economics* 142(1), 377–403. https://doi.org/10.1016/j.jfineco.2021.04.029 (accessed 2026-08-12)
37 +
38 +## §4.1 — Post-publication decay
39 +
40 +- Schwert, G. W. (2003). Anomalies and Market Efficiency. *Handbook of the Economics of Finance*, 939–974. https://doi.org/10.1016/S1574-0102(03)01024-0 (accessed 2026-08-12)
41 +- McLean, R. D. & Pontiff, J. (2016). Does Academic Research Destroy Stock Return Predictability? *Journal of Finance* 71(1), 5–32. https://doi.org/10.1111/jofi.12365 (accessed 2026-08-12)
42 +- Marquering, W., Nisser, J. & Valla, T. (2006). Disappearing anomalies: a dynamic analysis of the persistence of anomalies. *Applied Financial Economics* 16(4), 291–302. https://doi.org/10.1080/09603100500400361 (accessed 2026-08-12)
43 +- Chordia, T., Subrahmanyam, A. & Tong, Q. (2014). Have capital market anomalies attenuated in the recent era of high liquidity and trading activity? *Journal of Accounting and Economics* 58(1), 41–58. https://doi.org/10.1016/j.jacceco.2014.06.001 (accessed 2026-08-12)
44 +- Jacobs, H. & Müller, S. (2020). Anomalies across the globe: Once public, no longer existent? *Journal of Financial Economics* 135(1), 213–230. https://doi.org/10.1016/j.jfineco.2019.06.004 (accessed 2026-08-12)
45 +
46 +## §4.2 — Multiple testing, data snooping, backtest overfitting
47 +
48 +- White, H. (2000). A Reality Check for Data Snooping. *Econometrica* 68(5), 1097–1126. https://doi.org/10.1111/1468-0262.00152 (accessed 2026-08-12)
49 +- Sullivan, R., Timmermann, A. & White, H. (2001). Dangers of data mining: The case of calendar effects in stock returns. *Journal of Econometrics* 105(1), 249–286. https://doi.org/10.1016/s0304-4076(01)00077-x (accessed 2026-08-12)
50 +- Hansen, P. R. (2005). A Test for Superior Predictive Ability. *Journal of Business & Economic Statistics* 23(4), 365–380. https://doi.org/10.1198/073500105000000063 (accessed 2026-08-12)
51 +- Romano, J. P. & Wolf, M. (2005). Stepwise Multiple Testing as Formalized Data Snooping. *Econometrica* 73(4), 1237–1282. https://doi.org/10.1111/j.1468-0262.2005.00615.x (accessed 2026-08-12)
52 +- Benjamini, Y. & Hochberg, Y. (1995). Controlling the False Discovery Rate: A Practical and Powerful Approach to Multiple Testing. *Journal of the Royal Statistical Society, Series B* 57(1), 289–300. https://doi.org/10.1111/j.2517-6161.1995.tb02031.x (accessed 2026-08-12)
53 +- Harvey, C. R., Liu, Y. & Zhu, H. (2016). …and the Cross-Section of Expected Returns. *Review of Financial Studies* 29(1), 5–68. https://doi.org/10.1093/rfs/hhv059 (accessed 2026-08-12)
54 +- Harvey, C. R. (2017). Presidential Address: The Scientific Outlook in Financial Economics. *Journal of Finance* 72(4), 1399–1440. https://doi.org/10.1111/jofi.12530 (accessed 2026-08-12)
55 +- Bailey, D. H. & López de Prado, M. (2014). The Deflated Sharpe Ratio: Correcting for Selection Bias, Backtest Overfitting, and Non-Normality. *Journal of Portfolio Management* 40(5), 94–107. https://doi.org/10.3905/jpm.2014.40.5.094 (accessed 2026-08-12)
56 +- Bailey, D. H., Borwein, J. M., López de Prado, M. & Zhu, Q. J. (2014). Pseudo-Mathematics and Financial Charlatanism: The Effects of Backtest Overfitting on Out-of-Sample Performance. *Notices of the AMS* 61(5), 458–471. https://doi.org/10.1090/noti1105 (accessed 2026-08-12)
57 +- Bailey, D. H., Borwein, J. M., López de Prado, M. & Zhu, Q. J. (2016). The probability of backtest overfitting. *Journal of Computational Finance* 20(4), 39–69. https://doi.org/10.21314/jcf.2016.322 (accessed 2026-08-12)
58 +- Hou, K., Xue, C. & Zhang, L. (2020). Replicating Anomalies. *Review of Financial Studies* 33(5), 2019–2133. https://doi.org/10.1093/rfs/hhy131 (accessed 2026-08-12)
59 +- Jensen, T. I., Kelly, B. & Pedersen, L. H. (2023). Is There a Replication Crisis in Finance? *Journal of Finance* 78(5), 2465–2518. https://doi.org/10.1111/jofi.13249 (accessed 2026-08-12)
60 +- Chordia, T., Goyal, A. & Saretto, A. (2020). Anomalies and False Rejections. *Review of Financial Studies* 33(5), 2134–2179. https://doi.org/10.1093/rfs/hhaa018 (accessed 2026-08-12)
61 +- Harvey, C. R. & Liu, Y. (2020). False (and Missed) Discoveries in Financial Economics. *Journal of Finance* 75(5), 2503–2553. https://doi.org/10.1111/jofi.12951 (accessed 2026-08-12)
62 +- Giglio, S., Liao, Y. & Xiu, D. (2021). Thousands of Alpha Tests. *Review of Financial Studies* 34(7), 3456–3496. https://doi.org/10.1093/rfs/hhaa111 (accessed 2026-08-12)
63 +
64 +## §4.3 — Microstructure & artifacts
65 +
66 +- Roll, R. (1984). A Simple Implicit Measure of the Effective Bid-Ask Spread in an Efficient Market. *Journal of Finance* 39(4), 1127–1139. https://doi.org/10.1111/j.1540-6261.1984.tb03897.x (accessed 2026-08-12)
67 +- Blume, M. E. & Stambaugh, R. F. (1983). Biases in computed returns: An application to the size effect. *Journal of Financial Economics* 12(3), 387–404. https://doi.org/10.1016/0304-405x(83)90056-9 (accessed 2026-08-12)
68 +- Fisher, L. (1966). Some New Stock-Market Indexes. *Journal of Business* 39(S1), 191–225. https://doi.org/10.1086/294848 (accessed 2026-08-12)
69 +- Zhang, L., Mykland, P. A. & Aït-Sahalia, Y. (2005). A Tale of Two Time Scales: Determining Integrated Volatility with Noisy High-Frequency Data. *JASA* 100(472), 1394–1411. https://doi.org/10.1198/016214505000000169 (accessed 2026-08-12)
70 +- Aït-Sahalia, Y., Mykland, P. A. & Zhang, L. (2005). How Often to Sample a Continuous-Time Process in the Presence of Market Microstructure Noise. *Review of Financial Studies* 18(2), 351–416. https://doi.org/10.1093/rfs/hhi016 (accessed 2026-08-12)
71 +- Hansen, P. R. & Lunde, A. (2006). Realized Variance and Market Microstructure Noise. *Journal of Business & Economic Statistics* 24(2), 127–161. https://doi.org/10.1198/073500106000000071 (accessed 2026-08-12)
72 +- Bandi, F. M. & Russell, J. R. (2006). Separating microstructure noise from volatility. *Journal of Financial Economics* 79(3), 655–692. https://doi.org/10.1016/j.jfineco.2005.01.005 (accessed 2026-08-12)
73 +- Bandi, F. M. & Russell, J. R. (2008). Microstructure Noise, Realized Variance, and Optimal Sampling. *Review of Economic Studies* 75(2), 339–369. https://doi.org/10.1111/j.1467-937x.2008.00474.x (accessed 2026-08-12)
74 +
75 +## §4.4 — Time-series methodology
76 +
77 +- Lo, A. W. (1991). Long-Term Memory in Stock Market Prices. *Econometrica* 59(5), 1279–1313. https://doi.org/10.2307/2938368 (accessed 2026-08-12)
78 +- Granger, C. W. J. (1969). Investigating Causal Relations by Econometric Models and Cross-spectral Methods. *Econometrica* 37(3), 424–438. https://doi.org/10.2307/1912791 (accessed 2026-08-12)
79 +- Künsch, H. R. (1989). The Jackknife and the Bootstrap for General Stationary Observations. *Annals of Statistics* 17(3), 1217–1241. https://doi.org/10.1214/aos/1176347265 (accessed 2026-08-12)
80 +- Politis, D. N. & Romano, J. P. (1994). The Stationary Bootstrap. *JASA* 89(428), 1303–1313. https://doi.org/10.1080/01621459.1994.10476870 (accessed 2026-08-12)
81 +- Bai, J. & Perron, P. (1998). Estimating and Testing Linear Models with Multiple Structural Changes. *Econometrica* 66(1), 47–78. https://doi.org/10.2307/2998540 (accessed 2026-08-12)
82 +
83 +## §4.5 — Transaction costs
84 +
85 +- Hasbrouck, J. (2009). Trading Costs and Returns for U.S. Equities: Estimating Effective Costs from Daily Data. *Journal of Finance* 64(3), 1445–1477. https://doi.org/10.1111/j.1540-6261.2009.01469.x (accessed 2026-08-12)
86 +- Corwin, S. A. & Schultz, P. (2012). A Simple Way to Estimate Bid-Ask Spreads from Daily High and Low Prices. *Journal of Finance* 67(2), 719–760. https://doi.org/10.1111/j.1540-6261.2012.01729.x (accessed 2026-08-12)
87 +- Abdi, F. & Ranaldo, A. (2017). A Simple Estimation of Bid-Ask Spreads from Daily Close, High, and Low Prices. *Review of Financial Studies* 30(12), 4437–4480. https://doi.org/10.1093/rfs/hhx084 (accessed 2026-08-12)
88 +- Ardia, D., Guidotti, E. & Kroencke, T. A. (2024). Efficient estimation of bid–ask spreads from open, high, low, and close prices. *Journal of Financial Economics* 161, 103916. https://doi.org/10.1016/j.jfineco.2024.103916 (accessed 2026-08-12)
89 +- Fong, K. Y. L., Holden, C. W. & Trzcinka, C. A. (2017). What Are the Best Liquidity Proxies for Global Research? *Review of Finance* 21(4), 1355–1401. https://doi.org/10.1093/rof/rfx003 (accessed 2026-08-12)
90 +- Lesmond, D. A., Ogden, J. P. & Trzcinka, C. A. (1999). A New Estimate of Transaction Costs. *Review of Financial Studies* 12(5), 1113–1141. https://doi.org/10.1093/rfs/12.5.1113 (accessed 2026-08-12)
91 +- Bessembinder, H. (2003). Issues in Assessing Trade Execution Costs. *Journal of Financial Markets* 6(3), 233–257. https://doi.org/10.1016/s1386-4181(02)00064-2 (accessed 2026-08-12)
92 +- Novy-Marx, R. & Velikov, M. (2016). A Taxonomy of Anomalies and Their Trading Costs. *Review of Financial Studies* 29(1), 104–147. https://doi.org/10.1093/rfs/hhv063 (accessed 2026-08-12)
93 +- Frazzini, A., Israel, R. & Moskowitz, T. J. (2012). Trading Costs of Asset Pricing Anomalies. SSRN working paper. https://doi.org/10.2139/ssrn.2294498 (accessed 2026-08-12)
94 +- Chen, A. Y. & Velikov, M. (2022). Zeroing In on the Expected Returns of Anomalies. *Journal of Financial and Quantitative Analysis* 58(3), 968–1004. https://doi.org/10.1017/s0022109022000874 (accessed 2026-08-12)
95 +- Detzel, A., Novy-Marx, R. & Velikov, M. (2023). Model Comparison with Transaction Costs. *Journal of Finance* 78(3), 1743–1775. https://doi.org/10.1111/jofi.13225 (accessed 2026-08-12)
96 +
97 +## §4.6 — This dataset
98 +
99 +- HF Market Data API documentation & OpenAPI schema. https://www.hfmarketdata.io/docs (accessed 2026-08-12) — capabilities established empirically in `research/data_source_profile.md` (Experiment A).
modified research/notes/README.md +8 −1
@@ -10,4 +10,11 @@ status: draft
10 10
11 11 # Phase 1 literature notes
12 12
13 One note file per theme (CLAUDE.md §4.1-4.6). None written yet.
13 +One note file per theme (CLAUDE.md §4.1-4.6), all sources verified via OpenAlex on 2026-08-12:
14 +
15 +- [mean_reversion_leadlag.md](mean_reversion_leadlag.md) — Q1/Q2 founding results, lead-lag vs its artifact twin, decay evidence
16 +- [calendar_effects.md](calendar_effects.md) — Q3 classics, intraday momentum, what survived (turn-of-month) vs what died
17 +- [multiple_testing_snooping.md](multiple_testing_snooping.md) — §4.2, the most important: RC/SPA/StepM/FDR/DSR/PBO + finance base rates
18 +- [microstructure_artifacts.md](microstructure_artifacts.md) — §4.3: Roll, Blume-Stambaugh, Fisher, microstructure noise
19 +- [timeseries_methodology.md](timeseries_methodology.md) — §4.4: VR pitfalls, Lo R/S, Granger caveats, block bootstrap, CPCV, Bai-Perron
20 +- [transaction_costs.md](transaction_costs.md) — §4.5: OHLC spread estimators (EDGE primary), cost-survival base rates
added research/notes/calendar_effects.md +59 −0
@@ -0,0 +1,59 @@
1 +---
2 +project: anomaly-atlas
3 +document: Phase 1 notes — calendar & intraday seasonal effects
4 +author: Simon-Pierre Boucher
5 +contact: contact@spboucher.ai
6 +data_source: hfmarketdata.io
7 +created: 2026-08-12
8 +status: reviewed
9 +---
10 +
11 +# Calendar & intraday seasonal effects (Q3)
12 +
13 +*Phase 1 literature notes. Citations verified via OpenAlex, accessed
14 +2026-08-12. The calendar space is the p-hacking capital of finance — every
15 +claim below must be read against the Sullivan–Timmermann–White result (see
16 +multiple-testing notes).*
17 +
18 +## 1. The classic effects
19 +
20 +- French, K. R. (1980). Stock returns and the weekend effect. *Journal of Financial Economics* 8(1), 55–69. https://doi.org/10.1016/0304-405X(80)90021-5 — Systematically negative Monday returns. **Status: decayed** — one of the cleanest documented post-publication disappearances.
21 +- Ariel, R. A. (1987). A monthly effect in stock returns. *Journal of Financial Economics* 18(1), 161–174. https://doi.org/10.1016/0304-405X(87)90066-3 — Returns concentrate in the first half of the month (turn-of-month). **Status: disputed/partially persistent.**
22 +- Lakonishok, J. & Smidt, S. (1988). Are Seasonal Anomalies Real? A Ninety-Year Perspective. *Review of Financial Studies* 1(4), 403–425. https://doi.org/10.1093/rfs/1.4.403 — 90 years of DJIA data confirm turn-of-week/month/year and holiday effects, while *explicitly warning about data-snooping* — remarkably ahead of its time. **Status: the effects it confirmed have mostly decayed; the warning is robust.**
23 +- Wood, R. A., McInish, T. H. & Ord, J. K. (1985). An Investigation of Transactions Data for NYSE Stocks. *Journal of Finance* 40(3), 723–739. https://doi.org/10.2307/2327796 — Minute-data intraday **U-shape** in returns/volatility at open and close. **Status: robust as a *volatility/spread* pattern; as a *return* anomaly, likely-artifact** (auction mechanics, spread seasonality).
24 +
25 +## 2. The modern intraday variant
26 +
27 +- Gao, L., Han, Y., Li, S. Z. & Zhou, G. (2018). Market intraday momentum. *Journal of Financial Economics* 129(2), 394–414. https://doi.org/10.1016/j.jfineco.2018.05.009 — First half-hour SPY return predicts the last half-hour. Testable 1:1 on our data (SPY 1min, 2000→2026). **Status: disputed — decay after publication is itself a good expE sub-hypothesis.**
28 +- Baltussen, G., Da, Z., Lammers, S. & Martens, M. (2021). Hedging demand and market intraday momentum. *Journal of Financial Economics* 142(1), 377–403. https://doi.org/10.1016/j.jfineco.2021.04.029 — Mechanism: option gamma-hedging flows; the effect varies with hedging conditions. Gives us a *mechanism prior* — and our options chains (67 quarters) can proxy the hedging-demand state.
29 +
30 +## 3. Decay evidence specific to calendars
31 +
32 +- Schwert, G. W. (2003). Anomalies and Market Efficiency. *Handbook of the Economics of Finance*, 939–974. https://doi.org/10.1016/S1574-0102(03)01024-0 — Size, weekend, dividend, and January effects weaken or vanish after publication.
33 +- Marquering, W., Nisser, J. & Valla, T. (2006). Disappearing anomalies: a dynamic analysis of the persistence of anomalies. *Applied Financial Economics* 16(4), 291–302. https://doi.org/10.1080/09603100500400361 — Rolling-window analysis: weekend, holiday, time-of-month, January all gone post-publication; **turn-of-month persisted** (as of 2006). Makes turn-of-month the single most interesting calendar hypothesis to re-test on 2000–2026 data.
34 +
35 +## 4. Artifact confounds specific to Q3 (feed the taxonomy)
36 +
37 +1. **Spread/staleness intraday seasonality**: spreads and staleness are
38 + widest at the open — a "first-30-minutes effect" in *returns* can be pure
39 + T1/T2 artifact. expE must measure the intraday profile of our bounce null
40 + first (taxonomy open item).
41 +2. **Auction-close convention (T4)**: close-anchored calendar effects change
42 + magnitude depending on which close is used — daily bars vs last 1min bar
43 + differ by ~8 bp even on calm days (expA).
44 +3. **Session-length changes & DST**: ET wall-clock timestamps mean DST weeks
45 + shift the UTC session; naive UTC grouping manufactures day-of-week
46 + effects.
47 +4. **Hypothesis-space explosion**: day-of-week (5) × month (12) × turn-of-X
48 + × holiday × intraday half-hours (13) → hundreds of implicit tests. expE's
49 + budget must be pre-counted and corrected (White RC / SPA — see
50 + multiple-testing notes).
51 +
52 +## 5. Implications for expE
53 +
54 +Pre-register a SMALL set: (i) turn-of-month (the survivor per Marquering et
55 +al.), (ii) Monday effect (the canonical corpse — expected negative finding),
56 +(iii) intraday momentum first→last half-hour (Gao et al., with the
57 +gamma-hedging state split of Baltussen et al.), (iv) the intraday U-shape as
58 +an *artifact demonstration*, not an anomaly claim. Everything else is
59 +exploratory Level 0, reported as such.
added research/notes/mean_reversion_leadlag.md +53 −0
@@ -0,0 +1,53 @@
1 +---
2 +project: anomaly-atlas
3 +document: Phase 1 notes — short-horizon mean-reversion & lead-lag
4 +author: Simon-Pierre Boucher
5 +contact: contact@spboucher.ai
6 +data_source: hfmarketdata.io
7 +created: 2026-08-12
8 +status: reviewed
9 +---
10 +
11 +# Short-horizon mean-reversion & lead-lag (Q1, Q2)
12 +
13 +*Phase 1 literature notes. All citations verified via OpenAlex, accessed
14 +2026-08-12. Epistemic status flags: **robust** / **decayed** / **disputed** /
15 +**likely-artifact**.*
16 +
17 +## 1. The founding results
18 +
19 +- Lo, A. W. & MacKinlay, A. C. (1988). Stock Market Prices Do Not Follow Random Walks: Evidence from a Simple Specification Test. *Review of Financial Studies* 1(1), 41–66. https://doi.org/10.1093/rfs/1.1.41 — The variance-ratio test rejects the random walk for weekly index returns 1962–1985 (VR > 1: *momentum* at the index level), explicitly not fully explained by infrequent trading. **Status: robust as a historical fact; the effect itself decayed.**
20 +- Lehmann, B. N. (1990). Fads, Martingales, and Market Efficiency. *Quarterly Journal of Economics* 105(1), 1–28. https://doi.org/10.2307/2937816 — Weekly winner/loser reversals in individual stocks survive his spread corrections. **Status: disputed** — later work attributes much of it to bid-ask bounce and liquidity provision compensation.
21 +- Jegadeesh, N. (1990). Evidence of Predictable Behavior of Security Returns. *Journal of Finance* 45(3), 881–898. https://doi.org/10.1111/j.1540-6261.1990.tb05110.x — Strong negative first-order *monthly* serial correlation in individual stocks. **Status: decayed** post-publication (see McLean–Pontiff below).
22 +
23 +Key structural insight for Q1: at the *index/portfolio* level early evidence
24 +showed VR > 1 (positive autocorrelation), while *individual* stocks showed
25 +reversal. The difference is exactly the cross-autocorrelation structure —
26 +which is where the artifacts live.
27 +
28 +## 2. Lead-lag: the anomaly and its artifact twin
29 +
30 +- Lo, A. W. & MacKinlay, A. C. (1990). When Are Contrarian Profits Due to Stock Market Overreaction? *Review of Financial Studies* 3(2), 175–205. https://doi.org/10.1093/rfs/3.2.175 — Contrarian profits come largely from **lead-lag cross-autocorrelations (large leads small)**, not overreaction. The canonical Q2 result.
31 +- Scholes, M. & Williams, J. (1977). Estimating betas from nonsynchronous data. *Journal of Financial Economics* 5(3), 309–327. https://doi.org/10.1016/0304-405X(77)90041-1 — Non-synchronous observation biases correlations/betas and manufactures spurious lead-lag. **The artifact null for all of Q2** — our expB measured exactly this on hfmarketdata.io (SPY "leads" stale tickers +0.047 at 1min; SPX lags SPY +0.065).
32 +- Epps, T. W. (1979). Comovements in Stock Prices in the Very Short Run. *Journal of the American Statistical Association* 74(366a), 291–298. https://doi.org/10.1080/01621459.1979.10482508 — Cross-correlations shrink as sampling gets finer (the **Epps effect**). At our 1min floor, contemporaneous correlations are mechanically attenuated and the "missing" correlation shows up at leads/lags. Any 1min lead-lag claim must model this.
33 +- Chordia, T., Roll, R. & Subrahmanyam, A. (2005). Evidence on the speed of convergence to market efficiency. *Journal of Financial Economics* 76(2), 271–292. https://doi.org/10.1016/j.jfineco.2004.06.004 — Predictability from order flow is arbitraged away within **5–60 minutes** (already by 2005). Sets the prior: minute-scale inefficiencies in liquid names should be tiny-to-absent in our 2020s data.
34 +
35 +## 3. What decayed
36 +
37 +- McLean, R. D. & Pontiff, J. (2016). Does Academic Research Destroy Stock Return Predictability? *Journal of Finance* 71(1), 5–32. https://doi.org/10.1111/jofi.12365 — Across 97 published predictors: −26 % out-of-sample, −58 % post-publication. **The base rate for our whole project.**
38 +- Chordia, T., Subrahmanyam, A. & Tong, Q. (2014). Have capital market anomalies attenuated in the recent era of high liquidity and trading activity? *Journal of Accounting and Economics* 58(1), 41–58. https://doi.org/10.1016/j.jacceco.2014.06.001 — Anomaly profits attenuate sharply post-decimalization as liquidity/arbitrage grows. Directly relevant: our data (2000→2026) spans this attenuation; sub-period stability checks (expH) are mandatory.
39 +- Jacobs, H. & Müller, S. (2020). Anomalies across the globe: Once public, no longer existent? *Journal of Financial Economics* 135(1), 213–230. https://doi.org/10.1016/j.jfineco.2019.06.004 — Decay is largely a U.S. phenomenon. Our data is U.S.-centric: expect the *fastest* decay regime.
40 +
41 +## 4. Implications for our experiments (expC, expD)
42 +
43 +1. Reversion at 1min in liquid names should be ≈ 0 (bounce-free mega-caps
44 + showed AC1 CI covering 0 in expB) — a *negative finding* here is the
45 + expected, publishable outcome.
46 +2. Any reversion in mid/low liquidity must beat the measured bounce null
47 + (expB: AC1 −0.05 to −0.23 with zero economics).
48 +3. Any lead-lag must beat the staleness-predicted cross-correlation AND
49 + survive on both-fresh subsamples (Scholes–Williams; Epps).
50 +4. Test sub-periods: pre/post 2010 and pre/post publication of the classic
51 + papers; expect attenuation à la Chordia et al.
52 +5. Horizon matters: 1min → daily aggregation sweeps let us find where (if
53 + anywhere) reversion exceeds its artifact floor.
added research/notes/microstructure_artifacts.md +55 −0
@@ -0,0 +1,55 @@
1 +---
2 +project: anomaly-atlas
3 +document: Phase 1 notes — microstructure noise & data artifacts
4 +author: Simon-Pierre Boucher
5 +contact: contact@spboucher.ai
6 +data_source: hfmarketdata.io
7 +created: 2026-08-12
8 +status: reviewed
9 +---
10 +
11 +# Market microstructure & artifacts (§4.3 — feeds the taxonomy)
12 +
13 +*Phase 1 literature notes. Citations verified via OpenAlex, accessed
14 +2026-08-12. Cross-references: research/artifact_taxonomy.md (T1–T7), expB
15 +measured magnitudes.*
16 +
17 +## 1. Bid-ask bounce (taxonomy T1)
18 +
19 +- Roll, R. (1984). A Simple Implicit Measure of the Effective Bid-Ask Spread in an Efficient Market. *Journal of Finance* 39(4), 1127–1139. https://doi.org/10.1111/j.1540-6261.1984.tb03897.x — spread = 2√(−autocov₁). Implemented in `validation/artifacts.py`; expB found it undefined (positive autocov) for mega-caps at 1min — a known Roll limitation, not an anomaly.
20 +- Blume, M. E. & Stambaugh, R. F. (1983). Biases in computed returns: an application to the size effect. *Journal of Financial Economics* 12(3), 387–404. https://doi.org/10.1016/0304-405x(83)90056-9 — bounce in *closing prices* inflates equal-weighted portfolio returns; roughly **halved the size effect**. Warning: portfolio-level statistics inherit ticker-level bounce.
21 +
22 +## 2. Stale prices & non-synchronicity (T2, T3)
23 +
24 +- Fisher, L. (1966). Some New Stock-Market Indexes. *Journal of Business* 39(S1), 191–225. https://doi.org/10.1086/294848 — stale constituent prices induce spurious *index* autocorrelation ("Fisher effect"). Our SPX-lags-SPY measurement (+0.065 at 1min) is this, live.
25 +- Scholes & Williams (1977), Epps (1979), Lo & MacKinlay (1990) — see mean-reversion/lead-lag notes; the artifact-vs-anomaly frontier for Q2.
26 +
27 +## 3. Microstructure noise & realized measures
28 +
29 +- Zhang, L., Mykland, P. A. & Aït-Sahalia, Y. (2005). A Tale of Two Time Scales. *JASA* 100(472), 1394–1411. https://doi.org/10.1198/016214505000000169 — realized variance from finest-scale returns is dominated by noise; two-scales estimator fixes it. For us: 1min realized-vol features must use subsampling or noise correction.
30 +- Aït-Sahalia, Y., Mykland, P. A. & Zhang, L. (2005). How Often to Sample a Continuous-Time Process in the Presence of Market Microstructure Noise. *Review of Financial Studies* 18(2), 351–416. https://doi.org/10.1093/rfs/hhi016 — naive answer: sample sparsely (5min); better: model the noise.
31 +- Hansen, P. R. & Lunde, A. (2006). Realized Variance and Market Microstructure Noise. *JBES* 24(2), 127–161. https://doi.org/10.1198/073500106000000071 — noise is time-varying and **correlated with the efficient price**; iid-noise corrections are themselves biased.
32 +- Bandi, F. M. & Russell, J. R. (2006). Separating microstructure noise from volatility. *Journal of Financial Economics* 79(3), 655–692. https://doi.org/10.1016/j.jfineco.2005.01.005 — moment-based separation of noise vs integrated variance.
33 +- Bandi, F. M. & Russell, J. R. (2008). Microstructure Noise, Realized Variance, and Optimal Sampling. *Review of Economic Studies* 75(2), 339–369. https://doi.org/10.1111/j.1467-937x.2008.00474.x — MSE-optimal sampling frequency.
34 +
35 +## 4. What our dataset adds/changes
36 +
37 +1. We have **bars, not quotes/trades**: all noise corrections must work from
38 + 1min OHLCV. Roll-type estimators (from closes) and high-low estimators
39 + (see transaction-costs notes) are the available instruments.
40 +2. expA's dataset-specific artifacts (auction-close mismatch T4, rolling
41 + adjustment anchor T5, volume-convention mismatch T6, session semantics
42 + T7) are NOT in this classic literature — they are vendor-layer artifacts
43 + and belong to our taxonomy as original documentation.
44 +3. Hansen–Lunde's endogenous-noise warning applies directly to expC:
45 + bounce-null subtraction assumes independence between noise and efficient
46 + price; report sensitivity.
47 +
48 +## 5. Implications
49 +
50 +- Realized-vol-based hypotheses (Q1/Q3 conditioning variables) must use
51 + 5min-subsampled or two-scales estimators, never raw 1min RV.
52 +- Any portfolio-level result needs the Blume–Stambaugh check: recompute with
53 + bounce-robust prices (e.g., mid-of-day anchors) before believing it.
54 +- Index-based lead-lag (SPX family) is presumptively Fisher-effect until
55 + proven otherwise.
added research/notes/multiple_testing_snooping.md +49 −0
@@ -0,0 +1,49 @@
1 +---
2 +project: anomaly-atlas
3 +document: Phase 1 notes — multiple testing, data snooping, backtest overfitting
4 +author: Simon-Pierre Boucher
5 +contact: contact@spboucher.ai
6 +data_source: hfmarketdata.io
7 +created: 2026-08-12
8 +status: reviewed
9 +---
10 +
11 +# The statistics of not fooling yourself (§4.2 — most important)
12 +
13 +*Phase 1 literature notes. Citations verified via OpenAlex, accessed
14 +2026-08-12. This literature IS our methodology (Phase 9): every technique
15 +below maps to a concrete step in expF.*
16 +
17 +## 1. The formal machinery
18 +
19 +- White, H. (2000). A Reality Check for Data Snooping. *Econometrica* 68(5), 1097–1126. https://doi.org/10.1111/1468-0262.00152 — Bootstrap test of "does the BEST rule in my searched universe beat the benchmark?", correcting for the search itself. **The core of expF.**
20 +- Hansen, P. R. (2005). A Test for Superior Predictive Ability. *Journal of Business & Economic Statistics* 23(4), 365–380. https://doi.org/10.1198/073500105000000063 — SPA: studentized, less sensitive to irrelevant alternatives in the universe than White RC. Preferred variant.
21 +- Romano, J. P. & Wolf, M. (2005). Stepwise Multiple Testing as Formalized Data Snooping. *Econometrica* 73(4), 1237–1282. https://doi.org/10.1111/j.1468-0262.2005.00615.x — StepM: identifies *which* rules beat the benchmark with FWE control — not just whether the best one does.
22 +- Benjamini, Y. & Hochberg, Y. (1995). Controlling the False Discovery Rate. *JRSS-B* 57(1), 289–300. https://doi.org/10.1111/j.2517-6161.1995.tb02031.x — FDR: the right error rate for *scans* (expC–E), where we tolerate a known fraction of false leads into the next stage.
23 +- Bailey, D. H. & López de Prado, M. (2014). The Deflated Sharpe Ratio. *Journal of Portfolio Management* 40(5), 94–107. https://doi.org/10.3905/jpm.2014.40.5.094 — DSR: deflate the best Sharpe by the number of trials + non-normality. Requires *honest trial counting* — our LOG's hypothesis budget is exactly that input.
24 +- Bailey, Borwein, López de Prado & Zhu (2014). Pseudo-Mathematics and Financial Charlatanism. *Notices of the AMS* 61(5), 458–471. https://doi.org/10.1090/noti1105 — With enough configurations, a stellar in-sample backtest is *guaranteed*. The quantitative case for pre-registration.
25 +- Bailey, Borwein, López de Prado & Zhu (2016). The probability of backtest overfitting. *Journal of Computational Finance* 20(4), 39–69. https://doi.org/10.21314/jcf.2016.322 — PBO via combinatorially symmetric cross-validation (CSCV) — the basis of CPCV; candidate for expH's split design.
26 +
27 +## 2. Applied to finance's own literature (the base rates)
28 +
29 +- Harvey, C. R., Liu, Y. & Zhu, H. (2016). …and the Cross-Section of Expected Returns. *Review of Financial Studies* 29(1), 5–68. https://doi.org/10.1093/rfs/hhv059 — 300+ published factors; with multiple testing, a new "discovery" needs **t > 3.0**.
30 +- Harvey, C. R. (2017). Presidential Address: The Scientific Outlook in Financial Economics. *Journal of Finance* 72(4), 1399–1440. https://doi.org/10.1111/jofi.12530 — p-hacking incentives; minimum Bayes factors; Bayesianized p-values.
31 +- Chordia, T., Goyal, A. & Saretto, A. (2020). Anomalies and False Rejections. *Review of Financial Studies* 33(5), 2134–2179. https://doi.org/10.1093/rfs/hhaa018 — 2M+ simulated strategies → multiple-testing-safe hurdles **t ≈ 3.4–3.8**; at classic thresholds ~45 % of "anomalies" are false.
32 +- Harvey, C. R. & Liu, Y. (2020). False (and Missed) Discoveries in Financial Economics. *Journal of Finance* 75(5), 2503–2553. https://doi.org/10.1111/jofi.12951 — double bootstrap to pick t-hurdles for a target FDR, balancing Type I *and* Type II.
33 +- Giglio, S., Liao, Y. & Xiu, D. (2021). Thousands of Alpha Tests. *Review of Financial Studies* 34(7), 3456–3496. https://doi.org/10.1093/rfs/hhaa111 — alpha multiple-testing robust to omitted factors/missing data, wild bootstrap.
34 +- Hou, K., Xue & Zhang (2020). Replicating Anomalies. *Review of Financial Studies* 33(5), 2019–2133. https://doi.org/10.1093/rfs/hhy131 — 452 anomalies replicated with microcap-robust methods: **~65 % fail** at |t|≥1.96.
35 +- Jensen, T. I., Kelly, B. & Pedersen, L. H. (2023). Is There a Replication Crisis in Finance? *Journal of Finance* 78(5), 2465–2518. https://doi.org/10.1111/jofi.13249 — the counterpoint: Bayesian hierarchical pooling says most factor *themes* replicate, incl. in 93 countries. Read together with Hou et al.: the answer depends on the unit of analysis (theme vs individual signal) and the prior.
36 +- Sullivan, R., Timmermann, A. & White, H. (2001). Dangers of data mining: the case of calendar effects in stock returns. *Journal of Econometrics* 105(1), 249–286. https://doi.org/10.1016/s0304-4076(01)00077-x — the full calendar-rule universe under the Reality Check: **calendar effects largely vanish**. The single most important prior for expE.
37 +
38 +## 3. Design consequences for anomaly-atlas
39 +
40 +1. **Error-rate ladder**: FDR (BH) for the Level-0 scans → SPA/StepM against
41 + the artifact-null benchmark for Level-1 promotion → DSR with the logged
42 + trial count for anything Sharpe-like.
43 +2. **t-hurdle**: adopt t ≥ 3 (Harvey–Liu–Zhu) as the *floor*, and report the
44 + Chordia–Goyal–Saretto 3.4–3.8 band alongside.
45 +3. **Trial counting is a first-class artifact**: the LOG's hypothesis budget
46 + (charter §12) is the DSR/PBO input; silent hypothesis-space expansion
47 + invalidates the correction — hence append-only logging.
48 +4. **expF deliverable**: the survival curve (naive → FDR → SPA → DSR →
49 + costs → OOS) is finding E of the charter, whatever survives.
added research/notes/timeseries_methodology.md +53 −0
@@ -0,0 +1,53 @@
1 +---
2 +project: anomaly-atlas
3 +document: Phase 1 notes — time-series methodology
4 +author: Simon-Pierre Boucher
5 +contact: contact@spboucher.ai
6 +data_source: hfmarketdata.io
7 +created: 2026-08-12
8 +status: reviewed
9 +---
10 +
11 +# Time-series methodology (§4.4)
12 +
13 +*Phase 1 literature notes. Citations verified via OpenAlex, accessed
14 +2026-08-12.*
15 +
16 +## 1. Variance ratios & long memory
17 +
18 +- Lo & MacKinlay (1988) — the VR test itself (see mean-reversion notes).
19 + Implementation lesson learned in OUR code: the §8.1 synthetic gate caught a
20 + double-division-by-q bug in our first VR implementation — the estimator
21 + read ~1/q on a pure random walk. Estimator subtleties are real risks, not
22 + textbook trivia.
23 +- Lo, A. W. (1991). Long-Term Memory in Stock Market Prices. *Econometrica* 59(5), 1279–1313. https://doi.org/10.2307/2938368 — modified R/S statistic robust to short-range dependence; classic R/S (and naive Hurst estimation) **mistakes short memory + heteroskedasticity for long memory**. Any Hurst-based Q1 claim must use Lo's correction and a bounce-aware null.
24 +
25 +## 2. Granger causality caveats
26 +
27 +- Granger, C. W. J. (1969). Investigating Causal Relations by Econometric Models and Cross-spectral Methods. *Econometrica* 37(3), 424–438. https://doi.org/10.2307/1912791 — predictive content, not causation. For Q2 at 1min: Granger "causality" from fresh to stale series is *guaranteed* by non-synchronicity (expB measured it); only both-fresh subsamples and staleness-matched nulls make the test meaningful.
28 +
29 +## 3. Bootstrap for dependent data
30 +
31 +- Künsch, H. R. (1989). The Jackknife and the Bootstrap for General Stationary Observations. *Annals of Statistics* 17(3), 1217–1241. https://doi.org/10.1214/aos/1176347265 — moving-block bootstrap (our `stats/bootstrap.py`).
32 +- Politis, D. N. & Romano, J. P. (1994). The Stationary Bootstrap. *JASA* 89(428), 1303–1313. https://doi.org/10.1080/01621459.1994.10476870 — geometric random block lengths → stationary resamples; the resampling engine inside White's Reality Check. To implement for expF.
33 +
34 +## 4. Splits, walk-forward, CPCV
35 +
36 +- Bailey, Borwein, López de Prado & Zhu (2016). The probability of backtest overfitting. *Journal of Computational Finance* 20(4), 39–69. https://doi.org/10.21314/jcf.2016.322 — CSCV/PBO: combinatorial splits measure how often the in-sample winner underperforms out-of-sample. Candidate for expH; must be combined with *purging* (no leakage across split boundaries — overlapping bars/labels).
37 +- Charter constraint: the final holdout is touched ONCE (§8.2) — CPCV
38 + operates strictly inside the train/validation region.
39 +
40 +## 5. Structural breaks & regimes
41 +
42 +- Bai, J. & Perron, P. (1998). Estimating and Testing Linear Models with Multiple Structural Changes. *Econometrica* 66(1), 47–78. https://doi.org/10.2307/2998540 — multiple unknown breakpoints. Relevance: 2000–2026 spans decimalization aftermath, Reg NMS (2007), the 2008 crisis, HFT rise, 2020 COVID, T+1 (2024). An "anomaly" that is really one regime's plumbing (e.g., pre-2010 latency) must be caught by sub-period analysis (expH), and Bai–Perron gives the formal tool.
43 +
44 +## 6. Methodological rules adopted (feed Phase 9)
45 +
46 +1. Every test statistic ships with a block/stationary-bootstrap CI, block
47 + length ≥ one trading day for intraday data.
48 +2. Hurst/long-memory claims: Lo (1991) modified R/S only, with bounce and
49 + staleness nulls.
50 +3. Granger tests only on both-fresh subsamples with staleness-matched nulls.
51 +4. Sub-period grid pre-specified: 2000–07 / 2008–14 / 2015–19 / 2020–26 +
52 + Bai–Perron endogenous breaks as robustness.
53 +5. CPCV inside train/validation; single-touch holdout untouched until expH.
added research/notes/transaction_costs.md +49 −0
@@ -0,0 +1,49 @@
1 +---
2 +project: anomaly-atlas
3 +document: Phase 1 notes — transaction-cost realism
4 +author: Simon-Pierre Boucher
5 +contact: contact@spboucher.ai
6 +data_source: hfmarketdata.io
7 +created: 2026-08-12
8 +status: reviewed
9 +---
10 +
11 +# Transaction-cost realism (§4.5 — Q5, the cost frontier)
12 +
13 +*Phase 1 literature notes. Citations verified via OpenAlex, accessed
14 +2026-08-12. Constraint: our data has NO quotes — every spread must be
15 +estimated from OHLCV bars. That makes the low-frequency-estimator literature
16 +load-bearing.*
17 +
18 +## 1. Spread estimation from bar data (our only instruments)
19 +
20 +- Roll (1984) — see microstructure notes; autocovariance-based, undefined when autocov ≥ 0 (expB: mega-caps).
21 +- Hasbrouck, J. (2009). Trading Costs and Returns for U.S. Equities: Estimating Effective Costs from Daily Data. *Journal of Finance* 64(3), 1445–1477. https://doi.org/10.1111/j.1540-6261.2009.01469.x — Gibbs-sampled Roll model; corr 0.965 with TAQ benchmarks.
22 +- Corwin, S. A. & Schultz, P. (2012). A Simple Way to Estimate Bid-Ask Spreads from Daily High and Low Prices. *Journal of Finance* 67(2), 719–760. https://doi.org/10.1111/j.1540-6261.2012.01729.x — high-low estimator; we have highs/lows at every timeframe.
23 +- Abdi, F. & Ranaldo, A. (2017). A Simple Estimation of Bid-Ask Spreads from Daily Close, High, and Low Prices. *Review of Financial Studies* 30(12), 4437–4480. https://doi.org/10.1093/rfs/hhx084 — CHL estimator; best for illiquid names.
24 +- Ardia, D., Guidotti, E. & Kroencke, T. A. (2024). Efficient estimation of bid–ask spreads from open, high, low, and close prices. *Journal of Financial Economics* 161, 103916. https://doi.org/10.1016/j.jfineco.2024.103916 — state-of-the-art OHLC estimator (EDGE), asymptotically unbiased, minimal variance. **Primary estimator for expG**; Roll/CS/CHL as cross-checks.
25 +- Fong, K. Y. L., Holden, C. W. & Trzcinka, C. A. (2017). What Are the Best Liquidity Proxies for Global Research? *Review of Finance* 21(4), 1355–1401. https://doi.org/10.1093/rof/rfx003 — horse race of low-frequency proxies vs intraday benchmarks; validates picking 2–3 complementary proxies.
26 +- Lesmond, D. A., Ogden, J. P. & Trzcinka, C. A. (1999). A New Estimate of Transaction Costs. *Review of Financial Studies* 12(5), 1113–1141. https://doi.org/10.1093/rfs/12.5.1113 — LOT: costs from the incidence of zero returns (1.2 %–10.3 % across deciles). Our no-trade minutes are the intraday analogue.
27 +- Bessembinder, H. (2003). Issues in Assessing Trade Execution Costs. *Journal of Financial Markets* 6(3), 233–257. https://doi.org/10.1016/s1386-4181(02)00064-2 — measurement pitfalls (trade signing, timing conventions) — reminder that even "measured" costs carry convention risk.
28 +
29 +## 2. Do anomalies survive costs?
30 +
31 +- Novy-Marx, R. & Velikov, M. (2016). A Taxonomy of Anomalies and Their Trading Costs. *Review of Financial Studies* 29(1), 104–147. https://doi.org/10.1093/rfs/hhv063 — low-turnover anomalies mostly survive (with mitigation); **high-turnover ones mostly do not**. Short-horizon effects (ours) are the highest-turnover class → strong prior that Q1/Q2 effects die at the cost frontier.
32 +- Frazzini, A., Israel, R. & Moskowitz, T. J. (2012). Trading Costs of Asset Pricing Anomalies. SSRN. https://doi.org/10.2139/ssrn.2294498 — real institutional executions: realized costs are far below TAQ-implied for patient flow. Gives the *lower* bound of the cost sweep.
33 +- Chen, A. Y. & Velikov, M. (2022). Zeroing In on the Expected Returns of Anomalies. *Journal of Financial and Quantitative Analysis* 58(3), 968–1004. https://doi.org/10.1017/s0022109022000874 — net of spreads + post-publication decay + data mining, the average anomaly earns **~4 bp/month**. The sobering calibration for our whole atlas.
34 +- Detzel, A., Novy-Marx, R. & Velikov, M. (2023). Model Comparison with Transaction Costs. *Journal of Finance* 78(3), 1743–1775. https://doi.org/10.1111/jofi.13225 — ignoring costs biases even *model comparisons*; cost-awareness is not optional at any stage.
35 +
36 +## 3. Design of expG (cost frontier)
37 +
38 +1. **Spread panel**: EDGE (Ardia et al.) per ticker-month from daily OHLC +
39 + Corwin–Schultz and CHL cross-checks + Roll where defined; validate the
40 + three against each other (Fong et al. protocol).
41 +2. **Options as auxiliary evidence**: our options chains carry real bid/ask —
42 + the only quoted spreads in the dataset; usable as a sanity anchor for the
43 + underlying's cost regime (with care).
44 +3. **Sweep, don't point-estimate**: report each surviving effect's net value
45 + across cost multipliers 0.25×–2× the estimated half-spread (Frazzini
46 + lower bound ↔ retail-taker upper bound), and the crossing point where net
47 + effect = 0.
48 +4. **Capacity is out of scope** (no volume-at-quote data) — state it as a
49 + limitation rather than pretend.
50