Data section with real panel facts, summary table, data figures; MCS memory fix; parallel RV build
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
10 changed files +220 −14
modified
.gitignore
+1 −0
@@ -1,6 +1,7 @@ | ||
| 1 | 1 | # Raw & processed data — downloaded market data are NOT committed |
| 2 | 2 | data/raw/*.parquet |
| 3 | 3 | data/raw/*.csv |
| 4 | +data/processed/**/*.parquet | |
| 4 | 5 | data/processed/*.parquet |
| 5 | 6 | |
| 6 | 7 | # Python |
added
figures/fig_acf.png
+0 −0
Binary file not shown.
added
figures/fig_measures.png
+0 −0
Binary file not shown.
added
figures/fig_rv_series.png
+0 −0
Binary file not shown.
added
paper/main.pdf
+0 −0
Binary file not shown.
modified
paper/sections/data.tex
+130 −1
@@ -4,5 +4,134 @@ | ||
| 4 | 4 | % Projet : Prévision de volatilité réalisée multi-actifs |
| 5 | 5 | % (HAR-RV vs GARCH vs Machine Learning) |
| 6 | 6 | % Fichier : data.tex |
| 7 | −% Description : Section « data » — à rédiger après exécution du pipeline. | |
| 7 | +% Description : Section 3 — Données, univers, statistiques | |
| 8 | +% descriptives, microstructure et spécificités. | |
| 8 | 9 | % ================================================================ |
| 10 | +\section{Data} | |
| 11 | +\label{sec:data} | |
| 12 | + | |
| 13 | +\subsection{Source and universe} | |
| 14 | + | |
| 15 | +All price data come from the HF Market Data API | |
| 16 | +(\url{https://www.hfmarketdata.io}), an open high-frequency data service | |
| 17 | +providing intraday and daily OHLCV bars for equities, ETFs, continuous | |
| 18 | +futures, cryptocurrencies, indices and FX. We download 1-minute and | |
| 19 | +daily bars from January 2013 through July 2026 for a universe of 25 | |
| 20 | +forecast targets spanning four asset classes, plus the VIX as a | |
| 21 | +predictor: | |
| 22 | + | |
| 23 | +\begin{itemize} | |
| 24 | +\item \emph{Equities and ETFs} (10): SPY, QQQ, IWM, GLD, TLT, AAPL, | |
| 25 | + MSFT, NVDA, JPM, XOM --- split- and dividend-adjusted; | |
| 26 | +\item \emph{Foreign exchange} (5): EURUSD, USDJPY, GBPUSD, USDCAD, | |
| 27 | + AUDUSD; | |
| 28 | +\item \emph{Cryptocurrencies} (5): BTC, ETH, SOL, ADA, XRP;\footnote{The | |
| 29 | + originally targeted BNB is not available in the data source and is | |
| 30 | + replaced by ADA. ETH, XRP, ADA and SOL enter the sample at their | |
| 31 | + first available dates (May 2016, March 2017, January 2019 and April | |
| 32 | + 2021, respectively); BTC is available from April 2013.} | |
| 33 | +\item \emph{Futures} (5): ES (S\&P 500), NQ (Nasdaq 100), CL (WTI | |
| 34 | + crude), GC (gold), ZN (10-year Treasury note) --- continuous | |
| 35 | + front-month series, ratio back-adjusted so that roll dates do not | |
| 36 | + generate spurious returns. | |
| 37 | +\end{itemize} | |
| 38 | + | |
| 39 | +In total the pipeline processes roughly 94 million 1-minute bars and | |
| 40 | +produces a panel of 84{,}898 asset-day observations of realized | |
| 41 | +measures. A reusable API client handles rate limiting, exponential | |
| 42 | +back-off retries, pagination and a local parquet cache, making the | |
| 43 | +entire pipeline reproducible end to end from raw bars. | |
| 44 | + | |
| 45 | +\subsection{Trading sessions and data handling} | |
| 46 | + | |
| 47 | +Timestamps are in exchange-local (US Eastern) time. For equities and | |
| 48 | +ETFs, whose 1-minute bars cover the extended session (04:00--20:00), we | |
| 49 | +retain regular trading hours only (09:30--16:00) when computing | |
| 50 | +realized measures, so that RV is not contaminated by thin pre- and | |
| 51 | +post-market trading. FX trades essentially continuously from Sunday | |
| 52 | +17:00 to Friday 17:00; its trading day is the calendar day, and | |
| 53 | +weekend fragments are eliminated by the minimum-bar filter. Futures | |
| 54 | +trade nearly 24 hours with short maintenance breaks; their day is | |
| 55 | +likewise the calendar day. Crypto trades continuously; its day is the | |
| 56 | +calendar day in UTC, and weekends are genuine trading days --- crypto | |
| 57 | +therefore contributes roughly 7/5 as many observations per calendar | |
| 58 | +year as the session-limited markets. Days with fewer than 700 valid | |
| 59 | +1-minute bars (250 for the 6.5-hour equity session) --- holidays, | |
| 60 | +half-days, and outage days --- are discarded. Missing minutes within a | |
| 61 | +day (no trade) are handled by computing returns between consecutive | |
| 62 | +observed bars, and the 5-minute grids are built on observed prices, so | |
| 63 | +no artificial zero returns are introduced. | |
| 64 | + | |
| 65 | +\subsection{Descriptive statistics} | |
| 66 | + | |
| 67 | +Table~\ref{tab:summary_stats} reports per-instrument descriptive | |
| 68 | +statistics of the daily realized measures; Figure~\ref{fig:rv_series} | |
| 69 | +plots the annualized RV series for one representative asset per class, | |
| 70 | +and Figure~\ref{fig:acf} the autocorrelation functions of log RV. Three | |
| 71 | +features of the data matter for the forecasting exercise. First, the | |
| 72 | +levels differ enormously across classes: mean annualized volatility is | |
| 73 | +7.9\% for FX, 15.9\% for equities/ETFs, 16.8\% for futures, and 108\% | |
| 74 | +for crypto (median 68.9\%) --- any pooled model must absorb these level | |
| 75 | +differences before it can share dynamics. Second, persistence is | |
| 76 | +uniformly high --- the first autocorrelation of log RV averages | |
| 77 | +0.70--0.82 across classes --- and decays slowly | |
| 78 | +(Figure~\ref{fig:acf}), the signature of long memory that motivates the | |
| 79 | +HAR cascade. Third, jump days identified by the BNS test at | |
| 80 | +$\alpha = 0.1\%$ range from 3.7\% of days for equities to 14.8\% for | |
| 81 | +crypto, so jump-robust measures and jump-aware models are potentially | |
| 82 | +material outside equities. | |
| 83 | + | |
| 84 | +\begin{table}[t] | |
| 85 | +\centering | |
| 86 | +\caption{Descriptive statistics of daily realized volatility by | |
| 87 | +instrument. Volatility columns annualize the subsampled 5-minute RV | |
| 88 | +($\sqrt{252 \cdot RV}$, in percent). Skewness and first-order | |
| 89 | +autocorrelation refer to $\ln RV$. ``Jump days'' is the share of days | |
| 90 | +with a significant BNS jump statistic at $\alpha = 0.1\%$.} | |
| 91 | +\label{tab:summary_stats} | |
| 92 | +{\small \input{../results/tables/summary_stats}} | |
| 93 | +\end{table} | |
| 94 | + | |
| 95 | +\begin{figure}[t] | |
| 96 | +\centering | |
| 97 | +\includegraphics[width=\textwidth]{fig_rv_series.png} | |
| 98 | +\caption{Annualized realized volatility (subsampled 5-minute RV), one | |
| 99 | +representative instrument per asset class, 2013--2026.} | |
| 100 | +\label{fig:rv_series} | |
| 101 | +\end{figure} | |
| 102 | + | |
| 103 | +\subsection{Microstructure noise across measures} | |
| 104 | + | |
| 105 | +Figure~\ref{fig:measures} compares yearly averages of 1-minute RV, | |
| 106 | +subsampled 5-minute RV and the realized kernel for SPY and BTC. For SPY | |
| 107 | +the three measures are nearly indistinguishable --- at SPY's liquidity, | |
| 108 | +1-minute sampling carries little noise, and the pairwise correlations | |
| 109 | +between measures exceed 0.98. For BTC the ordering | |
| 110 | +$RV^{(1\mathrm{min})} > RV^{ss} > RK$ in the early sample (2013--2015) | |
| 111 | +is exactly the pattern implied by i.i.d.\ microstructure noise, whose | |
| 112 | +variance contribution grows linearly in the sampling frequency: early | |
| 113 | +Bitcoin trading was thin and fragmented, and plain 1-minute RV | |
| 114 | +overstates variance by a factor of two or more relative to the | |
| 115 | +noise-robust kernel. The gap closes progressively as crypto market | |
| 116 | +quality improves after 2017. This pattern justifies our choice of the | |
| 117 | +subsampled 5-minute RV as the headline measure --- following the | |
| 118 | +multi-asset evidence of \citet{LiuPattonSheppard2015} --- with the | |
| 119 | +1-minute RV and the realized kernel used to assess robustness to the | |
| 120 | +sampling scheme in Section~\ref{sec:robustness}. | |
| 121 | + | |
| 122 | +\begin{figure}[t] | |
| 123 | +\centering | |
| 124 | +\includegraphics[width=0.95\textwidth]{fig_measures.png} | |
| 125 | +\caption{Yearly mean annualized volatility by realized measure. The | |
| 126 | +divergence between the 1-minute RV and the noise-robust measures for | |
| 127 | +early-sample BTC reflects market microstructure noise.} | |
| 128 | +\label{fig:measures} | |
| 129 | +\end{figure} | |
| 130 | + | |
| 131 | +\begin{figure}[t] | |
| 132 | +\centering | |
| 133 | +\includegraphics[width=0.8\textwidth]{fig_acf.png} | |
| 134 | +\caption{Autocorrelation functions of daily $\ln RV$ (subsampled | |
| 135 | +5-minute measure), representative instruments.} | |
| 136 | +\label{fig:acf} | |
| 137 | +\end{figure} | |
added
results/tables/summary_stats.tex
+35 −0
@@ -0,0 +1,35 @@ | ||
| 1 | +\begin{tabular}{lrrrrrrr} | |
| 2 | +\toprule | |
| 3 | +Ticker & Days & \makecell{Mean vol.\\(\%)} & \makecell{Median vol.\\(\%)} & \makecell{P95 vol.\\(\%)} & \makecell{Skew.\\$\ln$RV} & \makecell{AC(1)\\$\ln$RV} & \makecell{Jump\\days (\%)} \\ | |
| 4 | +\midrule | |
| 5 | +\multicolumn{8}{l}{\itshape Equities/ETFs}\\ | |
| 6 | +\quad AAPL & 3,413 & 18.8 & 16.6 & 34.2 & 0.55 & 0.73 & 4.6 \\ | |
| 7 | +\quad GLD & 3,396 & 8.8 & 7.8 & 16.4 & 0.65 & 0.59 & 3.9 \\ | |
| 8 | +\quad IWM & 3,405 & 15.3 & 13.3 & 28.5 & 0.77 & 0.77 & 2.5 \\ | |
| 9 | +\quad JPM & 3,388 & 17.6 & 15.6 & 30.7 & 0.98 & 0.74 & 4.6 \\ | |
| 10 | +\quad MSFT & 3,397 & 17.8 & 15.9 & 32.0 & 0.65 & 0.75 & 4.7 \\ | |
| 11 | +\quad NVDA & 3,401 & 30.3 & 26.4 & 57.0 & 0.60 & 0.75 & 4.4 \\ | |
| 12 | +\quad QQQ & 3,410 & 13.4 & 11.3 & 27.2 & 0.47 & 0.78 & 2.8 \\ | |
| 13 | +\quad SPY & 3,414 & 10.2 & 8.4 & 21.3 & 0.53 & 0.79 & 2.8 \\ | |
| 14 | +\quad TLT & 3,394 & 8.5 & 7.7 & 15.0 & 0.73 & 0.70 & 2.8 \\ | |
| 15 | +\quad XOM & 3,389 & 18.0 & 15.7 & 33.6 & 0.59 & 0.84 & 4.3 \\ | |
| 16 | +\multicolumn{8}{l}{\itshape FX}\\ | |
| 17 | +\quad AUDUSD & 3,521 & 9.6 & 8.9 & 16.1 & 0.41 & 0.68 & 8.9 \\ | |
| 18 | +\quad EURUSD & 3,522 & 7.2 & 6.6 & 12.8 & 0.30 & 0.69 & 7.7 \\ | |
| 19 | +\quad GBPUSD & 3,521 & 8.1 & 7.4 & 13.5 & 0.65 & 0.70 & 7.7 \\ | |
| 20 | +\quad USDCAD & 3,521 & 6.7 & 6.2 & 11.4 & 0.24 & 0.69 & 6.3 \\ | |
| 21 | +\quad USDJPY & 3,522 & 7.9 & 7.0 & 14.7 & 0.24 & 0.72 & 10.2 \\ | |
| 22 | +\multicolumn{8}{l}{\itshape Crypto}\\ | |
| 23 | +\quad ADA & 2,750 & 106.8 & 81.4 & 221.7 & -2.75 & 0.65 & 25.2 \\ | |
| 24 | +\quad BTC & 4,498 & 97.8 & 50.3 & 372.7 & 0.95 & 0.90 & 11.8 \\ | |
| 25 | +\quad ETH & 3,528 & 108.4 & 75.2 & 289.7 & 0.46 & 0.88 & 12.1 \\ | |
| 26 | +\quad SOL & 1,917 & 81.4 & 69.7 & 162.2 & 0.53 & 0.76 & 5.9 \\ | |
| 27 | +\quad XRP & 3,373 & 137.3 & 77.2 & 478.6 & 0.85 & 0.90 & 18.3 \\ | |
| 28 | +\multicolumn{8}{l}{\itshape Futures}\\ | |
| 29 | +\quad CL & 3,502 & 33.4 & 28.3 & 64.7 & 0.78 & 0.87 & 10.5 \\ | |
| 30 | +\quad ES & 3,501 & 13.4 & 10.9 & 28.3 & 0.60 & 0.81 & 6.6 \\ | |
| 31 | +\quad GC & 3,501 & 14.1 & 12.5 & 25.1 & 0.71 & 0.70 & 8.9 \\ | |
| 32 | +\quad NQ & 3,498 & 16.9 & 14.2 & 35.2 & 0.38 & 0.80 & 6.3 \\ | |
| 33 | +\quad ZN & 3,216 & 5.2 & 4.6 & 9.3 & 0.82 & 0.68 & 19.5 \\ | |
| 34 | +\bottomrule | |
| 35 | +\end{tabular} | |
modified
scripts/02_build_rv.py
+42 −7
@@ -7,8 +7,10 @@ Projet : Prévision de volatilité réalisée multi-actifs | ||
| 7 | 7 | (HAR-RV vs GARCH vs Machine Learning) |
| 8 | 8 | Fichier : 02_build_rv.py |
| 9 | 9 | Description : Étape 02 — Construction des mesures de volatilité |
| 10 | − réalisée journalières et du panel multi-actifs. | |
| 11 | −Usage : python scripts/02_build_rv.py [--refresh] | |
| 10 | + réalisée journalières (parallélisée par actif) et | |
| 11 | + assemblage du panel multi-actifs. | |
| 12 | +Usage : python scripts/02_build_rv.py [--refresh] [--workers 8] | |
| 13 | + [--only-cached] | |
| 12 | 14 | ================================================================ |
| 13 | 15 | """ |
| 14 | 16 | |
@@ -17,26 +19,59 @@ from __future__ import annotations | ||
| 17 | 19 | import argparse |
| 18 | 20 | import logging |
| 19 | 21 | import sys |
| 22 | +from concurrent.futures import ProcessPoolExecutor, as_completed | |
| 20 | 23 | from pathlib import Path |
| 21 | 24 | |
| 25 | +import pandas as pd | |
| 26 | + | |
| 22 | 27 | sys.path.insert(0, str(Path(__file__).resolve().parents[1] / "src")) |
| 23 | 28 | |
| 24 | −from wp12 import config, data_pipeline # noqa: E402 | |
| 29 | +from wp12 import config # noqa: E402 | |
| 25 | 30 | |
| 26 | 31 | logger = logging.getLogger("02_build_rv") |
| 27 | 32 | |
| 28 | 33 | |
| 34 | +def build_one(inst: dict[str, str], refresh: bool) -> str: | |
| 35 | + """Worker: build the daily RV file for one instrument.""" | |
| 36 | + from wp12 import data_pipeline | |
| 37 | + from wp12.api_client import HFMarketDataClient | |
| 38 | + | |
| 39 | + config.setup_logging(logging.WARNING) | |
| 40 | + client = HFMarketDataClient() | |
| 41 | + df = data_pipeline.build_instrument_rv(client, inst, refresh=refresh) | |
| 42 | + return f"{inst['ticker']}: {len(df)} days" | |
| 43 | + | |
| 44 | + | |
| 29 | 45 | def main() -> None: |
| 30 | − """Build daily realized measures for the universe and save the panel.""" | |
| 46 | + """Build per-ticker RV files in parallel, then assemble the panel.""" | |
| 31 | 47 | ap = argparse.ArgumentParser() |
| 32 | 48 | ap.add_argument("--refresh", action="store_true", help="recompute RV files") |
| 49 | + ap.add_argument("--workers", type=int, default=8) | |
| 50 | + ap.add_argument("--only-cached", action="store_true", | |
| 51 | + help="only tickers whose 1-min bars are already cached") | |
| 33 | 52 | args = ap.parse_args() |
| 34 | 53 | |
| 35 | 54 | config.setup_logging() |
| 36 | 55 | logging.getLogger("urllib3").setLevel(logging.WARNING) |
| 37 | − panel = data_pipeline.build_panel(refresh=args.refresh) | |
| 38 | − logger.info("panel summary:\n%s", | |
| 39 | − panel.groupby("ticker")["rv5ss"].agg(["count", "mean"])) | |
| 56 | + univ = config.universe(include_predictors=True) | |
| 57 | + if args.only_cached: | |
| 58 | + raw = config.path("raw") | |
| 59 | + univ = [i for i in univ | |
| 60 | + if (raw / f"{i['asset']}_{i['ticker']}_1min.parquet").exists()] | |
| 61 | + logger.info("building RV for %d instruments", len(univ)) | |
| 62 | + with ProcessPoolExecutor(max_workers=args.workers) as pool: | |
| 63 | + futs = {pool.submit(build_one, i, args.refresh): i["ticker"] for i in univ} | |
| 64 | + for fut in as_completed(futs): | |
| 65 | + try: | |
| 66 | + logger.info("done %s", fut.result()) | |
| 67 | + except Exception as exc: # noqa: BLE001 | |
| 68 | + logger.error("FAILED %s: %s", futs[fut], exc) | |
| 69 | + | |
| 70 | + if not args.only_cached: | |
| 71 | + from wp12 import data_pipeline | |
| 72 | + panel = data_pipeline.build_panel() | |
| 73 | + logger.info("panel: %d rows, %d tickers", len(panel), | |
| 74 | + panel["ticker"].nunique()) | |
| 40 | 75 | |
| 41 | 76 | |
| 42 | 77 | if __name__ == "__main__": |
modified
scripts/07_make_figures.py
+5 −2
@@ -59,11 +59,14 @@ def fig_rv_series() -> None: | ||
| 59 | 59 | fig, axes = plt.subplots(4, 1, figsize=(TEXTWIDTH, 6.6), sharex=True) |
| 60 | 60 | for ax, (cls, tk) in zip(axes, REPRESENTATIVE.items()): |
| 61 | 61 | df = panel[panel["ticker"] == tk].sort_index() |
| 62 | − ax.plot(df.index, annualized_vol(df["rv5ss"]), lw=0.4, color=BLUE) | |
| 62 | + vol = annualized_vol(df["rv5ss"]) | |
| 63 | + ax.plot(df.index, vol, lw=0.4, color=BLUE) | |
| 63 | 64 | panel_label(ax, f"Panel {'ABCD'[list(REPRESENTATIVE).index(cls)]}. " |
| 64 | 65 | f"{tk} ({CLS_LABEL[cls]})") |
| 65 | 66 | ax.set_ylabel("Ann. vol. (%)") |
| 66 | − ax.set_ylim(bottom=0) | |
| 67 | + # cap the axis at the 99.5th percentile so isolated early-sample | |
| 68 | + # noise spikes (e.g. 2013 BTC) do not compress the whole panel | |
| 69 | + ax.set_ylim(0, float(vol.quantile(0.995)) * 1.1) | |
| 67 | 70 | axes[-1].set_xlabel("") |
| 68 | 71 | fig.tight_layout() |
| 69 | 72 | fig.savefig(FIG / "fig_rv_series.png") |
modified
src/wp12/evaluation.py
+7 −4
@@ -150,17 +150,20 @@ def model_confidence_set( | ||
| 150 | 150 | n = len(L) |
| 151 | 151 | rng = np.random.default_rng(seed) |
| 152 | 152 | idx = _stationary_bootstrap_idx(n, n_boot, avg_block, rng) |
| 153 | + # bootstrap means per model, computed once column-by-column (memory-safe) | |
| 154 | + all_boot_means = np.empty((n_boot, len(models))) | |
| 155 | + for j in range(len(models)): | |
| 156 | + all_boot_means[:, j] = L[idx, j].mean(axis=1) | |
| 157 | + all_dbar = L.mean(axis=0) | |
| 153 | 158 | |
| 154 | 159 | included = list(range(len(models))) |
| 155 | 160 | pvals: dict[str, float] = {} |
| 156 | 161 | running_max_p = 0.0 |
| 157 | 162 | while len(included) > 1: |
| 158 | − Ls = L[:, included] | |
| 159 | − dbar_i = Ls.mean(axis=0) # mean loss per model | |
| 163 | + dbar_i = all_dbar[included] # mean loss per model | |
| 160 | 164 | d_i_dot = dbar_i - dbar_i.mean() # vs. set average |
| 161 | 165 | # bootstrap distribution of the max t-stat |
| 162 | − boot_means = Ls[idx].mean(axis=1) # (b, k) | |
| 163 | − boot_center = boot_means - dbar_i | |
| 166 | + boot_center = all_boot_means[:, included] - dbar_i | |
| 164 | 167 | var_i = (boot_center - boot_center.mean(axis=0)).var(axis=0) + EPS |
| 165 | 168 | t_i = d_i_dot / np.sqrt(var_i) |
| 166 | 169 | t_max = float(t_i.max()) |
| 167 | 170 | |