SPB Git forge

spb/wp12_uqo

Public
5commits 1branches 0releases
1.2 MBsize
maindefault branch
1 mo agolast push
Python 75.9% TeX 24%

Data section with real panel facts, summary table, data figures; MCS memory fix; parallel RV build

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Simon-Pierre Boucher committed 1 mo ago (Aug 11, 2026) parent 72b4e9e

10 changed files +220 −14

modified .gitignore +1 −0
@@ -1,6 +1,7 @@
1 1 # Raw & processed data — downloaded market data are NOT committed
2 2 data/raw/*.parquet
3 3 data/raw/*.csv
4 +data/processed/**/*.parquet
4 5 data/processed/*.parquet
5 6
6 7 # Python
added figures/fig_acf.png +0 −0

Binary file not shown.

added figures/fig_measures.png +0 −0

Binary file not shown.

added figures/fig_rv_series.png +0 −0

Binary file not shown.

added paper/main.pdf +0 −0

Binary file not shown.

modified paper/sections/data.tex +130 −1
@@ -4,5 +4,134 @@
4 4 % Projet : Prévision de volatilité réalisée multi-actifs
5 5 % (HAR-RV vs GARCH vs Machine Learning)
6 6 % Fichier : data.tex
7 −% Description : Section « data » — à rédiger après exécution du pipeline.
7 +% Description : Section 3 — Données, univers, statistiques
8 +% descriptives, microstructure et spécificités.
8 9 % ================================================================
10 +\section{Data}
11 +\label{sec:data}
12 +
13 +\subsection{Source and universe}
14 +
15 +All price data come from the HF Market Data API
16 +(\url{https://www.hfmarketdata.io}), an open high-frequency data service
17 +providing intraday and daily OHLCV bars for equities, ETFs, continuous
18 +futures, cryptocurrencies, indices and FX. We download 1-minute and
19 +daily bars from January 2013 through July 2026 for a universe of 25
20 +forecast targets spanning four asset classes, plus the VIX as a
21 +predictor:
22 +
23 +\begin{itemize}
24 +\item \emph{Equities and ETFs} (10): SPY, QQQ, IWM, GLD, TLT, AAPL,
25 + MSFT, NVDA, JPM, XOM --- split- and dividend-adjusted;
26 +\item \emph{Foreign exchange} (5): EURUSD, USDJPY, GBPUSD, USDCAD,
27 + AUDUSD;
28 +\item \emph{Cryptocurrencies} (5): BTC, ETH, SOL, ADA, XRP;\footnote{The
29 + originally targeted BNB is not available in the data source and is
30 + replaced by ADA. ETH, XRP, ADA and SOL enter the sample at their
31 + first available dates (May 2016, March 2017, January 2019 and April
32 + 2021, respectively); BTC is available from April 2013.}
33 +\item \emph{Futures} (5): ES (S\&P 500), NQ (Nasdaq 100), CL (WTI
34 + crude), GC (gold), ZN (10-year Treasury note) --- continuous
35 + front-month series, ratio back-adjusted so that roll dates do not
36 + generate spurious returns.
37 +\end{itemize}
38 +
39 +In total the pipeline processes roughly 94 million 1-minute bars and
40 +produces a panel of 84{,}898 asset-day observations of realized
41 +measures. A reusable API client handles rate limiting, exponential
42 +back-off retries, pagination and a local parquet cache, making the
43 +entire pipeline reproducible end to end from raw bars.
44 +
45 +\subsection{Trading sessions and data handling}
46 +
47 +Timestamps are in exchange-local (US Eastern) time. For equities and
48 +ETFs, whose 1-minute bars cover the extended session (04:00--20:00), we
49 +retain regular trading hours only (09:30--16:00) when computing
50 +realized measures, so that RV is not contaminated by thin pre- and
51 +post-market trading. FX trades essentially continuously from Sunday
52 +17:00 to Friday 17:00; its trading day is the calendar day, and
53 +weekend fragments are eliminated by the minimum-bar filter. Futures
54 +trade nearly 24 hours with short maintenance breaks; their day is
55 +likewise the calendar day. Crypto trades continuously; its day is the
56 +calendar day in UTC, and weekends are genuine trading days --- crypto
57 +therefore contributes roughly 7/5 as many observations per calendar
58 +year as the session-limited markets. Days with fewer than 700 valid
59 +1-minute bars (250 for the 6.5-hour equity session) --- holidays,
60 +half-days, and outage days --- are discarded. Missing minutes within a
61 +day (no trade) are handled by computing returns between consecutive
62 +observed bars, and the 5-minute grids are built on observed prices, so
63 +no artificial zero returns are introduced.
64 +
65 +\subsection{Descriptive statistics}
66 +
67 +Table~\ref{tab:summary_stats} reports per-instrument descriptive
68 +statistics of the daily realized measures; Figure~\ref{fig:rv_series}
69 +plots the annualized RV series for one representative asset per class,
70 +and Figure~\ref{fig:acf} the autocorrelation functions of log RV. Three
71 +features of the data matter for the forecasting exercise. First, the
72 +levels differ enormously across classes: mean annualized volatility is
73 +7.9\% for FX, 15.9\% for equities/ETFs, 16.8\% for futures, and 108\%
74 +for crypto (median 68.9\%) --- any pooled model must absorb these level
75 +differences before it can share dynamics. Second, persistence is
76 +uniformly high --- the first autocorrelation of log RV averages
77 +0.70--0.82 across classes --- and decays slowly
78 +(Figure~\ref{fig:acf}), the signature of long memory that motivates the
79 +HAR cascade. Third, jump days identified by the BNS test at
80 +$\alpha = 0.1\%$ range from 3.7\% of days for equities to 14.8\% for
81 +crypto, so jump-robust measures and jump-aware models are potentially
82 +material outside equities.
83 +
84 +\begin{table}[t]
85 +\centering
86 +\caption{Descriptive statistics of daily realized volatility by
87 +instrument. Volatility columns annualize the subsampled 5-minute RV
88 +($\sqrt{252 \cdot RV}$, in percent). Skewness and first-order
89 +autocorrelation refer to $\ln RV$. ``Jump days'' is the share of days
90 +with a significant BNS jump statistic at $\alpha = 0.1\%$.}
91 +\label{tab:summary_stats}
92 +{\small \input{../results/tables/summary_stats}}
93 +\end{table}
94 +
95 +\begin{figure}[t]
96 +\centering
97 +\includegraphics[width=\textwidth]{fig_rv_series.png}
98 +\caption{Annualized realized volatility (subsampled 5-minute RV), one
99 +representative instrument per asset class, 2013--2026.}
100 +\label{fig:rv_series}
101 +\end{figure}
102 +
103 +\subsection{Microstructure noise across measures}
104 +
105 +Figure~\ref{fig:measures} compares yearly averages of 1-minute RV,
106 +subsampled 5-minute RV and the realized kernel for SPY and BTC. For SPY
107 +the three measures are nearly indistinguishable --- at SPY's liquidity,
108 +1-minute sampling carries little noise, and the pairwise correlations
109 +between measures exceed 0.98. For BTC the ordering
110 +$RV^{(1\mathrm{min})} > RV^{ss} > RK$ in the early sample (2013--2015)
111 +is exactly the pattern implied by i.i.d.\ microstructure noise, whose
112 +variance contribution grows linearly in the sampling frequency: early
113 +Bitcoin trading was thin and fragmented, and plain 1-minute RV
114 +overstates variance by a factor of two or more relative to the
115 +noise-robust kernel. The gap closes progressively as crypto market
116 +quality improves after 2017. This pattern justifies our choice of the
117 +subsampled 5-minute RV as the headline measure --- following the
118 +multi-asset evidence of \citet{LiuPattonSheppard2015} --- with the
119 +1-minute RV and the realized kernel used to assess robustness to the
120 +sampling scheme in Section~\ref{sec:robustness}.
121 +
122 +\begin{figure}[t]
123 +\centering
124 +\includegraphics[width=0.95\textwidth]{fig_measures.png}
125 +\caption{Yearly mean annualized volatility by realized measure. The
126 +divergence between the 1-minute RV and the noise-robust measures for
127 +early-sample BTC reflects market microstructure noise.}
128 +\label{fig:measures}
129 +\end{figure}
130 +
131 +\begin{figure}[t]
132 +\centering
133 +\includegraphics[width=0.8\textwidth]{fig_acf.png}
134 +\caption{Autocorrelation functions of daily $\ln RV$ (subsampled
135 +5-minute measure), representative instruments.}
136 +\label{fig:acf}
137 +\end{figure}
added results/tables/summary_stats.tex +35 −0
@@ -0,0 +1,35 @@
1 +\begin{tabular}{lrrrrrrr}
2 +\toprule
3 +Ticker & Days & \makecell{Mean vol.\\(\%)} & \makecell{Median vol.\\(\%)} & \makecell{P95 vol.\\(\%)} & \makecell{Skew.\\$\ln$RV} & \makecell{AC(1)\\$\ln$RV} & \makecell{Jump\\days (\%)} \\
4 +\midrule
5 +\multicolumn{8}{l}{\itshape Equities/ETFs}\\
6 +\quad AAPL & 3,413 & 18.8 & 16.6 & 34.2 & 0.55 & 0.73 & 4.6 \\
7 +\quad GLD & 3,396 & 8.8 & 7.8 & 16.4 & 0.65 & 0.59 & 3.9 \\
8 +\quad IWM & 3,405 & 15.3 & 13.3 & 28.5 & 0.77 & 0.77 & 2.5 \\
9 +\quad JPM & 3,388 & 17.6 & 15.6 & 30.7 & 0.98 & 0.74 & 4.6 \\
10 +\quad MSFT & 3,397 & 17.8 & 15.9 & 32.0 & 0.65 & 0.75 & 4.7 \\
11 +\quad NVDA & 3,401 & 30.3 & 26.4 & 57.0 & 0.60 & 0.75 & 4.4 \\
12 +\quad QQQ & 3,410 & 13.4 & 11.3 & 27.2 & 0.47 & 0.78 & 2.8 \\
13 +\quad SPY & 3,414 & 10.2 & 8.4 & 21.3 & 0.53 & 0.79 & 2.8 \\
14 +\quad TLT & 3,394 & 8.5 & 7.7 & 15.0 & 0.73 & 0.70 & 2.8 \\
15 +\quad XOM & 3,389 & 18.0 & 15.7 & 33.6 & 0.59 & 0.84 & 4.3 \\
16 +\multicolumn{8}{l}{\itshape FX}\\
17 +\quad AUDUSD & 3,521 & 9.6 & 8.9 & 16.1 & 0.41 & 0.68 & 8.9 \\
18 +\quad EURUSD & 3,522 & 7.2 & 6.6 & 12.8 & 0.30 & 0.69 & 7.7 \\
19 +\quad GBPUSD & 3,521 & 8.1 & 7.4 & 13.5 & 0.65 & 0.70 & 7.7 \\
20 +\quad USDCAD & 3,521 & 6.7 & 6.2 & 11.4 & 0.24 & 0.69 & 6.3 \\
21 +\quad USDJPY & 3,522 & 7.9 & 7.0 & 14.7 & 0.24 & 0.72 & 10.2 \\
22 +\multicolumn{8}{l}{\itshape Crypto}\\
23 +\quad ADA & 2,750 & 106.8 & 81.4 & 221.7 & -2.75 & 0.65 & 25.2 \\
24 +\quad BTC & 4,498 & 97.8 & 50.3 & 372.7 & 0.95 & 0.90 & 11.8 \\
25 +\quad ETH & 3,528 & 108.4 & 75.2 & 289.7 & 0.46 & 0.88 & 12.1 \\
26 +\quad SOL & 1,917 & 81.4 & 69.7 & 162.2 & 0.53 & 0.76 & 5.9 \\
27 +\quad XRP & 3,373 & 137.3 & 77.2 & 478.6 & 0.85 & 0.90 & 18.3 \\
28 +\multicolumn{8}{l}{\itshape Futures}\\
29 +\quad CL & 3,502 & 33.4 & 28.3 & 64.7 & 0.78 & 0.87 & 10.5 \\
30 +\quad ES & 3,501 & 13.4 & 10.9 & 28.3 & 0.60 & 0.81 & 6.6 \\
31 +\quad GC & 3,501 & 14.1 & 12.5 & 25.1 & 0.71 & 0.70 & 8.9 \\
32 +\quad NQ & 3,498 & 16.9 & 14.2 & 35.2 & 0.38 & 0.80 & 6.3 \\
33 +\quad ZN & 3,216 & 5.2 & 4.6 & 9.3 & 0.82 & 0.68 & 19.5 \\
34 +\bottomrule
35 +\end{tabular}
modified scripts/02_build_rv.py +42 −7
@@ -7,8 +7,10 @@ Projet : Prévision de volatilité réalisée multi-actifs
7 7 (HAR-RV vs GARCH vs Machine Learning)
8 8 Fichier : 02_build_rv.py
9 9 Description : Étape 02 — Construction des mesures de volatilité
10 − réalisée journalières et du panel multi-actifs.
11 −Usage : python scripts/02_build_rv.py [--refresh]
10 + réalisée journalières (parallélisée par actif) et
11 + assemblage du panel multi-actifs.
12 +Usage : python scripts/02_build_rv.py [--refresh] [--workers 8]
13 + [--only-cached]
12 14 ================================================================
13 15 """
14 16
@@ -17,26 +19,59 @@ from __future__ import annotations
17 19 import argparse
18 20 import logging
19 21 import sys
22 +from concurrent.futures import ProcessPoolExecutor, as_completed
20 23 from pathlib import Path
21 24
25 +import pandas as pd
26 +
22 27 sys.path.insert(0, str(Path(__file__).resolve().parents[1] / "src"))
23 28
24 −from wp12 import config, data_pipeline # noqa: E402
29 +from wp12 import config # noqa: E402
25 30
26 31 logger = logging.getLogger("02_build_rv")
27 32
28 33
34 +def build_one(inst: dict[str, str], refresh: bool) -> str:
35 + """Worker: build the daily RV file for one instrument."""
36 + from wp12 import data_pipeline
37 + from wp12.api_client import HFMarketDataClient
38 +
39 + config.setup_logging(logging.WARNING)
40 + client = HFMarketDataClient()
41 + df = data_pipeline.build_instrument_rv(client, inst, refresh=refresh)
42 + return f"{inst['ticker']}: {len(df)} days"
43 +
44 +
29 45 def main() -> None:
30 − """Build daily realized measures for the universe and save the panel."""
46 + """Build per-ticker RV files in parallel, then assemble the panel."""
31 47 ap = argparse.ArgumentParser()
32 48 ap.add_argument("--refresh", action="store_true", help="recompute RV files")
49 + ap.add_argument("--workers", type=int, default=8)
50 + ap.add_argument("--only-cached", action="store_true",
51 + help="only tickers whose 1-min bars are already cached")
33 52 args = ap.parse_args()
34 53
35 54 config.setup_logging()
36 55 logging.getLogger("urllib3").setLevel(logging.WARNING)
37 − panel = data_pipeline.build_panel(refresh=args.refresh)
38 − logger.info("panel summary:\n%s",
39 − panel.groupby("ticker")["rv5ss"].agg(["count", "mean"]))
56 + univ = config.universe(include_predictors=True)
57 + if args.only_cached:
58 + raw = config.path("raw")
59 + univ = [i for i in univ
60 + if (raw / f"{i['asset']}_{i['ticker']}_1min.parquet").exists()]
61 + logger.info("building RV for %d instruments", len(univ))
62 + with ProcessPoolExecutor(max_workers=args.workers) as pool:
63 + futs = {pool.submit(build_one, i, args.refresh): i["ticker"] for i in univ}
64 + for fut in as_completed(futs):
65 + try:
66 + logger.info("done %s", fut.result())
67 + except Exception as exc: # noqa: BLE001
68 + logger.error("FAILED %s: %s", futs[fut], exc)
69 +
70 + if not args.only_cached:
71 + from wp12 import data_pipeline
72 + panel = data_pipeline.build_panel()
73 + logger.info("panel: %d rows, %d tickers", len(panel),
74 + panel["ticker"].nunique())
40 75
41 76
42 77 if __name__ == "__main__":
modified scripts/07_make_figures.py +5 −2
@@ -59,11 +59,14 @@ def fig_rv_series() -> None:
59 59 fig, axes = plt.subplots(4, 1, figsize=(TEXTWIDTH, 6.6), sharex=True)
60 60 for ax, (cls, tk) in zip(axes, REPRESENTATIVE.items()):
61 61 df = panel[panel["ticker"] == tk].sort_index()
62 − ax.plot(df.index, annualized_vol(df["rv5ss"]), lw=0.4, color=BLUE)
62 + vol = annualized_vol(df["rv5ss"])
63 + ax.plot(df.index, vol, lw=0.4, color=BLUE)
63 64 panel_label(ax, f"Panel {'ABCD'[list(REPRESENTATIVE).index(cls)]}. "
64 65 f"{tk} ({CLS_LABEL[cls]})")
65 66 ax.set_ylabel("Ann. vol. (%)")
66 − ax.set_ylim(bottom=0)
67 + # cap the axis at the 99.5th percentile so isolated early-sample
68 + # noise spikes (e.g. 2013 BTC) do not compress the whole panel
69 + ax.set_ylim(0, float(vol.quantile(0.995)) * 1.1)
67 70 axes[-1].set_xlabel("")
68 71 fig.tight_layout()
69 72 fig.savefig(FIG / "fig_rv_series.png")
modified src/wp12/evaluation.py +7 −4
@@ -150,17 +150,20 @@ def model_confidence_set(
150 150 n = len(L)
151 151 rng = np.random.default_rng(seed)
152 152 idx = _stationary_bootstrap_idx(n, n_boot, avg_block, rng)
153 + # bootstrap means per model, computed once column-by-column (memory-safe)
154 + all_boot_means = np.empty((n_boot, len(models)))
155 + for j in range(len(models)):
156 + all_boot_means[:, j] = L[idx, j].mean(axis=1)
157 + all_dbar = L.mean(axis=0)
153 158
154 159 included = list(range(len(models)))
155 160 pvals: dict[str, float] = {}
156 161 running_max_p = 0.0
157 162 while len(included) > 1:
158 − Ls = L[:, included]
159 − dbar_i = Ls.mean(axis=0) # mean loss per model
163 + dbar_i = all_dbar[included] # mean loss per model
160 164 d_i_dot = dbar_i - dbar_i.mean() # vs. set average
161 165 # bootstrap distribution of the max t-stat
162 − boot_means = Ls[idx].mean(axis=1) # (b, k)
163 − boot_center = boot_means - dbar_i
166 + boot_center = all_boot_means[:, included] - dbar_i
164 167 var_i = (boot_center - boot_center.mean(axis=0)).var(axis=0) + EPS
165 168 t_i = d_i_dot / np.sqrt(var_i)
166 169 t_max = float(t_i.max())
167 170