SPB Git

spb/wp7_uqo Public

UQO Working Paper No. 7 — Options-implied information for cross-asset return and volatility prediction: evidence from 3.8B option contracts.

Python 66.5% TeX 32.7% Makefile 0.8%
10.1 KB · 216 lines latex
Raw Blame History
1% =============================================================================2% Author: Simon-Pierre Boucher3% Contact: contact@spboucher.ai4% =============================================================================5% ============================================================================6% Results7% ============================================================================8\section{Results}\label{sec:results}910\subsection{RQ1: Return Predictability}1112\subsubsection{Panel Regressions}1314Table~\ref{tab:rq1} reports pooled OLS results. Predictive power is dramatically stronger at the 5-day horizon: $R^2$ increases from 0.08\% (stocks, 1-day) to 4.8\% (stocks, 5-day), 12.4\% (ETFs), and 19.3\% (indices).1516\begin{table}[H]17\centering18\caption{Return predictability: pooled OLS panel regressions.}19\label{tab:rq1}20\begin{threeparttable}21\begin{tabular}{@{}ld{1.4}d{1.4}d{1.4}d{1.4}d{1.4}d{1.4}@{}}22\toprule23& \multicolumn{3}{c}{1-Day return} & \multicolumn{3}{c}{5-Day return} \\24\cmidrule(lr){2-4}\cmidrule(lr){5-7}25& \multicolumn{1}{c}{Stocks} & \multicolumn{1}{c}{ETFs} & \multicolumn{1}{c}{Idx} & \multicolumn{1}{c}{Stocks} & \multicolumn{1}{c}{ETFs} & \multicolumn{1}{c}{Idx} \\26\midrule27$R^2$       & 0.0008 & 0.0014 & 0.0043 & 0.0476 & 0.1242 & 0.1930 \\28Adj.\ $R^2$ & 0.0007 & 0.0011 & 0.0032 & 0.0475 & 0.1239 & 0.1921 \\29$N$          & 79943 & 30152 & 8986 & 79943 & 30152 & 8986 \\30\# sig.\ (5\%) & 2 & 2 & 0 & 7 & 9 & 9 \\31\bottomrule32\end{tabular}33\begin{tablenotes}[flushleft]\footnotesize34\item \textit{Notes.} Ten predictors: $IV_{ATM,30d}$, IV term slope, skew$_{25\delta}$, implied skewness, implied kurtosis, PC vol.\ ratio, PC OI ratio, net gamma exposure, $RV_t$, $RV_t^{(w)}$. HC1 standard errors. All variables winsorized at 1\%/99\% and standardized.35\end{tablenotes}36\end{threeparttable}37\end{table}3839\subsubsection{Portfolio Sorts}4041Table~\ref{tab:sorts} reports quintile long-short portfolio performance for 5-day stock returns. The implied kurtosis sort delivers an annualized Sharpe of 2.33 ($t=19.84$). Sorting on the put-call volume ratio yields Sharpe $=-2.48$: stocks with the highest put-call ratio (bearish sentiment) earn the lowest future returns.4243\begin{table}[H]44\centering45\caption{Long-short portfolio performance (Q5$-$Q1, 5-day returns).}46\label{tab:sorts}47\begin{threeparttable}48\begin{tabular}{@{}ld{2.2}d{2.2}d{1.3}d{2.3}r@{}}49\toprule50Sorting variable & \multicolumn{1}{c}{\shortstack{Ann.\\ret.\,(\%)}} & \multicolumn{1}{c}{\shortstack{Ann.\\vol.\,(\%)}} & \multicolumn{1}{c}{Sharpe} & \multicolumn{1}{c}{$t$-stat} & $N_{days}$ \\51\midrule52Implied kurtosis       &  44.04 & 18.91 &  2.328 &  19.843 & 3,777 \\53IV term slope          &  11.33 & 20.30 &  0.558 &   4.004 & 2,676 \\54PC OI ratio            &   5.27 & 13.27 &  0.397 &   3.492 & 4,024 \\[3pt]55ATM IV (30d)           &  -1.76 & 26.13 & -0.067 &  -0.575 & 3,777 \\56Implied skewness       &  -9.91 & 18.81 & -0.527 &  -4.493 & 3,777 \\57Put-call vol.\ ratio   & -30.29 & 12.22 & -2.478 & -21.796 & 4,024 \\58Skew$_{25\delta}$      & -37.53 & 20.10 & -1.868 & -15.918 & 3,778 \\59\bottomrule60\end{tabular}61\begin{tablenotes}[flushleft]\footnotesize62\item \textit{Notes.} Equal-weighted quintile portfolios formed each trading day. Returns are 5-day (weekly). Annualization uses $\sqrt{52}$. $t$-statistics test $H_0$: mean daily L/S return~$=0$.63\end{tablenotes}64\end{threeparttable}65\end{table}6667\subsubsection{Double Sorts}6869Table~\ref{tab:dsort} shows a $3\times 3$ sort on ATM~IV and implied skewness. Among high-IV stocks, those with high skewness earn $-47.8$~bps/week ($t=-6.17$), while those with low skewness earn $+41.8$~bps/week ($t=13.12$). The information content of skewness is amplified in high-uncertainty environments.7071\begin{table}[H]72\centering73\caption{Double sort: ATM\,IV $\times$ implied skewness $\to$ 5-day returns (bps/week).}74\label{tab:dsort}75\begin{threeparttable}76\begin{tabular}{@{}ld{2.2}d{2.2}d{2.2}@{}}77\toprule78& \multicolumn{1}{c}{Low skew} & \multicolumn{1}{c}{Med.} & \multicolumn{1}{c}{High skew} \\79\midrule80Low IV    &   8.15 & 24.12 &  35.90 \\81Medium IV &  22.11 & 23.80 &  15.62 \\82High IV   &  41.76 &  2.72 & -47.82 \\83\midrule84\textit{High$-$Low IV} & \textit{33.61} & \textit{$-$21.40} & \textit{$-$83.72} \\85\bottomrule86\end{tabular}87\begin{tablenotes}[flushleft]\footnotesize88\item \textit{Notes.} Tercile breakpoints computed cross-sectionally each day.89\end{tablenotes}90\end{threeparttable}91\end{table}929394\subsection{RQ2: IV Surface vs.\ HAR-RV}9596Table~\ref{tab:rq2} reports in-sample $R^2$. The combined HAR-RV\,$+$\,IV-surface model achieves $R^2=0.442$ for 1-day~RV, a 23.3\% improvement over HAR-RV (0.359). The IV-only model ($R^2=0.390$) outperforms HAR-RV at the daily horizon.9798\begin{table}[H]99\centering100\caption{In-sample $R^2$: realized variance forecasting models.}101\label{tab:rq2}102\begin{threeparttable}103\begin{tabular}{@{}ld{1.3}d{1.3}r@{}}104\toprule105Model & \multicolumn{1}{c}{1-Day RV} & \multicolumn{1}{c}{5-Day RV} & \multicolumn{1}{c}{$N$} \\106\midrule107GARCH proxy           & 0.326 & 0.648 & 264,380 \\108HAR-RV                & 0.359 & 0.888 & 264,345 \\109IV only               & 0.390 & 0.495 & 120,280 \\110IV surface            & 0.391 & 0.496 & 119,093 \\111HAR-RV $+$ IV surface & 0.442 & 0.896 & 119,093 \\112\bottomrule113\end{tabular}114\end{threeparttable}115\end{table}116117Diebold-Mariano tests show statistically significant MSE improvements for AAPL ($+10.3\%$, DM\,$=2.54$), AMD ($+11.7\%$, DM\,$=3.13$), CAT ($+32.7\%$, DM\,$=4.27$), and COST ($+24.4\%$, DM\,$=2.94$).118119120\subsection{RQ3: Implied Correlation and Stress Prediction}121122Table~\ref{tab:rq3} shows that the correlation \textit{ratio} ($\hat{\rho}_{impl}/\hat{\rho}_{real}$) significantly predicts stress at all horizons ($t$: 3.91--8.30). The raw divergence is not significant; the multiplicative form captures the signal. Appendix~\ref{app:crises} details the behavior of the implied-correlation index around nine major market events.123124\begin{table}[H]125\centering126\caption{Stress prediction: linear probability model ($t$-statistics).}127\label{tab:rq3}128\begin{threeparttable}129\begin{tabular}{@{}ld{2.2}d{2.2}d{2.2}@{}}130\toprule131Predictor & \multicolumn{1}{c}{5-Day} & \multicolumn{1}{c}{10-Day} & \multicolumn{1}{c}{20-Day} \\132\midrule133Corr.\ ratio ($\hat{\rho}_{impl}/\hat{\rho}_{real}$) & 4.07\sym{***} & 8.30\sym{***} & 3.91\sym{***} \\134SPX ATM IV        & -2.23\sym{**}  & -4.07\sym{***} & -7.30\sym{***} \\135VIX level         &  5.27\sym{***} &  6.69\sym{***} &  9.56\sym{***} \\136Corr.\ divergence &  0.00          &  0.00          &  0.00          \\137\midrule138$R^2$             &  0.228         &  0.156         &  0.088         \\139$N$               &  2486          &  2481          &  2471          \\140\bottomrule141\end{tabular}142\begin{tablenotes}[flushleft]\footnotesize143\item \textit{Notes.} HC1 standard errors. \sym{***}$p<0.01$; \sym{**}$p<0.05$.144\end{tablenotes}145\end{threeparttable}146\end{table}147148149\subsection{RQ4: Greeks Information Decay and Price Magnets}150151The information content of Greeks for RV forecasting increases monotonically with DTE: $R^2$ rises from 3.8\% (1-week) to 9.5\% (3-month), then declines slightly to 8.2\% (6-month). Return predictability is negligible across all buckets ($R^2<0.2\%$).152153For the price magnet analysis, prices moved toward the max-OI strike in only 47.0\% of 10,480 cases---below the 50\% null. The OI concentration coefficient is not significant ($t=1.06$), contradicting the ``max pain'' narrative.154155156\subsection{RQ5: ML Models for SPX RV}157158Table~\ref{tab:rq5} reports the out-of-sample comparison. OLS HAR-RV dominates at every horizon ($R^2_{OOS}$: 0.437, 0.948, 0.994). ML models underperform (RF: 0.143; GBM: 0.095 for 1-day). RF feature importance shows that 2-week ATM~IV accounts for 50.8\% of total importance; the full ranking is reported in Appendix~\ref{app:importance}.159160\begin{table}[H]161\centering162\caption{Out-of-sample SPX realized variance forecasting (2020--2025).}163\label{tab:rq5}164\begin{threeparttable}165\begin{tabular}{@{}lld{4.3}d{2.3}d{1.3}@{}}166\toprule167Model & Target & \multicolumn{1}{c}{MSE\,$(\times10^7)$} & \multicolumn{1}{c}{MAE\,$(\times10^4)$} & \multicolumn{1}{c}{$R^2_{OOS}$} \\168\midrule169VIX benchmark      & 1-Day &  1.835 & 1.371 & 0.336 \\170OLS: HAR-RV        & 1-Day &  1.556 & 1.018 & 0.437 \\171Random Forest      & 1-Day &  2.368 & 1.148 & 0.143 \\172Gradient Boosting  & 1-Day &  2.501 & 1.282 & 0.095 \\173\midrule174VIX benchmark      & 5-Day & 18.870 & 5.425 & 0.597 \\175OLS: HAR-RV        & 5-Day &  2.452 & 1.431 & 0.948 \\176\midrule177VIX benchmark      & 22-Day & 211.1 & 21.61 & 0.618 \\178OLS: HAR-RV        & 22-Day &  3.518 &  1.709 & 0.994 \\179\bottomrule180\end{tabular}181\begin{tablenotes}[flushleft]\footnotesize182\item \textit{Notes.} Training: 2010--2019 (2,515 days). Test: 2020--2025 (1,507 days).183\end{tablenotes}184\end{threeparttable}185\end{table}186187188\subsection{Granger Causality and VAR Analysis}189190Table~\ref{tab:granger} reports the Granger causality results. ATM~IV strongly Granger-causes RV ($F=62.4$, significant in 100\% of tickers). Returns Granger-cause skew ($F=13.6$, 95.7\%), indicating asymmetric volatility feedback.191192\begin{table}[H]193\centering194\caption{Granger causality tests (5 lags).}195\label{tab:granger}196\begin{threeparttable}197\begin{tabular}{@{}ld{2.3}d{1.4}d{2.1}@{}}198\toprule199Hypothesis & \multicolumn{1}{c}{Avg.\ $F$} & \multicolumn{1}{c}{Avg.\ $p$} & \multicolumn{1}{c}{\% sig.} \\200\midrule201ATM IV $\to$ RV        & 62.381 & 0.0000 & 100.0 \\202RV $\to$ ATM IV        & 17.933 & 0.1501 &  60.9 \\203Skew $\to$ RV          & 20.792 & 0.0177 &  94.2 \\204Return $\to$ Skew      & 13.576 & 0.0133 &  95.7 \\205ATM IV $\to$ Return    &  1.799 & 0.2688 &  24.6 \\206PC ratio $\to$ Return  &  1.010 & 0.4999 &   5.8 \\207\bottomrule208\end{tabular}209\begin{tablenotes}[flushleft]\footnotesize210\item \textit{Notes.} Per-ticker bivariate $F$-tests. \% sig.\ is the share of tickers rejecting the null at 5\%.211\end{tablenotes}212\end{threeparttable}213\end{table}214215The bivariate VAR(5) impulse response shows that a one-SD shock to ATM~IV produces a cumulative RV response of 0.67~SD at horizon~2, decaying to 0.27~SD at horizon~20. The forecast error variance decomposition reveals that IV shocks explain an increasing share of RV forecast error variance, reaching \textbf{73.8\%} at the 20-day horizon (Appendix~\ref{app:fevd}).216