spb/wp7_uqo Public
UQO Working Paper No. 7 — Options-implied information for cross-asset return and volatility prediction: evidence from 3.8B option contracts.
Python 66.5%
TeX 32.7%
Makefile 0.8%
1% =============================================================================2% Author: Simon-Pierre Boucher3% Contact: contact@spboucher.ai4% =============================================================================5% ============================================================================6% Results7% ============================================================================8\section{Results}\label{sec:results}910\subsection{RQ1: Return Predictability}1112\subsubsection{Panel Regressions}1314Table~\ref{tab:rq1} reports pooled OLS results. Predictive power is dramatically stronger at the 5-day horizon: $R^2$ increases from 0.08\% (stocks, 1-day) to 4.8\% (stocks, 5-day), 12.4\% (ETFs), and 19.3\% (indices).1516\begin{table}[H]17\centering18\caption{Return predictability: pooled OLS panel regressions.}19\label{tab:rq1}20\begin{threeparttable}21\begin{tabular}{@{}ld{1.4}d{1.4}d{1.4}d{1.4}d{1.4}d{1.4}@{}}22\toprule23& \multicolumn{3}{c}{1-Day return} & \multicolumn{3}{c}{5-Day return} \\24\cmidrule(lr){2-4}\cmidrule(lr){5-7}25& \multicolumn{1}{c}{Stocks} & \multicolumn{1}{c}{ETFs} & \multicolumn{1}{c}{Idx} & \multicolumn{1}{c}{Stocks} & \multicolumn{1}{c}{ETFs} & \multicolumn{1}{c}{Idx} \\26\midrule27$R^2$ & 0.0008 & 0.0014 & 0.0043 & 0.0476 & 0.1242 & 0.1930 \\28Adj.\ $R^2$ & 0.0007 & 0.0011 & 0.0032 & 0.0475 & 0.1239 & 0.1921 \\29$N$ & 79943 & 30152 & 8986 & 79943 & 30152 & 8986 \\30\# sig.\ (5\%) & 2 & 2 & 0 & 7 & 9 & 9 \\31\bottomrule32\end{tabular}33\begin{tablenotes}[flushleft]\footnotesize34\item \textit{Notes.} Ten predictors: $IV_{ATM,30d}$, IV term slope, skew$_{25\delta}$, implied skewness, implied kurtosis, PC vol.\ ratio, PC OI ratio, net gamma exposure, $RV_t$, $RV_t^{(w)}$. HC1 standard errors. All variables winsorized at 1\%/99\% and standardized.35\end{tablenotes}36\end{threeparttable}37\end{table}3839\subsubsection{Portfolio Sorts}4041Table~\ref{tab:sorts} reports quintile long-short portfolio performance for 5-day stock returns. The implied kurtosis sort delivers an annualized Sharpe of 2.33 ($t=19.84$). Sorting on the put-call volume ratio yields Sharpe $=-2.48$: stocks with the highest put-call ratio (bearish sentiment) earn the lowest future returns.4243\begin{table}[H]44\centering45\caption{Long-short portfolio performance (Q5$-$Q1, 5-day returns).}46\label{tab:sorts}47\begin{threeparttable}48\begin{tabular}{@{}ld{2.2}d{2.2}d{1.3}d{2.3}r@{}}49\toprule50Sorting variable & \multicolumn{1}{c}{\shortstack{Ann.\\ret.\,(\%)}} & \multicolumn{1}{c}{\shortstack{Ann.\\vol.\,(\%)}} & \multicolumn{1}{c}{Sharpe} & \multicolumn{1}{c}{$t$-stat} & $N_{days}$ \\51\midrule52Implied kurtosis & 44.04 & 18.91 & 2.328 & 19.843 & 3,777 \\53IV term slope & 11.33 & 20.30 & 0.558 & 4.004 & 2,676 \\54PC OI ratio & 5.27 & 13.27 & 0.397 & 3.492 & 4,024 \\[3pt]55ATM IV (30d) & -1.76 & 26.13 & -0.067 & -0.575 & 3,777 \\56Implied skewness & -9.91 & 18.81 & -0.527 & -4.493 & 3,777 \\57Put-call vol.\ ratio & -30.29 & 12.22 & -2.478 & -21.796 & 4,024 \\58Skew$_{25\delta}$ & -37.53 & 20.10 & -1.868 & -15.918 & 3,778 \\59\bottomrule60\end{tabular}61\begin{tablenotes}[flushleft]\footnotesize62\item \textit{Notes.} Equal-weighted quintile portfolios formed each trading day. Returns are 5-day (weekly). Annualization uses $\sqrt{52}$. $t$-statistics test $H_0$: mean daily L/S return~$=0$.63\end{tablenotes}64\end{threeparttable}65\end{table}6667\subsubsection{Double Sorts}6869Table~\ref{tab:dsort} shows a $3\times 3$ sort on ATM~IV and implied skewness. Among high-IV stocks, those with high skewness earn $-47.8$~bps/week ($t=-6.17$), while those with low skewness earn $+41.8$~bps/week ($t=13.12$). The information content of skewness is amplified in high-uncertainty environments.7071\begin{table}[H]72\centering73\caption{Double sort: ATM\,IV $\times$ implied skewness $\to$ 5-day returns (bps/week).}74\label{tab:dsort}75\begin{threeparttable}76\begin{tabular}{@{}ld{2.2}d{2.2}d{2.2}@{}}77\toprule78& \multicolumn{1}{c}{Low skew} & \multicolumn{1}{c}{Med.} & \multicolumn{1}{c}{High skew} \\79\midrule80Low IV & 8.15 & 24.12 & 35.90 \\81Medium IV & 22.11 & 23.80 & 15.62 \\82High IV & 41.76 & 2.72 & -47.82 \\83\midrule84\textit{High$-$Low IV} & \textit{33.61} & \textit{$-$21.40} & \textit{$-$83.72} \\85\bottomrule86\end{tabular}87\begin{tablenotes}[flushleft]\footnotesize88\item \textit{Notes.} Tercile breakpoints computed cross-sectionally each day.89\end{tablenotes}90\end{threeparttable}91\end{table}929394\subsection{RQ2: IV Surface vs.\ HAR-RV}9596Table~\ref{tab:rq2} reports in-sample $R^2$. The combined HAR-RV\,$+$\,IV-surface model achieves $R^2=0.442$ for 1-day~RV, a 23.3\% improvement over HAR-RV (0.359). The IV-only model ($R^2=0.390$) outperforms HAR-RV at the daily horizon.9798\begin{table}[H]99\centering100\caption{In-sample $R^2$: realized variance forecasting models.}101\label{tab:rq2}102\begin{threeparttable}103\begin{tabular}{@{}ld{1.3}d{1.3}r@{}}104\toprule105Model & \multicolumn{1}{c}{1-Day RV} & \multicolumn{1}{c}{5-Day RV} & \multicolumn{1}{c}{$N$} \\106\midrule107GARCH proxy & 0.326 & 0.648 & 264,380 \\108HAR-RV & 0.359 & 0.888 & 264,345 \\109IV only & 0.390 & 0.495 & 120,280 \\110IV surface & 0.391 & 0.496 & 119,093 \\111HAR-RV $+$ IV surface & 0.442 & 0.896 & 119,093 \\112\bottomrule113\end{tabular}114\end{threeparttable}115\end{table}116117Diebold-Mariano tests show statistically significant MSE improvements for AAPL ($+10.3\%$, DM\,$=2.54$), AMD ($+11.7\%$, DM\,$=3.13$), CAT ($+32.7\%$, DM\,$=4.27$), and COST ($+24.4\%$, DM\,$=2.94$).118119120\subsection{RQ3: Implied Correlation and Stress Prediction}121122Table~\ref{tab:rq3} shows that the correlation \textit{ratio} ($\hat{\rho}_{impl}/\hat{\rho}_{real}$) significantly predicts stress at all horizons ($t$: 3.91--8.30). The raw divergence is not significant; the multiplicative form captures the signal. Appendix~\ref{app:crises} details the behavior of the implied-correlation index around nine major market events.123124\begin{table}[H]125\centering126\caption{Stress prediction: linear probability model ($t$-statistics).}127\label{tab:rq3}128\begin{threeparttable}129\begin{tabular}{@{}ld{2.2}d{2.2}d{2.2}@{}}130\toprule131Predictor & \multicolumn{1}{c}{5-Day} & \multicolumn{1}{c}{10-Day} & \multicolumn{1}{c}{20-Day} \\132\midrule133Corr.\ ratio ($\hat{\rho}_{impl}/\hat{\rho}_{real}$) & 4.07\sym{***} & 8.30\sym{***} & 3.91\sym{***} \\134SPX ATM IV & -2.23\sym{**} & -4.07\sym{***} & -7.30\sym{***} \\135VIX level & 5.27\sym{***} & 6.69\sym{***} & 9.56\sym{***} \\136Corr.\ divergence & 0.00 & 0.00 & 0.00 \\137\midrule138$R^2$ & 0.228 & 0.156 & 0.088 \\139$N$ & 2486 & 2481 & 2471 \\140\bottomrule141\end{tabular}142\begin{tablenotes}[flushleft]\footnotesize143\item \textit{Notes.} HC1 standard errors. \sym{***}$p<0.01$; \sym{**}$p<0.05$.144\end{tablenotes}145\end{threeparttable}146\end{table}147148149\subsection{RQ4: Greeks Information Decay and Price Magnets}150151The information content of Greeks for RV forecasting increases monotonically with DTE: $R^2$ rises from 3.8\% (1-week) to 9.5\% (3-month), then declines slightly to 8.2\% (6-month). Return predictability is negligible across all buckets ($R^2<0.2\%$).152153For the price magnet analysis, prices moved toward the max-OI strike in only 47.0\% of 10,480 cases---below the 50\% null. The OI concentration coefficient is not significant ($t=1.06$), contradicting the ``max pain'' narrative.154155156\subsection{RQ5: ML Models for SPX RV}157158Table~\ref{tab:rq5} reports the out-of-sample comparison. OLS HAR-RV dominates at every horizon ($R^2_{OOS}$: 0.437, 0.948, 0.994). ML models underperform (RF: 0.143; GBM: 0.095 for 1-day). RF feature importance shows that 2-week ATM~IV accounts for 50.8\% of total importance; the full ranking is reported in Appendix~\ref{app:importance}.159160\begin{table}[H]161\centering162\caption{Out-of-sample SPX realized variance forecasting (2020--2025).}163\label{tab:rq5}164\begin{threeparttable}165\begin{tabular}{@{}lld{4.3}d{2.3}d{1.3}@{}}166\toprule167Model & Target & \multicolumn{1}{c}{MSE\,$(\times10^7)$} & \multicolumn{1}{c}{MAE\,$(\times10^4)$} & \multicolumn{1}{c}{$R^2_{OOS}$} \\168\midrule169VIX benchmark & 1-Day & 1.835 & 1.371 & 0.336 \\170OLS: HAR-RV & 1-Day & 1.556 & 1.018 & 0.437 \\171Random Forest & 1-Day & 2.368 & 1.148 & 0.143 \\172Gradient Boosting & 1-Day & 2.501 & 1.282 & 0.095 \\173\midrule174VIX benchmark & 5-Day & 18.870 & 5.425 & 0.597 \\175OLS: HAR-RV & 5-Day & 2.452 & 1.431 & 0.948 \\176\midrule177VIX benchmark & 22-Day & 211.1 & 21.61 & 0.618 \\178OLS: HAR-RV & 22-Day & 3.518 & 1.709 & 0.994 \\179\bottomrule180\end{tabular}181\begin{tablenotes}[flushleft]\footnotesize182\item \textit{Notes.} Training: 2010--2019 (2,515 days). Test: 2020--2025 (1,507 days).183\end{tablenotes}184\end{threeparttable}185\end{table}186187188\subsection{Granger Causality and VAR Analysis}189190Table~\ref{tab:granger} reports the Granger causality results. ATM~IV strongly Granger-causes RV ($F=62.4$, significant in 100\% of tickers). Returns Granger-cause skew ($F=13.6$, 95.7\%), indicating asymmetric volatility feedback.191192\begin{table}[H]193\centering194\caption{Granger causality tests (5 lags).}195\label{tab:granger}196\begin{threeparttable}197\begin{tabular}{@{}ld{2.3}d{1.4}d{2.1}@{}}198\toprule199Hypothesis & \multicolumn{1}{c}{Avg.\ $F$} & \multicolumn{1}{c}{Avg.\ $p$} & \multicolumn{1}{c}{\% sig.} \\200\midrule201ATM IV $\to$ RV & 62.381 & 0.0000 & 100.0 \\202RV $\to$ ATM IV & 17.933 & 0.1501 & 60.9 \\203Skew $\to$ RV & 20.792 & 0.0177 & 94.2 \\204Return $\to$ Skew & 13.576 & 0.0133 & 95.7 \\205ATM IV $\to$ Return & 1.799 & 0.2688 & 24.6 \\206PC ratio $\to$ Return & 1.010 & 0.4999 & 5.8 \\207\bottomrule208\end{tabular}209\begin{tablenotes}[flushleft]\footnotesize210\item \textit{Notes.} Per-ticker bivariate $F$-tests. \% sig.\ is the share of tickers rejecting the null at 5\%.211\end{tablenotes}212\end{threeparttable}213\end{table}214215The bivariate VAR(5) impulse response shows that a one-SD shock to ATM~IV produces a cumulative RV response of 0.67~SD at horizon~2, decaying to 0.27~SD at horizon~20. The forecast error variance decomposition reveals that IV shocks explain an increasing share of RV forecast error variance, reaching \textbf{73.8\%} at the 20-day horizon (Appendix~\ref{app:fevd}).216