spb/wp7_uqo Public
UQO Working Paper No. 7 — Options-implied information for cross-asset return and volatility prediction: evidence from 3.8B option contracts.
Python 66.5%
TeX 32.7%
Makefile 0.8%
1% =============================================================================2% Author: Simon-Pierre Boucher3% Contact: contact@spboucher.ai4% =============================================================================5% ============================================================================6% Robustness7% ============================================================================8\section{Robustness}\label{sec:robust}910\subsection{Alternative Standard Errors}1112Table~\ref{tab:robust_se} compares inference for the 5-day return regression under three SE specifications \citep{white1980heteroskedasticity,newey1987simple,cameron2011robust}; \citet{petersen2009estimating} shows that two-way clustering is the most conservative choice for finance panels of this type. Of the ten predictors, nine remain significant at 5\% under Newey-West HAC(5), and eight under double-clustered (ticker\,$+$\,date) SEs. Implied skewness ($t_{DC}=-4.94$), implied kurtosis ($t_{DC}=9.02$), PC volume ratio ($t_{DC}=-7.96$), and net gamma ($t_{DC}=5.39$) remain highly significant under all specifications.1314\begin{table}[H]15\centering16\caption{Robustness of 5-day return predictability to SE specification.}17\label{tab:robust_se}18\begin{threeparttable}19\begin{tabular}{@{}ld{2.2}d{2.2}d{2.2}@{}}20\toprule21Variable & \multicolumn{1}{c}{HC1} & \multicolumn{1}{c}{NW(5)} & \multicolumn{1}{c}{DC} \\22\midrule23$IV_{ATM,30d}$ & 3.86\sym{***} & 3.32\sym{***} & 1.40 \\24IV term slope & 4.83\sym{***} & 4.59\sym{***} & 2.12\sym{**} \\25Skew$_{25\delta}$ & -2.03\sym{**} & -0.93 & -0.39 \\26Implied skewness & -15.19\sym{***}& -13.66\sym{***}& -4.94\sym{***} \\27Implied kurtosis & 25.87\sym{***}& 25.64\sym{***}& 9.02\sym{***} \\28PC vol.\ ratio & -13.72\sym{***}& -16.69\sym{***}& -7.96\sym{***} \\29PC OI ratio & 12.33\sym{***}& 13.12\sym{***}& 4.61\sym{***} \\30Net gamma exp. & 15.55\sym{***}& 17.83\sym{***}& 5.39\sym{***} \\31$RV_{daily}$ & -17.18\sym{***}& -17.87\sym{***}& -7.27\sym{***} \\32$RV_{weekly}$ & 10.10\sym{***}& 7.95\sym{***}& 3.53\sym{***} \\33\bottomrule34\end{tabular}35\begin{tablenotes}[flushleft]\footnotesize36\item \textit{Notes.} NW(5) = Newey-West with 5 lags. DC = double-clustered by ticker and date following \citet{cameron2011robust}. \sym{***}$p<0.01$; \sym{**}$p<0.05$.37\end{tablenotes}38\end{threeparttable}39\end{table}4041\subsection{Subperiod Stability}4243Table~\ref{tab:subperiod} shows that both return predictability and the HAR+IV improvement are stable across nine subperiods. The 5-day return $R^2$ ranges from 3.1\% (post-GFC) to 14.4\% (COVID). The HAR+IV improvement over HAR-RV is positive in eight of nine subperiods, peaking at $+35.5\%$ during the low-vol era (2017--18) and $+34.4\%$ during the recovery (2023--24).4445\begin{table}[H]46\centering47\caption{Subperiod stability.}48\label{tab:subperiod}49\begin{threeparttable}50\small51\begin{tabular}{@{}ld{1.3}d{1.3}d{1.3}d{2.1}@{}}52\toprule53Subperiod & \multicolumn{1}{c}{\shortstack{5D ret\\$R^2$}} & \multicolumn{1}{c}{\shortstack{HAR-RV\\$R^2$}} & \multicolumn{1}{c}{\shortstack{HAR+IV\\$R^2$}} & \multicolumn{1}{c}{$\Delta R^2\,(\%)$} \\54\midrule55Post-GFC (2010--12) & 0.031 & 0.344 & 0.392 & +14.0 \\56Bull (2013--16) & 0.075 & 0.342 & 0.336 & -1.7 \\57Low vol (2017--18) & 0.136 & 0.309 & 0.419 & +35.5 \\58Pre-COVID (2019) & 0.080 & 0.280 & 0.363 & +29.9 \\59COVID (2020) & 0.144 & 0.604 & 0.721 & +19.4 \\60Post-COVID (2021) & 0.036 & 0.376 & 0.440 & +16.9 \\61Rate hike (2022) & 0.084 & 0.387 & 0.471 & +21.8 \\62Recovery (2023--24) & 0.052 & 0.290 & 0.390 & +34.4 \\63Recent (2025) & 0.037 & 0.321 & 0.402 & +25.2 \\64\bottomrule65\end{tabular}66\begin{tablenotes}[flushleft]\footnotesize67\item \textit{Notes.} $\Delta R^2 = (R^2_{HAR+IV} - R^2_{HAR}) / R^2_{HAR} \times 100$.68\end{tablenotes}69\end{threeparttable}70\end{table}7172\subsection{VIX Regime Conditioning}7374Table~\ref{tab:vix_regime} shows that predictability increases with VIX level. The 5-day return $R^2$ rises from 2.7\% (VIX\,$<$\,15) to 19.8\% (VIX\,$>$\,35). Option-implied information is most valuable in high-uncertainty environments.7576\begin{table}[H]77\centering78\caption{Predictability by VIX regime.}79\label{tab:vix_regime}80\begin{threeparttable}81\begin{tabular}{@{}ld{1.3}d{1.3}d{2.1}r@{}}82\toprule83VIX regime & \multicolumn{1}{c}{\shortstack{5D ret\\$R^2$}} & \multicolumn{1}{c}{\shortstack{HAR+IV\\$R^2$ (RV)}} & \multicolumn{1}{c}{\shortstack{Pct.\ of\\sample (\%)}} & $N$ \\84\midrule85Very low ($<$15) & 0.027 & 0.359 & 35.7 & 43,971 \\86Low (15--20) & 0.040 & 0.341 & 35.0 & 40,221 \\87Medium (20--25) & 0.073 & 0.394 & 15.8 & 19,210 \\88High (25--35) & 0.080 & 0.385 & 11.0 & 13,139 \\89Crisis ($>$35) & 0.198 & 0.618 & 2.5 & 2,540 \\90\bottomrule91\end{tabular}92\end{threeparttable}93\end{table}9495\subsection{Additional Controls, Quantile Regressions, and Ticker-Level Analysis}9697Adding controls (log volume, log OI, lagged returns) increases 5-day return $R^2$ from 5.3\% to 5.6\%; all core predictors retain significance. Quantile regressions at $\tau \in \{0.10, 0.25, 0.50, 0.75, 0.90\}$ \citep{koenker1978regression} show that the implied kurtosis and put-call ratio coefficients retain a stable sign and magnitude across the entire conditional return distribution.9899The distribution of ticker-level $R^2$ values for 5-day returns has mean 5.9\%, median 4.7\%, and interquartile range [2.0\%, 8.2\%], confirming that panel results are not driven by outliers.100101Rolling-window (252-day) analysis yields average $R^2$ of 8.0\% (std.\ 4.5\%, range [2.1\%, 17.9\%]) for 5-day returns and 41.2\% (std.\ 10.7\%) for 1-day RV forecasting.102103\subsection{Estimator and Specification Sensitivity}104105Two implementation choices could mechanically drive the baseline results: the winsorization cutoff and the HAC lag length. Table~\ref{tab:winsor} shows that neither does. Predictability rises modestly with heavier trimming---from $R^2 = 5.1\%$ at 0.5\% cutoffs to 6.2\% at 5\% cutoffs---because winsorization suppresses noise in the extreme tails, and even wholly unwinsorized data leave the key predictors highly significant ($t_{kurt} = 13.8$). Lengthening the Newey-West window from 5 to 22 lags \citep{newey1987simple}, enough to absorb the full overlap of the 5-day return plus a trading month, barely moves the $t$-statistics: implied kurtosis declines only from 25.6 to 22.1, and the set of significant predictors (9 of 10) is unchanged at every lag length.106107\begin{table}[H]108\centering109\caption{Winsorization and HAC-lag sensitivity (5-day returns).}110\label{tab:winsor}111\begin{threeparttable}112\small113\begin{tabular}{@{}ld{1.3}d{2.2}d{3.2}r@{\hskip 2em}ld{2.2}@{}}114\toprule115\multicolumn{5}{c}{Panel A: winsorization cutoff} & \multicolumn{2}{c}{Panel B: NW lags} \\116\cmidrule(r){1-5}\cmidrule(l){6-7}117Cutoff & \multicolumn{1}{c}{$R^2$} & \multicolumn{1}{c}{$t_{kurt}$} & \multicolumn{1}{c}{$t_{PC}$} & \multicolumn{1}{c}{\#sig.} & Lags & \multicolumn{1}{c}{$t_{kurt}$} \\118\midrule119None & 0.034 & 13.76 & -3.59 & 8/10 & NW(5) & 25.64 \\1200.5\% & 0.051 & 36.51 & -21.31 & 10/10 & NW(10) & 23.89 \\1211\% (baseline) & 0.053 & 40.07 & -21.48 & 9/10 & NW(22) & 22.09 \\1222.5\% & 0.058 & 43.66 & -21.32 & 9/10 & & \\1235\% & 0.062 & 45.73 & -20.86 & 9/10 & & \\124\bottomrule125\end{tabular}126\begin{tablenotes}[flushleft]\footnotesize127\item \textit{Notes.} Panel A: pooled 5-day return regression (all 69 tickers) with HC1 $t$-statistics under alternative symmetric winsorization cutoffs. $t_{kurt}$ and $t_{PC}$ refer to implied kurtosis and the put-call volume ratio. \#sig.\ counts predictors significant at 5\%. Panel B: Newey-West $t$-statistic of implied kurtosis at alternative lag lengths (baseline data treatment).128\end{tablenotes}129\end{threeparttable}130\end{table}131132\subsection{Rank-Based Information Coefficients}133134Parametric regressions may be sensitive to outliers and functional form even after winsorization. As a fully nonparametric check, I compute daily cross-sectional Spearman information coefficients (ICs)---the rank correlation between each predictor and the subsequent 5-day stock return---and test whether the time-series mean IC differs from zero. Table~\ref{tab:ic} confirms the regression evidence. Implied kurtosis carries a mean IC of $+0.098$ ($t = 16.8$), positive on 64.3\% of the 4{,}024 trading days; the put-call volume ratio ($-0.059$, $t = -13.8$) and the 25$\delta$ skew ($-0.056$, $t = -9.7$) are reliably negative. At the 1-day horizon no predictor achieves a mean IC above 0.02 in absolute value, mirroring the near-zero daily $R^2$ of Section~\ref{sec:results}.135136\begin{table}[H]137\centering138\caption{Daily cross-sectional Spearman information coefficients (stocks, 5-day returns).}139\label{tab:ic}140\begin{threeparttable}141\begin{tabular}{@{}ld{1.4}d{3.2}d{2.1}@{}}142\toprule143Predictor & \multicolumn{1}{c}{Mean IC} & \multicolumn{1}{c}{$t$-stat} & \multicolumn{1}{c}{\% days $>0$} \\144\midrule145Net gamma exposure & 0.1702 & 33.99 & 77.8 \\146Implied kurtosis & 0.0978 & 16.77 & 64.3 \\147IV term slope & 0.0328 & 6.30 & 54.6 \\148PC OI ratio & 0.0170 & 3.63 & 52.4 \\149$RV_{daily}$ & 0.0161 & 2.41 & 52.7 \\150$RV_{weekly}$ & 0.0120 & 1.73 & 51.2 \\151$IV_{ATM,30d}$ & 0.0085 & 1.09 & 53.0 \\152Implied skewness & -0.0188 & -3.30 & 47.4 \\153Skew$_{25\delta}$ & -0.0558 & -9.66 & 41.9 \\154PC vol.\ ratio & -0.0591 & -13.81 & 37.2 \\155\bottomrule156\end{tabular}157\begin{tablenotes}[flushleft]\footnotesize158\item \textit{Notes.} For each trading day with at least 20 stocks, the Spearman rank correlation between the predictor and the 5-day forward return is computed across stocks; the table reports the time-series mean, its Fama-MacBeth-style $t$-statistic, and the share of days with a positive IC ($N = 4{,}024$ days maximum).159\end{tablenotes}160\end{threeparttable}161\end{table}162163\subsection{Decile Sorts}164165The baseline quintile sorts could mask nonlinearity in the extreme tails of the sorting variables \citep{fama2008dissecting}. Table~\ref{tab:decile} repeats the three strongest sorts with decile portfolios. Sharpening the sort raises the spread in annualized returns---implied kurtosis rises from 44.0\% to 52.1\%, and the 25$\delta$ skew from $-37.5\%$ to $-52.5\%$---while Sharpe ratios remain essentially unchanged (2.33 vs.\ 2.19 for kurtosis), since the finer portfolios are noisier. The premia are therefore monotone rather than driven by a single extreme quintile.166167\begin{table}[H]168\centering169\caption{Quintile vs.\ decile long-short sorts (5-day returns).}170\label{tab:decile}171\begin{threeparttable}172\begin{tabular}{@{}lld{3.2}d{1.3}d{3.2}@{}}173\toprule174Sorting variable & Scheme & \multicolumn{1}{c}{Ann.\ ret.\ (\%)} & \multicolumn{1}{c}{Sharpe} & \multicolumn{1}{c}{$t$-stat} \\175\midrule176Implied kurtosis & Q5$-$Q1 & 44.04 & 2.328 & 19.84 \\177 & D10$-$D1 & 52.08 & 2.185 & 18.56 \\[3pt]178Put-call vol.\ ratio & Q5$-$Q1 & -30.29 & -2.478 & -21.80 \\179 & D10$-$D1 & -33.37 & -2.099 & -18.47 \\[3pt]180Skew$_{25\delta}$ & Q5$-$Q1 & -37.53 & -1.868 & -15.92 \\181 & D10$-$D1 & -52.51 & -1.896 & -16.11 \\182\bottomrule183\end{tabular}184\begin{tablenotes}[flushleft]\footnotesize185\item \textit{Notes.} Equal-weighted daily-rebalanced portfolios of individual stocks; long-short is the top minus the bottom portfolio. Annualization uses $\sqrt{52}$.186\end{tablenotes}187\end{threeparttable}188\end{table}189190\subsection{Temporal Stability and a Placebo Test}191192A recurrent concern with in-sample predictability is that it is concentrated in a handful of episodes and vanishes when they are excluded \citep{welch2008comprehensive,campbell2008predicting}. Table~\ref{tab:loyo} therefore re-estimates the two headline regressions sixteen times, excluding one calendar year at a time. The 5-day return $R^2$ stays within [4.7\%, 5.7\%] regardless of which year is dropped---removing 2018 (Volmageddon) has the largest effect---and the HAR$+$IV realized-variance $R^2$ stays within [0.478, 0.492] for every exclusion except 2020, whose removal lowers it to 0.414 by taking out the COVID variance burst. No single year drives either result.193194Finally, a placebo experiment permutes the entire block of option-implied features within each ticker (preserving their cross-feature correlation while destroying their alignment with dates) and re-runs the 5-day return regression. Across ten seeded draws, the placebo $R^2$ averages 0.0008 (maximum 0.0010) against an actual $R^2$ of 0.0529---a 66-fold gap confirming that the measured predictability reflects genuine temporal information rather than mechanical panel structure. Together with $t$-statistics that comfortably clear the conservative $t > 3.0$ multiple-testing hurdle of \citet{harvey2016and} for every key predictor, these checks make a data-mining explanation of the findings implausible.195196\begin{table}[H]197\centering198\caption{Leave-one-year-out panel $R^2$.}199\label{tab:loyo}200\begin{threeparttable}201\small202\begin{tabular}{@{}ld{1.4}d{1.4}@{\hskip 2.5em}ld{1.4}d{1.4}@{}}203\toprule204Excl.\ year & \multicolumn{1}{c}{\shortstack{5D ret\\$R^2$}} & \multicolumn{1}{c}{\shortstack{HAR+IV\\$R^2$}} &205Excl.\ year & \multicolumn{1}{c}{\shortstack{5D ret\\$R^2$}} & \multicolumn{1}{c}{\shortstack{HAR+IV\\$R^2$}} \\206\midrule2072010 & 0.0537 & 0.4852 & 2018 & 0.0469 & 0.4861 \\2082011 & 0.0535 & 0.4810 & 2019 & 0.0540 & 0.4858 \\2092012 & 0.0537 & 0.4831 & 2020 & 0.0505 & 0.4142 \\2102013 & 0.0538 & 0.4800 & 2021 & 0.0557 & 0.4817 \\2112014 & 0.0529 & 0.4833 & 2022 & 0.0500 & 0.4786 \\2122015 & 0.0514 & 0.4922 & 2023 & 0.0532 & 0.4883 \\2132016 & 0.0520 & 0.4844 & 2024 & 0.0572 & 0.4895 \\2142017 & 0.0547 & 0.4839 & 2025 & 0.0573 & 0.4900 \\215\bottomrule216\end{tabular}217\begin{tablenotes}[flushleft]\footnotesize218\item \textit{Notes.} Each row re-estimates the pooled 5-day return regression (ten predictors) and the HAR$+$IV 1-day RV regression (specification of Table~\ref{tab:subperiod}) on the full panel excluding the indicated calendar year. Full-sample baselines: 0.053 and 0.483.219\end{tablenotes}220\end{threeparttable}221\end{table}222