% ============================================================================= % Author: Simon-Pierre Boucher % Contact: contact@spboucher.ai % ============================================================================= % ============================================================================ % Robustness % ============================================================================ \section{Robustness}\label{sec:robust} \subsection{Alternative Standard Errors} Table~\ref{tab:robust_se} compares inference for the 5-day return regression under three SE specifications \citep{white1980heteroskedasticity,newey1987simple,cameron2011robust}; \citet{petersen2009estimating} shows that two-way clustering is the most conservative choice for finance panels of this type. Of the ten predictors, nine remain significant at 5\% under Newey-West HAC(5), and eight under double-clustered (ticker\,$+$\,date) SEs. Implied skewness ($t_{DC}=-4.94$), implied kurtosis ($t_{DC}=9.02$), PC volume ratio ($t_{DC}=-7.96$), and net gamma ($t_{DC}=5.39$) remain highly significant under all specifications. \begin{table}[H] \centering \caption{Robustness of 5-day return predictability to SE specification.} \label{tab:robust_se} \begin{threeparttable} \begin{tabular}{@{}ld{2.2}d{2.2}d{2.2}@{}} \toprule Variable & \multicolumn{1}{c}{HC1} & \multicolumn{1}{c}{NW(5)} & \multicolumn{1}{c}{DC} \\ \midrule $IV_{ATM,30d}$ & 3.86\sym{***} & 3.32\sym{***} & 1.40 \\ IV term slope & 4.83\sym{***} & 4.59\sym{***} & 2.12\sym{**} \\ Skew$_{25\delta}$ & -2.03\sym{**} & -0.93 & -0.39 \\ Implied skewness & -15.19\sym{***}& -13.66\sym{***}& -4.94\sym{***} \\ Implied kurtosis & 25.87\sym{***}& 25.64\sym{***}& 9.02\sym{***} \\ PC vol.\ ratio & -13.72\sym{***}& -16.69\sym{***}& -7.96\sym{***} \\ PC OI ratio & 12.33\sym{***}& 13.12\sym{***}& 4.61\sym{***} \\ Net gamma exp. & 15.55\sym{***}& 17.83\sym{***}& 5.39\sym{***} \\ $RV_{daily}$ & -17.18\sym{***}& -17.87\sym{***}& -7.27\sym{***} \\ $RV_{weekly}$ & 10.10\sym{***}& 7.95\sym{***}& 3.53\sym{***} \\ \bottomrule \end{tabular} \begin{tablenotes}[flushleft]\footnotesize \item \textit{Notes.} NW(5) = Newey-West with 5 lags. DC = double-clustered by ticker and date following \citet{cameron2011robust}. \sym{***}$p<0.01$; \sym{**}$p<0.05$. \end{tablenotes} \end{threeparttable} \end{table} \subsection{Subperiod Stability} Table~\ref{tab:subperiod} shows that both return predictability and the HAR+IV improvement are stable across nine subperiods. The 5-day return $R^2$ ranges from 3.1\% (post-GFC) to 14.4\% (COVID). The HAR+IV improvement over HAR-RV is positive in eight of nine subperiods, peaking at $+35.5\%$ during the low-vol era (2017--18) and $+34.4\%$ during the recovery (2023--24). \begin{table}[H] \centering \caption{Subperiod stability.} \label{tab:subperiod} \begin{threeparttable} \small \begin{tabular}{@{}ld{1.3}d{1.3}d{1.3}d{2.1}@{}} \toprule Subperiod & \multicolumn{1}{c}{\shortstack{5D ret\\$R^2$}} & \multicolumn{1}{c}{\shortstack{HAR-RV\\$R^2$}} & \multicolumn{1}{c}{\shortstack{HAR+IV\\$R^2$}} & \multicolumn{1}{c}{$\Delta R^2\,(\%)$} \\ \midrule Post-GFC (2010--12) & 0.031 & 0.344 & 0.392 & +14.0 \\ Bull (2013--16) & 0.075 & 0.342 & 0.336 & -1.7 \\ Low vol (2017--18) & 0.136 & 0.309 & 0.419 & +35.5 \\ Pre-COVID (2019) & 0.080 & 0.280 & 0.363 & +29.9 \\ COVID (2020) & 0.144 & 0.604 & 0.721 & +19.4 \\ Post-COVID (2021) & 0.036 & 0.376 & 0.440 & +16.9 \\ Rate hike (2022) & 0.084 & 0.387 & 0.471 & +21.8 \\ Recovery (2023--24) & 0.052 & 0.290 & 0.390 & +34.4 \\ Recent (2025) & 0.037 & 0.321 & 0.402 & +25.2 \\ \bottomrule \end{tabular} \begin{tablenotes}[flushleft]\footnotesize \item \textit{Notes.} $\Delta R^2 = (R^2_{HAR+IV} - R^2_{HAR}) / R^2_{HAR} \times 100$. \end{tablenotes} \end{threeparttable} \end{table} \subsection{VIX Regime Conditioning} Table~\ref{tab:vix_regime} shows that predictability increases with VIX level. The 5-day return $R^2$ rises from 2.7\% (VIX\,$<$\,15) to 19.8\% (VIX\,$>$\,35). Option-implied information is most valuable in high-uncertainty environments. \begin{table}[H] \centering \caption{Predictability by VIX regime.} \label{tab:vix_regime} \begin{threeparttable} \begin{tabular}{@{}ld{1.3}d{1.3}d{2.1}r@{}} \toprule VIX regime & \multicolumn{1}{c}{\shortstack{5D ret\\$R^2$}} & \multicolumn{1}{c}{\shortstack{HAR+IV\\$R^2$ (RV)}} & \multicolumn{1}{c}{\shortstack{Pct.\ of\\sample (\%)}} & $N$ \\ \midrule Very low ($<$15) & 0.027 & 0.359 & 35.7 & 43,971 \\ Low (15--20) & 0.040 & 0.341 & 35.0 & 40,221 \\ Medium (20--25) & 0.073 & 0.394 & 15.8 & 19,210 \\ High (25--35) & 0.080 & 0.385 & 11.0 & 13,139 \\ Crisis ($>$35) & 0.198 & 0.618 & 2.5 & 2,540 \\ \bottomrule \end{tabular} \end{threeparttable} \end{table} \subsection{Additional Controls, Quantile Regressions, and Ticker-Level Analysis} Adding controls (log volume, log OI, lagged returns) increases 5-day return $R^2$ from 5.3\% to 5.6\%; all core predictors retain significance. Quantile regressions at $\tau \in \{0.10, 0.25, 0.50, 0.75, 0.90\}$ \citep{koenker1978regression} show that the implied kurtosis and put-call ratio coefficients retain a stable sign and magnitude across the entire conditional return distribution. The distribution of ticker-level $R^2$ values for 5-day returns has mean 5.9\%, median 4.7\%, and interquartile range [2.0\%, 8.2\%], confirming that panel results are not driven by outliers. Rolling-window (252-day) analysis yields average $R^2$ of 8.0\% (std.\ 4.5\%, range [2.1\%, 17.9\%]) for 5-day returns and 41.2\% (std.\ 10.7\%) for 1-day RV forecasting. \subsection{Estimator and Specification Sensitivity} Two implementation choices could mechanically drive the baseline results: the winsorization cutoff and the HAC lag length. Table~\ref{tab:winsor} shows that neither does. Predictability rises modestly with heavier trimming---from $R^2 = 5.1\%$ at 0.5\% cutoffs to 6.2\% at 5\% cutoffs---because winsorization suppresses noise in the extreme tails, and even wholly unwinsorized data leave the key predictors highly significant ($t_{kurt} = 13.8$). Lengthening the Newey-West window from 5 to 22 lags \citep{newey1987simple}, enough to absorb the full overlap of the 5-day return plus a trading month, barely moves the $t$-statistics: implied kurtosis declines only from 25.6 to 22.1, and the set of significant predictors (9 of 10) is unchanged at every lag length. \begin{table}[H] \centering \caption{Winsorization and HAC-lag sensitivity (5-day returns).} \label{tab:winsor} \begin{threeparttable} \small \begin{tabular}{@{}ld{1.3}d{2.2}d{3.2}r@{\hskip 2em}ld{2.2}@{}} \toprule \multicolumn{5}{c}{Panel A: winsorization cutoff} & \multicolumn{2}{c}{Panel B: NW lags} \\ \cmidrule(r){1-5}\cmidrule(l){6-7} Cutoff & \multicolumn{1}{c}{$R^2$} & \multicolumn{1}{c}{$t_{kurt}$} & \multicolumn{1}{c}{$t_{PC}$} & \multicolumn{1}{c}{\#sig.} & Lags & \multicolumn{1}{c}{$t_{kurt}$} \\ \midrule None & 0.034 & 13.76 & -3.59 & 8/10 & NW(5) & 25.64 \\ 0.5\% & 0.051 & 36.51 & -21.31 & 10/10 & NW(10) & 23.89 \\ 1\% (baseline) & 0.053 & 40.07 & -21.48 & 9/10 & NW(22) & 22.09 \\ 2.5\% & 0.058 & 43.66 & -21.32 & 9/10 & & \\ 5\% & 0.062 & 45.73 & -20.86 & 9/10 & & \\ \bottomrule \end{tabular} \begin{tablenotes}[flushleft]\footnotesize \item \textit{Notes.} Panel A: pooled 5-day return regression (all 69 tickers) with HC1 $t$-statistics under alternative symmetric winsorization cutoffs. $t_{kurt}$ and $t_{PC}$ refer to implied kurtosis and the put-call volume ratio. \#sig.\ counts predictors significant at 5\%. Panel B: Newey-West $t$-statistic of implied kurtosis at alternative lag lengths (baseline data treatment). \end{tablenotes} \end{threeparttable} \end{table} \subsection{Rank-Based Information Coefficients} Parametric regressions may be sensitive to outliers and functional form even after winsorization. As a fully nonparametric check, I compute daily cross-sectional Spearman information coefficients (ICs)---the rank correlation between each predictor and the subsequent 5-day stock return---and test whether the time-series mean IC differs from zero. Table~\ref{tab:ic} confirms the regression evidence. Implied kurtosis carries a mean IC of $+0.098$ ($t = 16.8$), positive on 64.3\% of the 4{,}024 trading days; the put-call volume ratio ($-0.059$, $t = -13.8$) and the 25$\delta$ skew ($-0.056$, $t = -9.7$) are reliably negative. At the 1-day horizon no predictor achieves a mean IC above 0.02 in absolute value, mirroring the near-zero daily $R^2$ of Section~\ref{sec:results}. \begin{table}[H] \centering \caption{Daily cross-sectional Spearman information coefficients (stocks, 5-day returns).} \label{tab:ic} \begin{threeparttable} \begin{tabular}{@{}ld{1.4}d{3.2}d{2.1}@{}} \toprule Predictor & \multicolumn{1}{c}{Mean IC} & \multicolumn{1}{c}{$t$-stat} & \multicolumn{1}{c}{\% days $>0$} \\ \midrule Net gamma exposure & 0.1702 & 33.99 & 77.8 \\ Implied kurtosis & 0.0978 & 16.77 & 64.3 \\ IV term slope & 0.0328 & 6.30 & 54.6 \\ PC OI ratio & 0.0170 & 3.63 & 52.4 \\ $RV_{daily}$ & 0.0161 & 2.41 & 52.7 \\ $RV_{weekly}$ & 0.0120 & 1.73 & 51.2 \\ $IV_{ATM,30d}$ & 0.0085 & 1.09 & 53.0 \\ Implied skewness & -0.0188 & -3.30 & 47.4 \\ Skew$_{25\delta}$ & -0.0558 & -9.66 & 41.9 \\ PC vol.\ ratio & -0.0591 & -13.81 & 37.2 \\ \bottomrule \end{tabular} \begin{tablenotes}[flushleft]\footnotesize \item \textit{Notes.} For each trading day with at least 20 stocks, the Spearman rank correlation between the predictor and the 5-day forward return is computed across stocks; the table reports the time-series mean, its Fama-MacBeth-style $t$-statistic, and the share of days with a positive IC ($N = 4{,}024$ days maximum). \end{tablenotes} \end{threeparttable} \end{table} \subsection{Decile Sorts} The baseline quintile sorts could mask nonlinearity in the extreme tails of the sorting variables \citep{fama2008dissecting}. Table~\ref{tab:decile} repeats the three strongest sorts with decile portfolios. Sharpening the sort raises the spread in annualized returns---implied kurtosis rises from 44.0\% to 52.1\%, and the 25$\delta$ skew from $-37.5\%$ to $-52.5\%$---while Sharpe ratios remain essentially unchanged (2.33 vs.\ 2.19 for kurtosis), since the finer portfolios are noisier. The premia are therefore monotone rather than driven by a single extreme quintile. \begin{table}[H] \centering \caption{Quintile vs.\ decile long-short sorts (5-day returns).} \label{tab:decile} \begin{threeparttable} \begin{tabular}{@{}lld{3.2}d{1.3}d{3.2}@{}} \toprule Sorting variable & Scheme & \multicolumn{1}{c}{Ann.\ ret.\ (\%)} & \multicolumn{1}{c}{Sharpe} & \multicolumn{1}{c}{$t$-stat} \\ \midrule Implied kurtosis & Q5$-$Q1 & 44.04 & 2.328 & 19.84 \\ & D10$-$D1 & 52.08 & 2.185 & 18.56 \\[3pt] Put-call vol.\ ratio & Q5$-$Q1 & -30.29 & -2.478 & -21.80 \\ & D10$-$D1 & -33.37 & -2.099 & -18.47 \\[3pt] Skew$_{25\delta}$ & Q5$-$Q1 & -37.53 & -1.868 & -15.92 \\ & D10$-$D1 & -52.51 & -1.896 & -16.11 \\ \bottomrule \end{tabular} \begin{tablenotes}[flushleft]\footnotesize \item \textit{Notes.} Equal-weighted daily-rebalanced portfolios of individual stocks; long-short is the top minus the bottom portfolio. Annualization uses $\sqrt{52}$. \end{tablenotes} \end{threeparttable} \end{table} \subsection{Temporal Stability and a Placebo Test} A recurrent concern with in-sample predictability is that it is concentrated in a handful of episodes and vanishes when they are excluded \citep{welch2008comprehensive,campbell2008predicting}. Table~\ref{tab:loyo} therefore re-estimates the two headline regressions sixteen times, excluding one calendar year at a time. The 5-day return $R^2$ stays within [4.7\%, 5.7\%] regardless of which year is dropped---removing 2018 (Volmageddon) has the largest effect---and the HAR$+$IV realized-variance $R^2$ stays within [0.478, 0.492] for every exclusion except 2020, whose removal lowers it to 0.414 by taking out the COVID variance burst. No single year drives either result. Finally, a placebo experiment permutes the entire block of option-implied features within each ticker (preserving their cross-feature correlation while destroying their alignment with dates) and re-runs the 5-day return regression. Across ten seeded draws, the placebo $R^2$ averages 0.0008 (maximum 0.0010) against an actual $R^2$ of 0.0529---a 66-fold gap confirming that the measured predictability reflects genuine temporal information rather than mechanical panel structure. Together with $t$-statistics that comfortably clear the conservative $t > 3.0$ multiple-testing hurdle of \citet{harvey2016and} for every key predictor, these checks make a data-mining explanation of the findings implausible. \begin{table}[H] \centering \caption{Leave-one-year-out panel $R^2$.} \label{tab:loyo} \begin{threeparttable} \small \begin{tabular}{@{}ld{1.4}d{1.4}@{\hskip 2.5em}ld{1.4}d{1.4}@{}} \toprule Excl.\ year & \multicolumn{1}{c}{\shortstack{5D ret\\$R^2$}} & \multicolumn{1}{c}{\shortstack{HAR+IV\\$R^2$}} & Excl.\ year & \multicolumn{1}{c}{\shortstack{5D ret\\$R^2$}} & \multicolumn{1}{c}{\shortstack{HAR+IV\\$R^2$}} \\ \midrule 2010 & 0.0537 & 0.4852 & 2018 & 0.0469 & 0.4861 \\ 2011 & 0.0535 & 0.4810 & 2019 & 0.0540 & 0.4858 \\ 2012 & 0.0537 & 0.4831 & 2020 & 0.0505 & 0.4142 \\ 2013 & 0.0538 & 0.4800 & 2021 & 0.0557 & 0.4817 \\ 2014 & 0.0529 & 0.4833 & 2022 & 0.0500 & 0.4786 \\ 2015 & 0.0514 & 0.4922 & 2023 & 0.0532 & 0.4883 \\ 2016 & 0.0520 & 0.4844 & 2024 & 0.0572 & 0.4895 \\ 2017 & 0.0547 & 0.4839 & 2025 & 0.0573 & 0.4900 \\ \bottomrule \end{tabular} \begin{tablenotes}[flushleft]\footnotesize \item \textit{Notes.} Each row re-estimates the pooled 5-day return regression (ten predictors) and the HAR$+$IV 1-day RV regression (specification of Table~\ref{tab:subperiod}) on the full panel excluding the indicated calendar year. Full-sample baselines: 0.053 and 0.483. \end{tablenotes} \end{threeparttable} \end{table}