% ============================================================================= % Author: Simon-Pierre Boucher % Contact: contact@spboucher.ai % ============================================================================= % ============================================================================ % Results % ============================================================================ \section{Results}\label{sec:results} \subsection{RQ1: Return Predictability} \subsubsection{Panel Regressions} Table~\ref{tab:rq1} reports pooled OLS results. Predictive power is dramatically stronger at the 5-day horizon: $R^2$ increases from 0.08\% (stocks, 1-day) to 4.8\% (stocks, 5-day), 12.4\% (ETFs), and 19.3\% (indices). \begin{table}[H] \centering \caption{Return predictability: pooled OLS panel regressions.} \label{tab:rq1} \begin{threeparttable} \begin{tabular}{@{}ld{1.4}d{1.4}d{1.4}d{1.4}d{1.4}d{1.4}@{}} \toprule & \multicolumn{3}{c}{1-Day return} & \multicolumn{3}{c}{5-Day return} \\ \cmidrule(lr){2-4}\cmidrule(lr){5-7} & \multicolumn{1}{c}{Stocks} & \multicolumn{1}{c}{ETFs} & \multicolumn{1}{c}{Idx} & \multicolumn{1}{c}{Stocks} & \multicolumn{1}{c}{ETFs} & \multicolumn{1}{c}{Idx} \\ \midrule $R^2$ & 0.0008 & 0.0014 & 0.0043 & 0.0476 & 0.1242 & 0.1930 \\ Adj.\ $R^2$ & 0.0007 & 0.0011 & 0.0032 & 0.0475 & 0.1239 & 0.1921 \\ $N$ & 79943 & 30152 & 8986 & 79943 & 30152 & 8986 \\ \# sig.\ (5\%) & 2 & 2 & 0 & 7 & 9 & 9 \\ \bottomrule \end{tabular} \begin{tablenotes}[flushleft]\footnotesize \item \textit{Notes.} Ten predictors: $IV_{ATM,30d}$, IV term slope, skew$_{25\delta}$, implied skewness, implied kurtosis, PC vol.\ ratio, PC OI ratio, net gamma exposure, $RV_t$, $RV_t^{(w)}$. HC1 standard errors. All variables winsorized at 1\%/99\% and standardized. \end{tablenotes} \end{threeparttable} \end{table} \subsubsection{Portfolio Sorts} Table~\ref{tab:sorts} reports quintile long-short portfolio performance for 5-day stock returns. The implied kurtosis sort delivers an annualized Sharpe of 2.33 ($t=19.84$). Sorting on the put-call volume ratio yields Sharpe $=-2.48$: stocks with the highest put-call ratio (bearish sentiment) earn the lowest future returns. \begin{table}[H] \centering \caption{Long-short portfolio performance (Q5$-$Q1, 5-day returns).} \label{tab:sorts} \begin{threeparttable} \begin{tabular}{@{}ld{2.2}d{2.2}d{1.3}d{2.3}r@{}} \toprule Sorting variable & \multicolumn{1}{c}{\shortstack{Ann.\\ret.\,(\%)}} & \multicolumn{1}{c}{\shortstack{Ann.\\vol.\,(\%)}} & \multicolumn{1}{c}{Sharpe} & \multicolumn{1}{c}{$t$-stat} & $N_{days}$ \\ \midrule Implied kurtosis & 44.04 & 18.91 & 2.328 & 19.843 & 3,777 \\ IV term slope & 11.33 & 20.30 & 0.558 & 4.004 & 2,676 \\ PC OI ratio & 5.27 & 13.27 & 0.397 & 3.492 & 4,024 \\[3pt] ATM IV (30d) & -1.76 & 26.13 & -0.067 & -0.575 & 3,777 \\ Implied skewness & -9.91 & 18.81 & -0.527 & -4.493 & 3,777 \\ Put-call vol.\ ratio & -30.29 & 12.22 & -2.478 & -21.796 & 4,024 \\ Skew$_{25\delta}$ & -37.53 & 20.10 & -1.868 & -15.918 & 3,778 \\ \bottomrule \end{tabular} \begin{tablenotes}[flushleft]\footnotesize \item \textit{Notes.} Equal-weighted quintile portfolios formed each trading day. Returns are 5-day (weekly). Annualization uses $\sqrt{52}$. $t$-statistics test $H_0$: mean daily L/S return~$=0$. \end{tablenotes} \end{threeparttable} \end{table} \subsubsection{Double Sorts} Table~\ref{tab:dsort} shows a $3\times 3$ sort on ATM~IV and implied skewness. Among high-IV stocks, those with high skewness earn $-47.8$~bps/week ($t=-6.17$), while those with low skewness earn $+41.8$~bps/week ($t=13.12$). The information content of skewness is amplified in high-uncertainty environments. \begin{table}[H] \centering \caption{Double sort: ATM\,IV $\times$ implied skewness $\to$ 5-day returns (bps/week).} \label{tab:dsort} \begin{threeparttable} \begin{tabular}{@{}ld{2.2}d{2.2}d{2.2}@{}} \toprule & \multicolumn{1}{c}{Low skew} & \multicolumn{1}{c}{Med.} & \multicolumn{1}{c}{High skew} \\ \midrule Low IV & 8.15 & 24.12 & 35.90 \\ Medium IV & 22.11 & 23.80 & 15.62 \\ High IV & 41.76 & 2.72 & -47.82 \\ \midrule \textit{High$-$Low IV} & \textit{33.61} & \textit{$-$21.40} & \textit{$-$83.72} \\ \bottomrule \end{tabular} \begin{tablenotes}[flushleft]\footnotesize \item \textit{Notes.} Tercile breakpoints computed cross-sectionally each day. \end{tablenotes} \end{threeparttable} \end{table} \subsection{RQ2: IV Surface vs.\ HAR-RV} Table~\ref{tab:rq2} reports in-sample $R^2$. The combined HAR-RV\,$+$\,IV-surface model achieves $R^2=0.442$ for 1-day~RV, a 23.3\% improvement over HAR-RV (0.359). The IV-only model ($R^2=0.390$) outperforms HAR-RV at the daily horizon. \begin{table}[H] \centering \caption{In-sample $R^2$: realized variance forecasting models.} \label{tab:rq2} \begin{threeparttable} \begin{tabular}{@{}ld{1.3}d{1.3}r@{}} \toprule Model & \multicolumn{1}{c}{1-Day RV} & \multicolumn{1}{c}{5-Day RV} & \multicolumn{1}{c}{$N$} \\ \midrule GARCH proxy & 0.326 & 0.648 & 264,380 \\ HAR-RV & 0.359 & 0.888 & 264,345 \\ IV only & 0.390 & 0.495 & 120,280 \\ IV surface & 0.391 & 0.496 & 119,093 \\ HAR-RV $+$ IV surface & 0.442 & 0.896 & 119,093 \\ \bottomrule \end{tabular} \end{threeparttable} \end{table} Diebold-Mariano tests show statistically significant MSE improvements for AAPL ($+10.3\%$, DM\,$=2.54$), AMD ($+11.7\%$, DM\,$=3.13$), CAT ($+32.7\%$, DM\,$=4.27$), and COST ($+24.4\%$, DM\,$=2.94$). \subsection{RQ3: Implied Correlation and Stress Prediction} Table~\ref{tab:rq3} shows that the correlation \textit{ratio} ($\hat{\rho}_{impl}/\hat{\rho}_{real}$) significantly predicts stress at all horizons ($t$: 3.91--8.30). The raw divergence is not significant; the multiplicative form captures the signal. Appendix~\ref{app:crises} details the behavior of the implied-correlation index around nine major market events. \begin{table}[H] \centering \caption{Stress prediction: linear probability model ($t$-statistics).} \label{tab:rq3} \begin{threeparttable} \begin{tabular}{@{}ld{2.2}d{2.2}d{2.2}@{}} \toprule Predictor & \multicolumn{1}{c}{5-Day} & \multicolumn{1}{c}{10-Day} & \multicolumn{1}{c}{20-Day} \\ \midrule Corr.\ ratio ($\hat{\rho}_{impl}/\hat{\rho}_{real}$) & 4.07\sym{***} & 8.30\sym{***} & 3.91\sym{***} \\ SPX ATM IV & -2.23\sym{**} & -4.07\sym{***} & -7.30\sym{***} \\ VIX level & 5.27\sym{***} & 6.69\sym{***} & 9.56\sym{***} \\ Corr.\ divergence & 0.00 & 0.00 & 0.00 \\ \midrule $R^2$ & 0.228 & 0.156 & 0.088 \\ $N$ & 2486 & 2481 & 2471 \\ \bottomrule \end{tabular} \begin{tablenotes}[flushleft]\footnotesize \item \textit{Notes.} HC1 standard errors. \sym{***}$p<0.01$; \sym{**}$p<0.05$. \end{tablenotes} \end{threeparttable} \end{table} \subsection{RQ4: Greeks Information Decay and Price Magnets} The information content of Greeks for RV forecasting increases monotonically with DTE: $R^2$ rises from 3.8\% (1-week) to 9.5\% (3-month), then declines slightly to 8.2\% (6-month). Return predictability is negligible across all buckets ($R^2<0.2\%$). For the price magnet analysis, prices moved toward the max-OI strike in only 47.0\% of 10,480 cases---below the 50\% null. The OI concentration coefficient is not significant ($t=1.06$), contradicting the ``max pain'' narrative. \subsection{RQ5: ML Models for SPX RV} Table~\ref{tab:rq5} reports the out-of-sample comparison. OLS HAR-RV dominates at every horizon ($R^2_{OOS}$: 0.437, 0.948, 0.994). ML models underperform (RF: 0.143; GBM: 0.095 for 1-day). RF feature importance shows that 2-week ATM~IV accounts for 50.8\% of total importance; the full ranking is reported in Appendix~\ref{app:importance}. \begin{table}[H] \centering \caption{Out-of-sample SPX realized variance forecasting (2020--2025).} \label{tab:rq5} \begin{threeparttable} \begin{tabular}{@{}lld{4.3}d{2.3}d{1.3}@{}} \toprule Model & Target & \multicolumn{1}{c}{MSE\,$(\times10^7)$} & \multicolumn{1}{c}{MAE\,$(\times10^4)$} & \multicolumn{1}{c}{$R^2_{OOS}$} \\ \midrule VIX benchmark & 1-Day & 1.835 & 1.371 & 0.336 \\ OLS: HAR-RV & 1-Day & 1.556 & 1.018 & 0.437 \\ Random Forest & 1-Day & 2.368 & 1.148 & 0.143 \\ Gradient Boosting & 1-Day & 2.501 & 1.282 & 0.095 \\ \midrule VIX benchmark & 5-Day & 18.870 & 5.425 & 0.597 \\ OLS: HAR-RV & 5-Day & 2.452 & 1.431 & 0.948 \\ \midrule VIX benchmark & 22-Day & 211.1 & 21.61 & 0.618 \\ OLS: HAR-RV & 22-Day & 3.518 & 1.709 & 0.994 \\ \bottomrule \end{tabular} \begin{tablenotes}[flushleft]\footnotesize \item \textit{Notes.} Training: 2010--2019 (2,515 days). Test: 2020--2025 (1,507 days). \end{tablenotes} \end{threeparttable} \end{table} \subsection{Granger Causality and VAR Analysis} Table~\ref{tab:granger} reports the Granger causality results. ATM~IV strongly Granger-causes RV ($F=62.4$, significant in 100\% of tickers). Returns Granger-cause skew ($F=13.6$, 95.7\%), indicating asymmetric volatility feedback. \begin{table}[H] \centering \caption{Granger causality tests (5 lags).} \label{tab:granger} \begin{threeparttable} \begin{tabular}{@{}ld{2.3}d{1.4}d{2.1}@{}} \toprule Hypothesis & \multicolumn{1}{c}{Avg.\ $F$} & \multicolumn{1}{c}{Avg.\ $p$} & \multicolumn{1}{c}{\% sig.} \\ \midrule ATM IV $\to$ RV & 62.381 & 0.0000 & 100.0 \\ RV $\to$ ATM IV & 17.933 & 0.1501 & 60.9 \\ Skew $\to$ RV & 20.792 & 0.0177 & 94.2 \\ Return $\to$ Skew & 13.576 & 0.0133 & 95.7 \\ ATM IV $\to$ Return & 1.799 & 0.2688 & 24.6 \\ PC ratio $\to$ Return & 1.010 & 0.4999 & 5.8 \\ \bottomrule \end{tabular} \begin{tablenotes}[flushleft]\footnotesize \item \textit{Notes.} Per-ticker bivariate $F$-tests. \% sig.\ is the share of tickers rejecting the null at 5\%. \end{tablenotes} \end{threeparttable} \end{table} The bivariate VAR(5) impulse response shows that a one-SD shock to ATM~IV produces a cumulative RV response of 0.67~SD at horizon~2, decaying to 0.27~SD at horizon~20. The forecast error variance decomposition reveals that IV shocks explain an increasing share of RV forecast error variance, reaching \textbf{73.8\%} at the 20-day horizon (Appendix~\ref{app:fevd}).