% Author: Simon-Pierre Boucher — contact@spboucher.ai % %====================================================================== \section{Robustness Analysis} \label{sec:robustness} %====================================================================== To assess the sensitivity of our findings to modeling assumptions, sample composition, and estimation techniques, we conduct a battery of robustness checks. These analyses address five potential concerns: heteroskedasticity, multicollinearity among semantic features, sensitivity to outliers, distributional heterogeneity across price quantiles, and variable selection stability. All statistics reported in this section are produced by the replication pipeline accompanying the paper. \subsection{Heteroskedasticity Diagnostics} The Breusch-Pagan test \citep{breusch1979simple} applied to Model~D strongly rejects the null hypothesis of homoskedasticity (LM $= 768.0$, $p < 0.001$). This result formally justifies our use of HC3 heteroskedasticity-consistent standard errors throughout. The pattern of heteroskedasticity---residual variance increasing at both tails of the fitted value distribution (Figure~\ref{fig:diagnostics}, panel a)---is typical of housing price models and reflects the greater price dispersion in both the lowest-quality and highest-quality segments of the market \citep{goodman1995age}. \subsection{Variance Inflation Factors} The strong correlations among the 20 cosine similarity measures (Figure~\ref{fig:corr_matrix}) translate into substantial multicollinearity by conventional standards: variance inflation factors (VIFs) computed for Model~D average 17.2 across the similarity variables, and 16 of the 20 exceed the conventional threshold of 10 \citep{obrien2007caution}, with maxima for Renovated (41.6) and Entry-Level (40.7). This is an expected consequence of the design: all 20 features are projections of the same embedding onto related quality concepts, and they share the common influence of description verbosity. Three considerations indicate that this collinearity inflates standard errors without invalidating the analysis. First, as \citet{obrien2007caution} emphasizes, a high VIF is not by itself grounds for respecification: it signals imprecision, not bias, and the relevant question is whether the affected coefficients remain informative. Despite VIF-inflated standard errors, 16 of the 20 similarity measures are individually significant at the 5\% level, indicating that each retains sufficient unique variance for estimation. Second, the block-level evidence---the joint \fstat{} of 99.53 and the 6-percentage-point $R^2$ gain---is unaffected by collinearity among the block's members. Third, the coefficient estimates are stable across the winsorized and bootstrap re-estimations reported below, which would not be the case if the estimates were fragile artifacts of a near-singular design. The practical implication, developed in Section~\ref{sec:discussion}, is that individual semantic coefficients should be interpreted as partial associations within a correlated block of quality signals, with block-level conclusions being the most robust. \subsection{Quantile Regression Analysis} To examine whether the implicit prices of semantic dimensions vary across the price distribution, we estimate Model~D using quantile regression \citep{koenker1978regression, koenker2005quantile} at the 25th, 50th, and 75th percentiles. Table~\ref{tab:quantile} reports the coefficients for the six most economically significant semantic dimensions. \input{tables/tab_quantile} Three patterns emerge. First, the Luxury premium \textit{increases} monotonically with property value: a one-standard-deviation increase in luxury similarity is associated with a 9.7\% premium at the 25th percentile but a 17.3\% premium at the 75th percentile. This is consistent with quality complementarities---luxury features are valued more in properties that already occupy the upper market segment. Second, the negative effects of Motivated Seller and Needs Renovation language \textit{attenuate} at higher quantiles (the Motivated Seller discount shrinks from 8.7\% at $\tau = 0.25$ to a statistically insignificant 3.8\% at $\tau = 0.75$), suggesting that urgency and condition discounts are proportionally smaller for high-value properties where buyer pools are less price-sensitive. Third, the Modern/Contemporary and Land \& Nature premiums are remarkably stable across quantiles, indicating that these dimensions carry value throughout the market rather than in a particular segment. These heterogeneous effects across the conditional price distribution confirm that the OLS estimates represent averages across meaningfully different market segments, and they strengthen the economic interpretation of the semantic dimensions. \subsection{Sensitivity to Outliers} We re-estimate Model~D after trimming prices below the 1st and above the 99th percentile, removing 341 observations from the tails of the price distribution. The trimmed model yields an adjusted $R^2$ of 0.486, somewhat lower than the full-sample estimate of 0.511---as expected, since trimming removes exactly the price variation that the model exploits. Importantly, no coefficient changes sign relative to the full model, and all 16 similarity measures that are significant in the full model remain significant in the trimmed specification. The largest coefficient shifts are for Entry-Level ($-0.024$) and Motivated Seller ($+0.023$), both within the confidence intervals of the full-model estimates. This stability confirms that our results are not driven by extreme observations in either tail of the price distribution. \subsection{Bootstrap Inference} To verify that our HC3 standard errors provide reliable inference, we estimate Model~D using 1,000 nonparametric bootstrap replications \citep{efron1993introduction}. For each replication, we resample $n = 17{,}087$ observations with replacement and re-estimate the model, obtaining the empirical distribution of each coefficient. Bootstrap standard errors are within 9\% of the HC3 estimates for every semantic dimension---and within 5\% for 15 of the 20---and the bootstrap 95\% percentile confidence intervals closely match the HC3-based intervals. This concordance confirms the reliability of our asymptotic inference under the observed heteroskedasticity pattern. \subsection{Regularized Variable Selection} As an alternative to significance-based variable selection (Model~E), we employ Lasso \citep{tibshirani1996regression} and elastic net \citep{zou2005regularization} regularization with five-fold cross-validation over the standardized Model~D design. At the cross-validated penalty, both procedures retain a sparse set of four similarity variables (Needs Renovation, Premium Location, Income/Investment, and Family-Friendly) and shrink the remainder to zero. This aggressive selection is the textbook behavior of $\ell_1$ penalties under strong collinearity: when regressors are correlated at 0.66--0.97, the Lasso selects one representative per correlated block and discards near-duplicates, so the retained variables act as proxies for their blocks rather than as a list of the ``true'' non-zero effects \citep{zou2005regularization}. The exercise therefore complements, rather than replicates, the significance-based selection in Model~E: it confirms that the semantic block contains genuine predictive signal that survives penalization, while reinforcing the message of the VIF analysis that variable-by-variable attributions within the block are sensitive to the selection criterion. Conclusions in this paper that rely on individual dimensions (e.g., the Luxury gradient) are accordingly cross-checked against the quantile and bootstrap evidence above. \subsection{Comparison with Alternative Text Representations} To contextualize the predictive and inferential value of the reference-based approach, Table~\ref{tab:text_comparison} compares it with alternative representations of the same text, estimated on identical samples. For each method, we report the $\Delta R^2$ gain over the structural-only Model~A together with an assessment of interpretability. \input{tables/tab_text_comparison} Two results stand out. First, \textit{what} the text says matters more than \textit{how much}: description length alone adds 1.2 percentage points of $R^2$, whereas content-based features add 4.4--7.8 points. Second, there is a measurable price of interpretability. Twenty principal components of the raw embeddings---which capture 49.5\% of the embedding variance but have no economic meaning---outperform the 20 reference similarities in pure fit ($\Delta R^2$ of $+0.070$ vs.\ $+0.044$ without the length control). The reference-based approach deliberately trades roughly two percentage points of $R^2$ for features that support coefficient-level economic inference; for the inferential goals of hedonic analysis, we consider this trade worthwhile, while applications that only require prediction may prefer the unrestricted embedding space. \subsection{Summary of Robustness Findings} Table~\ref{tab:robustness_summary} summarizes the robustness checks. \input{tables/tab_robustness_summary} The overall picture is that the paper's central findings---the joint explanatory power of semantic features and the sign and approximate magnitude of the main implicit price gradients---are stable across estimation methods, sample definitions, and distributional assumptions. The diagnostics also delineate the limits of the evidence: the semantic dimensions are strongly collinear, so individual coefficients are measured imprecisely relative to the block as a whole, and regularized selection does not single out the same variables as significance testing. Extensions that would further strengthen the analysis---spatial fixed effects, alternative embedding models, and systematic reference-description sensitivity analysis---are discussed in Sections~\ref{sec:discussion} and~\ref{sec:conclusion}.