% Author: Simon-Pierre Boucher — contact@spboucher.ai % % ═══════════════════════════════════════════════════════════════════════ % 7. DISCUSSION % ═══════════════════════════════════════════════════════════════════════ \section{Discussion} \label{sec:discussion} The results support four main messages, each associated with a specific modeling framework. For each, we interpret our estimates against the prior literature, noting where they agree with published findings, where they diverge, and what the divergences imply. \subsection{Message 1: OLS Remains Useful for Interpretable Associations} The semi-log OLS model explains 63.4\% of the variation in log listing prices using 62 regressors---within the range typical of large cross-sectional hedonic studies \citep{sirmans2005composition, malpezzi2003hedonic}---and its headline gradients align well with the meta-analytic record. The living-area elasticity of 0.63 is consistent with the concave size-price relationship documented across the 125 studies synthesized by \citet{sirmans2005composition}; the positive bathroom and garage gradients and the negative age gradient likewise match the modal signs in that synthesis. The negative conditional bedroom gradient---often treated as an anomaly---is in fact the standard result once total living area is held constant, reflecting a room-size trade-off rather than a distaste for bedrooms. The 33.4\% foreclosure discount is larger than the roughly 27\% forced-sale discount estimated from Massachusetts transactions by \citet{campbell2011forced}, as expected given that our estimate reflects asking-price positioning of distressed listings rather than realized sale prices. Regularized linear benchmarks (Ridge, Lasso, Elastic Net) achieve essentially identical performance ($R^2 = 0.628$--$0.630$), confirming that the parametric specification is not overfit and that the 20-percentage-point gap relative to XGBoost reflects genuine non-linearities and interactions, not estimation noise. The fixed-effects progression carries a sharper lesson. Moving from four Census regions ($R^2 = 0.634$) to state fixed effects ($0.678$) to 886 ZIP3 fixed effects ($0.725$) shows that a large share of "unexplained" variation is spatial, echoing the simulation-based advice of \citet{kuminoff2010which} that spatial fixed effects are the single most consequential specification choice in hedonic work, and the best-practice guidance of \citet{bishop2020best}. The ZIP3 model closes roughly half the OLS-to-XGBoost gap, implying that fine geographic partitioning ---not flexible functional form---accounts for much of what tree ensembles add. This dovetails with \citet{pace2020examining}, who find that machine-learning methods extract predictive content from the residuals of hedonic models precisely because those residuals retain spatial structure. The residual Moran's $I$ of 0.2745 after regional controls---comparable in magnitude to the residual autocorrelation documented in metropolitan samples by \citet{basu1998analysis} ---confirms that even our richest interpretable specification leaves the spatial error structure that the spatial econometrics literature has long modeled explicitly \citep{anselin1988spatial, lesage2009introduction, dubin1998spatial}. Two coefficient-level findings deserve confrontation with the capitalization literature. The negative school-rating coefficient contradicts the positive valuation of school quality identified by boundary-discontinuity designs \citep{black1999better, bayer2007unified}. We read this divergence not as evidence against school-quality capitalization but as a textbook illustration of why those designs exist: in a national cross-section, school ratings are collinear with property-tax rates (which enter separately), regional price levels, and unobserved neighborhood composition, so the partial coefficient is uninformative about the structural valuation. Analogously, the negative Walk Score coefficient---despite a documented walkability premium \citep{pivo2011walkability}---reflects conditioning simultaneously on bike and transit accessibility; the three scores are highly correlated, and their premia cannot be attributed variable-by-variable. Both cases reinforce the paper's interpretive stance: hedonic coefficients are conditional associations whose structural content depends on identification that a national listing cross-section does not provide \citep{kuminoff2013new}. Finally, the imputation sensitivity analysis shows the baseline lot-size gradient (0.02) is attenuated roughly threefold by state-median imputation, with the complete-case estimate (0.06) more plausible though subject to selection. The broader point---data-quality audits are not optional in scraped-data hedonics---is a practical extension of the measurement cautions in \citet{bishop2020best}. \subsection{Message 2: Quantile Regression Adds Distributional Insight} Quantile regression reveals that mean effects mask economically large distributional heterogeneity, and the shape of that heterogeneity is consistent with the housing literature. The declining living-area gradient (0.358 at $\tau = 0.10$ versus 0.251 at $\tau = 0.90$) mirrors the pattern in \citet{zietz2008determinants}, where square footage is valued relatively more in lower-priced homes. The rising pool and lot-size gradients toward the top of the distribution are consistent with amenity-driven luxury pricing documented at upper quantiles in Hong Kong \citep{mak2010quantile} and Sydney \citep{waltl2019variation}. The garage gradient---ten times larger at the bottom decile than the top ($z = 28.9$)---extends this logic: basic functional amenities differentiate lower-priced properties where they are scarce, and become nearly universal (hence unpriced) upstream. Such quantile-dependent pricing is one empirical face of housing-market segmentation \citep{goodman1998andrew}: submarkets defined by price tier value the same attribute bundle differently. Our inter-quantile Wald tests formalize what earlier housing QR studies typically showed graphically: coefficient equality is rejected for 11 of 13 attributes, so the conditional listing-price distribution is differentially stretched and compressed by attributes, not merely shifted. This is the cross-sectional analogue of the finding in \citet{mcmillen2008changes} that house price distributions move through coefficients rather than characteristics. For policy, the pattern implies that improvements to basic amenities are associated with the largest proportional listing-price differences at the lower end of the market, while luxury amenities differentiate the top---though we emphasize, again, that these are associations in asking prices, filtered through the seller-behavior mechanisms of \citet{genesove2001loss} and \citet{han2016role}, not causal renovation returns. The subsample stability analysis shows most quantile coefficients are highly robust (CV $< 8\%$, 100\% sign stability), with age the notable exception (CV = 78.6\%). Geographically heterogeneous vintage effects---historic premia in some markets, obsolescence discounts in others---are the natural reading, and they caution against interpreting any single national age gradient too literally. \subsection{Message 3: ML Performance Depends Critically on Validation Design} XGBoost improves on OLS by 20 percentage points of $R^2$ under random validation (0.833 vs.\ 0.630)---squarely within the 15--25\% error-reduction range reported in the ML-valuation literature \citep{bourassa2010predicting, park2015using, chen2016xgboost}. Had we stopped there, the paper would read as one more confirmation that boosting beats hedonics. The geographic holdout overturns that reading: under state-level holdout XGBoost attains only $R^2 = 0.425$--$0.547$ depending on specification. The 29--41-point collapse is strikingly similar in kind to what \citet{ploton2020spatial} document for ecological mapping models, and it validates, in a housing context, the warnings of \citet{roberts2017cross} and \citet{meyer2021predicting} about evaluating spatial predictions on randomly held-out data. The ablation analysis identifies \emph{what} fails to transfer. Neighborhood features contribute the largest random-validation gain ($+17.6$ pp) but only $+3.5$ pp under geographic holdout---consistent with the conjecture that scores like Walk Score and school ratings carry market-specific meaning. Validation research on Walk Score itself finds that its correspondence with underlying walkability constructs varies across metropolitan areas \citep{duncan2011validation}, which supports the mechanism, though we do not test it directly and flag it as plausible rather than established. More striking, the features most obviously "about" geography are actively harmful out-of-market: the full model with region dummies loses 7.6 pp of geographic holdout $R^2$ relative to the market-status stage, and adding raw coordinates ---the most flexible location encoding---produces the largest interpolation gain ($+3.7$ pp random) alongside a large extrapolation loss. Region dummies and coordinates let trees partition price levels by place, which is exactly what \citet{pace2020examining} show residual-based ML exploits, and exactly what cannot transfer to states never seen in training. We stress the scope of this conclusion, in light of the debate opened by \citet{wadoux2021spatial}: for \emph{interpolation}---valuing a property in a market represented in training data, the typical AVM production setting---random validation is informative and the geographic features earn their keep. Our claim concerns \emph{extrapolation} to unrepresented markets, where blocked validation is the appropriate benchmark \citep{valavi2019blockcv, meyer2021predicting} and where the AVM-evaluation literature already counsels reporting more than headline accuracy \citep{steurer2021metrics}. The practical rule for deployment follows: match the validation protocol to the deployment question, and if the question involves new markets, prefer the leaner, geography-free specification---it costs 0.3 pp of interpolation accuracy and buys 9.4 pp of extrapolation accuracy. \subsection{Message 4: SHAP Interprets Prediction, Not Economics} SHAP values identify living area, bathrooms, school quality, lot size, and region as the dominant predictive contributors, and this ranking is remarkably stable across three tree ensembles ($\rho = 0.89$--$0.99$). The stability result addresses a genuine concern in the explainability literature---that importance rankings are artifacts of a particular fitted model \citep{ribeiro2016should}---and parallels the motivation for model-agnostic explanation methods. But stability is not structure. Three cautions from the literature apply directly. First, SHAP decomposes predictions, not the data-generating process; a feature may earn a large SHAP value by proxying unobservables \citep{lundberg2017unified, lundberg2020local}. Second, correlated features share contributions in ways that defeat attribute-level economic interpretation---our school-rating case is again illustrative, ranking 3rd by SHAP while its OLS sign is negative and its credible structural valuation \citep{black1999better, bayer2007unified} is positive. Third, as \citet{rudin2019stop} argues, post-hoc explanation of a black box is not a substitute for an interpretable model when stakes are high; in our framework, the interpretable model (OLS/QR) and the black box answer different questions, and SHAP does not convert the latter into the former. In Rosen's terms: SHAP values are not implicit prices, and the stability we document should raise confidence in SHAP as a description of \emph{this prediction technology}, not as a measurement of \emph{market valuation}. \subsection{Synthesis} Across the four messages, a single theme recurs: each framework is reliable precisely within the question it was built to answer. OLS with rich spatial controls yields stable, literature-consistent conditional associations; quantile regression reveals formally significant distributional structure; boosting delivers real interpolation gains whose extrapolation content must be established by design, not assumed; and SHAP describes the predictive machine without licensing economic claims. The methodological corollary---that validation design and feature choice interact, and that random-split comparisons overstate ML advantages for out-of-market questions---is, we believe, the paper's most transferable lesson for both the hedonic and the AVM literatures.