SPB Git

spb/wp3_uqo Public

UQO Working Paper No. 3 — Hedonic housing price models for the US: parametric, quantile, and machine-learning approaches.

TeX 77.8% Python 22.1%
15.4 KB · 229 lines latex
Raw Blame History
1% Author: Simon-Pierre Boucher — contact@spboucher.ai2%3% ═══════════════════════════════════════════════════════════════════════4% 2. LITERATURE REVIEW5% ═══════════════════════════════════════════════════════════════════════6\section{Literature Review}7\label{sec:literature}89This paper sits at the intersection of four literatures: the hedonic pricing10tradition in housing economics, the econometrics of distributional heterogeneity,11the rapidly growing machine-learning strand of property valuation, and the12methodological literature on validating predictive models under spatial dependence.13We review each in turn, emphasizing in every case how it bears on the design choices14of this study and where the gap addressed by this paper lies.1516\subsection{Hedonic Pricing: Origins, Theory, and the Interpretation of Implicit Prices}1718The practice of regressing prices on product characteristics predates its19theoretical justification by several decades. \citet{waugh1928quality} priced20quality attributes of vegetables, \citet{court1939hedonic} constructed hedonic21price indexes for automobiles, and \citet{griliches1961hedonic} revived the method22for quality-adjusted price measurement. In housing, \citet{ridker1967determinants}23provided the first regression-based estimates of how a local disamenity---air24pollution---is capitalized into residential property values, inaugurating the25property-value approach to valuing local public goods. The theoretical foundation26arrived with \citet{lancaster1966new}, who recast goods as bundles of27characteristics, and \citet{rosen1974hedonic}, who showed that in a competitive28market the equilibrium price schedule $P(\mathbf{z})$ traces out a double envelope29of bid and offer functions, so that the gradient $\partial P/\partial z_k$ equals30the marginal implicit price of attribute $z_k$. \citet{sirmans2005composition}31synthesize 125 empirical applications, documenting robust positive gradients for32living area, bathrooms, and garages and negative gradients for property age---the33same qualitative pattern our conditional-mean estimates reproduce at national scale.3435A central and often underappreciated distinction is that first-stage hedonic36gradients are equilibrium price phenomena, not preference parameters. Recovering37willingness-to-pay functions requires a second stage whose identification demands38exclusion restrictions rarely available in practice \citep{bartik1987estimation,39ekeland2004identification}. Modern syntheses formalize when hedonic estimates can40be trusted: \citet{kuminoff2013new} survey the equilibrium-sorting framework that41now underpins structural interpretation of housing-market regressions, and42\citet{bishop2020best} distill best practices for credible hedonic estimation,43emphasizing fine spatial controls, robustness to specification, and transparent44treatment of data limitations. Particularly relevant to our fixed-effects45progression, \citet{kuminoff2010which} show in large-scale simulations that46specifications with rich spatial fixed effects dramatically outperform sparse47cross-sectional specifications in recovering marginal willingness to pay when48omitted spatial variables correlate with amenities. Our region~$\rightarrow$49state~$\rightarrow$ ZIP3 exercise provides complementary, purely predictive50evidence on the same point. Throughout, we follow this literature's counsel and51interpret coefficients as conditional associations---listing-price capitalization52gradients---rather than structural demand parameters.5354\subsection{Functional Form and Specification}5556The semi-log specification became the workhorse of applied hedonics after57\citet{cropper1988choice} showed in Monte Carlo experiments that simple forms58outperform flexible ones when some attributes are omitted or measured with59error---precisely the situation of scraped listing data; \citet{hill2013hedonic}60surveys the subsequent evolution of hedonic specification practice in residential61applications. Earlier work had already62cautioned against atheoretic flexibility: \citet{halvorsen1981choice} demonstrated63the sensitivity of Box-Cox hedonic estimates, and \citet{anglin1996semiparametric}64documented gains from semiparametric estimation of the price function. This65literature frames one of our central questions: how much of the predictive66advantage of gradient-boosted trees reflects genuine non-linearity that the67semi-log form misses, and how much reflects something else---in our case, spatial68memorization.6970\subsection{Spatial Dependence in Housing Markets}7172Housing data are spatial data. \citet{can1992specification} showed that ignoring73spatial autocorrelation biases hedonic coefficient estimates;74\citet{dubin1998spatial} provides an accessible treatment of why residential75prices are spatially correlated (shared unobserved neighborhood attributes,76market-mediated spillovers); and \citet{basu1998analysis} document strong77spatial autocorrelation in Dallas transaction prices even after conditioning on78rich attribute sets. The formal apparatus of spatial econometrics---spatial lag,79spatial error, and spatial Durbin models---is developed in \citet{anselin1988spatial}80and \citet{lesage2009introduction}, with \citet{anselin2010thirty} providing a81retrospective. We deliberately stop short of estimating spatial econometric models:82our contribution is diagnostic. Using Moran's~$I$ \citep{moran1950notes} on OLS83residuals, we quantify the spatial dependence that remains after 62 attributes and84regional controls, and we show that it is a stable, structural feature of the data85rather than a sampling artifact. \citet{pace2020examining} pursue a complementary86strategy, using trees and forests to examine what information hedonic and spatial87model residuals still contain; their finding---that machine-learning methods88extract signal from residual spatial structure---anticipates our result that89tree-based models exploit location rather than only attribute non-linearity.9091\subsection{Quantile Regression and Distributional Heterogeneity}9293Quantile regression \citep{koenker1978regression, koenker2001quantile} replaces the94conditional mean with a family of conditional quantiles, allowing implicit prices95to vary across the price distribution. \citet{zietz2008determinants} introduced the96approach to housing, finding that square footage and other attributes are priced97differently at different quantiles; \citet{mak2010quantile} document quantile-varying98gradients in Hong Kong; and \citet{liao2012hedonic} combine quantile regression with99spatial methods. \citet{mcmillen2008changes} shows that changes in the100\emph{distribution} of house prices over time are driven largely by changes in101coefficients rather than in characteristics---evidence that quantile-specific102pricing is economically meaningful, not a statistical curiosity. More recently,103\citet{waltl2019variation} provides a comprehensive quantile analysis of the Sydney104market, documenting systematic variation across both price segments and locations.105Our contribution to this strand is scale and formality: we estimate five quantiles106on a national sample, verify coefficient stability across ten independent107subsamples, and subject the tail differences to formal inter-quantile Wald tests,108rejecting coefficient equality for 11 of 13 key attributes.109110\subsection{Listing Prices, Search, and Seller Behavior}111112Because our dependent variable is an asking price, the microstructure of listing113behavior matters for interpretation. Theory and evidence establish that list prices114are strategic objects: \citet{horowitz1992role} models the list price as a115commitment device in seller search; \citet{knight2002listing} shows that116overpricing lengthens time-on-market and reduces eventual sale prices;117\citet{genesove2001loss} demonstrate that loss aversion leads sellers---especially118those facing nominal losses---to set systematically higher asking prices; and119\citet{han2016role} show that the asking price plays a directing role in buyer120search, so that its information content varies across market segments. Two121implications follow for our estimates. First, listing-price gradients need not122equal transaction-price gradients, and the wedge is likely correlated with123attributes (luxury segments and distressed properties exhibit different list-to-sale124gaps). Second, this wedge affects all three modeling frameworks identically, so125\emph{comparisons across frameworks}---our primary object of interest---are126unaffected even where levels must be interpreted cautiously.127128\subsection{Capitalization of Local Public Goods and Amenities}129130Several of our regressors proxy local public goods, for which a mature131capitalization literature exists. \citet{oates1969effects} initiated the study of132property-tax and public-spending capitalization; \citet{black1999better} used133school-attendance boundaries to isolate the value of school quality, finding that134parents pay 2.5\% more for a 5\% increase in test scores---about half the naive135hedonic estimate; and136\citet{bayer2007unified} embed boundary discontinuities in an equilibrium sorting137model, showing that naive cross-sectional estimates confound school quality with138neighbor characteristics. This literature disciplines our reading of the139counterintuitive negative school-rating coefficient in the OLS results: without140boundary-style identification, school ratings in a national cross-section absorb141correlated neighborhood and fiscal variation (the property-tax rate enters142separately), and the coefficient should not be read as the value of school quality.143Similarly, \citet{pivo2011walkability} document a walkability premium using Walk144Score---the same measure we employ---while our specification, which conditions145simultaneously on bike and transit accessibility, illustrates how collinear146accessibility measures split the premium in ways that resist attribute-by-attribute147interpretation. Finally, the growing climate-capitalization literature148\citep{baldauf2020does, bernstein2019disaster, murfin2020risk} identifies a class149of price-relevant risk variables that are entirely missing from our data---a150limitation we return to in Section~\ref{sec:limitations}.151152\subsection{Machine Learning in Property Valuation}153154The econometrics profession has converged on a division of labor in which machine155learning excels at prediction ($\hat{y}$) problems while classical methods target156parameter ($\hat{\beta}$) problems \citep{mullainathan2017machine, varian2014big,157athey2019machine}. In real estate, this maps onto automated valuation models158(AVMs): \citet{bourassa2010predicting} compare methods for exploiting spatial159dependence in prediction; \citet{park2015using} document early machine-learning160gains in county-level housing data; \citet{kok2017big} describe the shift from161manual appraisal to big-data valuation; and \citet{steurer2021metrics} catalogue162the metrics appropriate for evaluating AVM performance, emphasizing that headline163$R^2$ figures conceal economically relevant tail behavior. The tree-ensemble164methods we deploy---random forests \citep{breiman2001random}, gradient boosting165\citep{friedman2001greedy}, and their modern implementations XGBoost166\citep{chen2016xgboost} and LightGBM \citep{ke2017lightgbm}---dominate tabular167prediction tasks of this kind. Against this backdrop, our contribution is to ask168not \emph{whether} boosting beats OLS (it does, by 20 percentage points of $R^2$169under random validation) but \emph{what that gap is made of}: the ablation design170decomposes it into feature-group contributions, and the geographic holdout reveals171that a substantial share reflects spatial memorization rather than transferable172attribute-price structure.173174\subsection{Explainability: SHAP and Its Limits}175176SHAP values \citep{lundberg2017unified}, rooted in the cooperative-game solution177concept of \citet{shapley1953value} and computable efficiently for trees via178TreeSHAP \citep{lundberg2020local}, have become the de facto standard for179interpreting ensemble predictions. Yet the explainability literature itself urges180caution: local surrogate explanations can be unstable \citep{ribeiro2016should},181and \citet{rudin2019stop} argues that post-hoc explanations of black-box models182should not be conflated with intrinsically interpretable modeling, particularly in183high-stakes settings. In the hedonic context the danger is specific: SHAP values184are prediction decompositions, sensitive to feature correlation, and do not satisfy185the equilibrium conditions under which Rosen's gradient equals a marginal implicit186price \citep{rosen1974hedonic, bishop2020best}. We therefore use SHAP for what it187can do---describe the predictive structure of the fitted model---and we probe the188robustness of that description by comparing importance rankings across three189different tree ensembles, finding rank correlations of 0.89--0.99.190191\subsection{Validation Under Spatial Dependence}192193A methodological literature largely developed in ecology and geostatistics warns194that random cross-validation overstates predictive skill whenever observations are195spatially dependent, because information leaks from training to test folds through196spatial proximity. \citet{roberts2017cross} systematize blocking strategies for197structured data; \citet{valavi2019blockcv} provide the standard software198implementation of spatial blocking; \citet{ploton2020spatial} demonstrate,199strikingly, that large-scale ecological mapping models with excellent random-CV200scores lose most of their skill under spatial validation; and201\citet{meyer2021predicting} formalize the ``area of applicability'' of spatial202prediction models. The lesson is not uncontested: \citet{wadoux2021spatial} show203that when the goal is map accuracy over a sampled region---an interpolation204problem---spatial cross-validation can be \emph{pessimistically} biased, and205design-based random validation is appropriate. This debate sharpens rather than206undermines our design: predicting prices in entirely unobserved states is an207\emph{extrapolation} task, for which held-out-region validation is the relevant208benchmark, while our random split answers the interpolation question. Reporting209both, and showing that feature sets rank differently under each, is precisely what210the debate prescribes. To our knowledge, this framing has not previously been211brought to bear on national-scale hedonic housing models.212213\subsection{Research Gap}214215Three gaps emerge from this review. First, the hedonic and AVM literatures rarely216meet: machine-learning papers report accuracy without engaging identification and217interpretation, while econometric papers report coefficients without assessing218predictive generalization. Comparative studies on a single large dataset that treat219OLS, quantile regression, and boosting as complementary lenses---each answering a220different question---remain scarce. Second, the spatial-validation insights of221ecology have barely penetrated housing economics, despite housing being a222canonically spatial asset; the interaction between geographic \emph{features} and223validation \emph{design} (our ablation and coordinate experiments) is essentially224unexplored at national scale. Third, the literature offers little guidance on how225explanation tools like SHAP behave across model families in housing applications.226This paper addresses all three, at the scale of 788,842 listings spanning every227U.S. state, while maintaining the interpretive discipline the identification228literature demands.229