SPB Git

spb/wp9_uqo Public

UQO Working Paper No. 9 — A grand hedonic model of the Canadian housing market: decomposing structure and location value.

TeX 60.1% Python 39.8%
9.3 KB · 137 lines latex
Raw Blame History
1% Author: Simon-Pierre Boucher — contact@spboucher.ai2% ============================================================================3\section{Discussion}4\label{sec:discussion}5% ============================================================================67\subsection{Magnitudes in the light of the literature}89\paragraph{The location share.} Our headline decomposition---structure explains 46\% of10log-price variance, location a further 30 points---is a cross-sectional, listing-level11statement, but its magnitude aligns closely with what aggregate approaches find from the12opposite direction. \citet{davis2007price} estimate that land accounts for roughly 46\%13of the value of the U.S. housing stock, with much higher shares in coastal metros, and14\citet{knoll2017no} attribute about 80\% of the global house-price run-up since 1950 to15land. Land and ``location'' are not identical objects---our FSA premia capitalize16amenities and access as well as physical land scarcity17\citep{cheshire1995price,albouy2016cities}---but both measurements say the same thing:18the dwelling itself is the minority component of housing value. The ×9 span of our19neighbourhood premia likewise echoes the enormous cross-city dispersion of land values20documented by \citet{albouy2016cities} and is exactly the pattern predicted where supply21constraints bind heterogeneously across locations22\citep{saiz2010geographic,gyourko2013superstar}: Vancouver---sea, mountains and an23agricultural land reserve---anchors the top of our premium distribution, as24superstar-city dynamics would predict.2526\paragraph{Structural implicit prices.} The living-area elasticity of 0.55 sits inside27the range catalogued by \citet{sirmans2005composition} across decades of published28hedonics, and its decline from 0.66 to 0.55 as location controls tighten is the29classic signature of omitted-locational-quality bias that within-neighbourhood designs30are built to remove \citep{kiel2008location}. The near-zero conditional bedroom31coefficient reproduces one of the most robust findings of that meta-literature: floor32space, not room count, carries value. Our bathroom premium (11--15\%) is at the upper33end of published estimates, which we attribute to bathrooms proxying for unobserved34renovation status and finish quality in listing data that lack an age variable. The35diminishing marginal elasticity of floor space (from $\sim$0.65 at 60~m$^2$ to36$\sim$0.45 at 350~m$^2$) gives parametric confirmation of the curvature that37\citet{mcmillen2010issues} detect nonparametrically.3839\paragraph{Spatial dependence.} The 82\% reduction of Moran's~$I$ (0.46 to 0.08) speaks40directly to the oldest empirical worry in housing hedonics41\citep{dubin1988estimation,basu1998analysis}. It quantifies, for a national market, the42claim of \citet{bourassa2007spatial} and \citet{gibbons2012mostly} that geographic43controls---here, three-character postal geography---do most of the work that parametric44spatial-lag structures are designed to do. Our result does not make spatial econometrics45redundant: the residual $I=0.08$ is within-FSA dependence that an explicit46spatio-temporal model \citep{pace1998spatiotemporal,dube2013spatiotemporal} could47exploit, particularly for prediction. It does, however, shift the burden of proof toward48parsimony: a fixed effect per neighbourhood is transparent, imposes no weight-matrix49assumptions, and removes five-sixths of the spatial signal.5051\paragraph{Quantile and provincial heterogeneity.} The rising size elasticity across the52price distribution (0.56 at $\tau=0.1$ to 0.60 at $\tau=0.9$) and the rising lot53elasticity mirror the distributional patterns of \citet{zietz2008determinants}, and the54coastal-versus-Prairies gradient in the size elasticity (0.49 in BC to 0.66 in MB) is55consistent with the submarket literature's core claim that implicit prices---not just56price levels---vary across segments \citep{goodman1998housing,bourassa2003submarkets}.57This heterogeneity qualifies our own grand model: the absorbed intercepts allow every58neighbourhood its own price \emph{level}, but the structural slopes are pooled, and59Section~\ref{subsec:heterogeneity} shows those slopes move within an economically60meaningful band. Fully interacted (submarket-specific) coefficient systems are the61natural next step, at a substantial cost in transparency.6263\paragraph{Valuation accuracy.} A held-out $R^2$ of 0.764 and a median absolute error of6415.8\% place the transparent hedonic model within the accuracy bands reported in the65mass-appraisal comparison literature \citep{mccluskey2013prediction} and close to the66performance that machine-learning methods deliver on comparable tasks67\citep{mullainathan2017machine}. Two caveats temper the comparison: our target is the68\emph{list} price rather than the transaction price, and our evaluation is restricted to69neighbourhoods observed in training. Within those bounds, the result supports the70position of \citet{clapp2003semiparametric}: most of the predictive content of71``black-box'' valuation lies in the location surface, which a fixed-effect design72captures explicitly and auditably rather than implicitly.7374\subsection{Points of tension with prior findings}7576Three tensions deserve explicit statement. First, the submarket literature sometimes77finds that \emph{how} submarkets are defined matters little for prediction78\citep{bourassa2003submarkets}; our single-country evidence cannot adjudicate the79optimal geography, and FSAs---postal artifacts---are surely not it. The 0.08 residual80Moran's~$I$ suggests finer geography would still add value. Second,81\citet{kuminoff2010which} warn that hedonic estimates can be fragile to specification;82our robustness battery (six samples, quantiles, quadratic terms) addresses the83first-stage version of this concern, but any second-stage welfare use of our premia84would inherit the well-known identification problems \citep{epple1987hedonic}. Third,85our LOPO exercise shows structural prices transfer imperfectly (mean held-out $R^2$ of860.36, negative for Alberta), which cuts against reading the grand model's pooled slopes87as universal constants---and aligns with the heterogeneity that segmented-market models88predict \citep{goodman1998housing}.8990\subsection{Implications}9192For \emph{assessment and property taxation}, the decomposition implies that93neighbourhood-level value---not structure---is where most of the assessable base lives,94so assessment uniformity depends first on getting location surfaces right; the FSA95premia we publish are directly usable as such a surface. For \emph{index construction},96the stability of structural implicit prices across provinces supports pooled hedonic97indices with local intercepts \citep{hill2013hedonic}, while the quantile results98caution that constant-quality adjustment differs across market segments. For99\emph{automated valuation}, the results quantify the price of transparency: a fully100inspectable model concedes little accuracy on within-support predictions, echoing101\citet{mullainathan2017machine}'s point that prediction tasks discipline, rather than102replace, economic structure. For \emph{housing policy}, a ×9 neighbourhood premium span103within one country, concentrated in two metropolitan systems, is the observable imprint104of restricted supply against agglomeration demand105\citep{saiz2010geographic,gyourko2013superstar,combes2015empirics}; policies that expand106supply where premia are highest attack the decomposition's dominant term.107108\subsection{Limitations}109110Five limitations bound the interpretation. (i)~Prices are \emph{list} prices: sellers111set them strategically \citep{genesove2001loss} and they interact with search and112bargaining \citep{han2015microstructure}, so implicit prices measured on listings can113differ from transaction-based ones, especially in tight markets. (ii)~The data lack year114of construction, renovation status and interior quality; the neighbourhood effects115absorb the cross-neighbourhood component of this unobserved quality, but the within-FSA116structural estimates remain exposed to within-neighbourhood quality sorting.117(iii)~Lot information is sparse and noisily parsed from free text, limiting the118precision of the land-versus-structure margin at the listing level. (iv)~The analysis is119a single cross-section: it characterises the spatial and structural \emph{level} of120prices, not their dynamics, and cannot speak to the efficiency questions of the121repeat-sales tradition \citep{case1989efficiency}. (v)~FSA boundaries are postal122conveniences; premia estimated at this resolution average over genuinely finer123neighbourhood variation, as the residual spatial autocorrelation confirms.124125\subsection{Future work}126127The natural agenda follows the limitations: linking listings to closing prices and128time-on-market to estimate the list-to-sale wedge \citep{genesove2001loss,han2015microstructure};129enriching the attribute set with age and quality; decomposing the neighbourhood premia130into capitalized amenities---schools, transit, environmental quality---in the131quasi-experimental tradition \citep{black1999better,chay2005does,linden2008estimates};132modelling the residual within-FSA dependence with spatio-temporal structure133\citep{dube2013spatiotemporal}; extending the cross-section to a panel to study134dynamics and policy incidence; and estimating submarket-specific coefficient systems to135map the heterogeneity that our provincial results only sketch136\citep{goodman1998housing,bourassa2003submarkets}.137