spb/wp9_uqo Public
UQO Working Paper No. 9 — A grand hedonic model of the Canadian housing market: decomposing structure and location value.
TeX 60.1%
Python 39.8%
1% Author: Simon-Pierre Boucher — contact@spboucher.ai2% ============================================================================3\section{Related Literature}4\label{sec:literature}5% ============================================================================67This paper sits at the intersection of six literatures: the hedonic framework and its8identification; functional form and distributional heterogeneity in housing hedonics;9the treatment of space---submarkets, spatial econometrics and fixed effects; the theory10of amenity capitalization; urban spatial structure and the land share of housing value;11and property valuation, from price indices to automated valuation models. We review each12in turn and state where the present study enters.1314\subsection{The hedonic framework and its identification}1516The hedonic method has two intellectual roots. \citet{lancaster1966new} recast consumer17theory in terms of the characteristics embodied in goods rather than the goods themselves,18while \citet{court1939hedonic} and \citet{griliches1961hedonic} developed the empirical19machinery of quality-adjusted price indices. \citet{rosen1974hedonic} unified these ideas20in an equilibrium model in which the observed price schedule traces out the envelope of21buyers' bid functions and sellers' offer functions; its gradient with respect to a22characteristic identifies, at the margin, the implicit price of that characteristic. A23subsequent literature clarified exactly what the second stage---recovering preference24parameters from the implicit prices---can and cannot deliver. \citet{epple1987hedonic}25formalised the simultaneity problem that plagues second-stage demand estimation;26\citet{ekeland2004identification} show that nonlinearity of the hedonic price function27aids identification of preferences; \citet{bajari2005hedonic} develop a tractable demand28estimator with unobserved product characteristics; and \citet{kuminoff2010which} document29how sensitive welfare estimates are to specification. We deliberately remain at the first30stage and target the implicit-price schedule itself, which is the relevant object for31valuation, assessment and index construction, and which is identified under much weaker32assumptions than the underlying preferences.3334\subsection{Functional form and heterogeneity in the implicit prices}3536A large applied literature studies which attributes belong in a housing hedonic and what37functional form to impose. \citet{sirmans2005composition} catalogue the regressors used38in decades of published models and document the central role of living area, bathrooms,39lot size, age and location; their meta-tabulation also shows that bedroom coefficients40are frequently small or negative once floor area is controlled---the pattern we recover41for Canada. \citet{malpezzi2003hedonic} surveys functional-form choices and argues that42the semi-logarithmic specification---log price on linear (or log) characteristics---is a43robust default: it accommodates the right-skew of prices, yields coefficients44interpretable as approximate percentage effects \citep{halvorsen1980interpretation}, and45mitigates heteroskedasticity. We adopt this specification throughout. A complementary46literature relaxes its restrictions in two directions. \citet{mcmillen2010issues}47estimate fully nonparametric hedonic surfaces and show that parametric functional forms48can mask economically meaningful curvature---a concern we address directly with a49quadratic-in-log-area specification that reveals diminishing returns to floor space.50Quantile-regression approaches \citep{koenker1978regression,zietz2008determinants} show51that implicit prices differ systematically along the price distribution, so that a single52conditional-mean coefficient can be an incomplete summary; we document this for the53Canadian market in Section~\ref{subsec:quantile}. \emph{Gap:} these strands are almost54exclusively metropolitan or regional in scope; national-scale evidence on functional form55and distributional heterogeneity, estimated on a single consistent dataset, is rare.5657\subsection{Space in housing models: submarkets, spatial econometrics, fixed effects}5859Because housing is immobile, location is intrinsic to its value, and unobserved local60amenities induce strong spatial dependence in prices. The empirical importance of this61dependence is long established: \citet{dubin1988estimation} showed that ignoring62spatially autocorrelated errors distorts hedonic inference, and \citet{basu1998analysis}63documented pervasive residual autocorrelation in Dallas transactions.64\citet{can1992specification} and \citet{anselin1988spatial} formalised spatial65autocorrelation in hedonic errors, while \citet{pace1998spatiotemporal} and, in the66Canadian context, \citet{dube2013spatiotemporal} developed spatio-temporal weight-matrix67approaches. A parallel strand argues that housing markets are segmented into68\emph{submarkets} within which prices are set: \citet{goodman1998housing} formalise69market segmentation tests, and \citet{bourassa2003submarkets} show that accounting for70submarkets materially improves prediction accuracy---with the practical finding that71simple geographic definitions perform as well as statistically derived ones.72\citet{bourassa2007spatial} and \citet{case2004modeling} compare these approaches73head-to-head and conclude that geographic submarket controls capture most of what74parametric spatial models capture, at a fraction of the complexity.7576Our approach takes this conclusion to its logical end: rather than imposing a spatial77weight matrix, we absorb a fixed effect for each of 1{,}153 FSA neighbourhoods, allowing78the data to assign an arbitrary price level to each area and identifying structural79implicit prices from \emph{within-neighbourhood} variation. This is the housing analogue80of the high-dimensional fixed-effects designs pioneered in labour economics by81\citet{abowd1999high} and now standard in applied microeconomics, with computation82following \citet{guimaraes2010simple} and \citet{correia2017reghdfe} and inference83clustered at the neighbourhood level \citep{cameron2015practitioner}. The choice also84speaks to a methodological debate: \citet{gibbons2012mostly} caution that parametric85spatial-lag models are often hard to interpret causally and that flexible fixed effects86are frequently preferable, while \citet{lesage2009introduction} develop the87spatial-autoregressive alternative. We side with the fixed-effects approach but validate88it directly, testing for residual spatial autocorrelation with Moran's~$I$89\citep{moran1950notes} before and after absorbing the neighbourhood effects.90\emph{Gap:} submarket and spatial-fixed-effect studies are overwhelmingly single-city;91we are not aware of a prior implementation at the scale of an entire national market92with more than a thousand absorbed neighbourhoods.9394\subsection{Amenity capitalization}9596Why should a neighbourhood intercept capture economic value at all? The theoretical97answer runs from \citet{tiebout1956pure}, in which households sort across jurisdictions98according to local public goods, through \citet{oates1969effects}, who first measured99the capitalization of taxes and school spending into house prices, to100\citet{roback1982wages}, who embedded amenity capitalization in a general spatial101equilibrium of wages and rents. Quasi-experimental hedonic studies have since measured102the capitalization of specific amenities: school quality \citep{black1999better}, air103quality \citep{chay2005does} and local crime risk \citep{linden2008estimates}, while104\citet{cheshire1995price} and \citet{albouy2016cities} quantify the total land-value105content of local amenities. Our neighbourhood fixed effects deliberately bundle all such106capitalized amenities into a single FSA-level premium rather than attempting to107disentangle them; the estimated premium is therefore an upper envelope on the value of108location that future work can decompose amenity by amenity.109110\subsection{Urban spatial structure and the land share of housing value}111112The monocentric tradition of \citet{alonso1964location}, \citet{muth1969cities} and113\citet{mills1967aggregative} predicts that land rents---and hence house prices, holding114structure constant---decline with distance from employment centres, a prediction whose115modern quantitative form is estimated by \citet{ahlfeldt2015density} on uniquely sharp116Berlin data. Our distance-to-metro gradient is the reduced-form counterpart of this117prediction at national scale, and the agglomeration literature reviewed by118\citet{combes2015empirics} supplies the economic mechanism. A related macro-oriented119literature measures how much of housing value is land rather than structure:120\citet{davis2007price} put the U.S. residential land share near 46\% in aggregate,121\citet{knoll2017no} show that land appreciation explains about 80\% of the global122house-price increase since World War~II, and \citet{albouy2016cities} document enormous123cross-city dispersion in land values. On the supply side, \citet{saiz2010geographic}124shows that geographic constraints generate exactly the kind of persistent cross-location125price dispersion that our fixed effects absorb, and \citet{gyourko2013superstar}126formalise the resulting ``superstar city'' dynamics---with Vancouver and Toronto as127textbook candidates \citep{glaeser2005why,cmhc2018canadian}. \emph{Gap:} this literature128measures the land/location share with aggregate or city-level data; a listing-level129decomposition for an entire national market, in which the location share is estimated130directly from within-neighbourhood variation, complements it from the bottom up.131132\subsection{Valuation: price indices, mass appraisal, and machine learning}133134The hedonic regression is one of two workhorses of constant-quality house-price135measurement, alongside the repeat-sales method of \citet{bailey1963regression} that136underlies the \citet{case1989efficiency} index tradition; \citet{hill2013hedonic}137surveys and taxonomises the hedonic index literature. The same estimated surface doubles138as a valuation engine: \citet{clapp2003semiparametric} develops a location-value surface139for automated valuation, \citet{kiel2008location} systematise the ``location, location,140location'' decomposition of house-price determinants, and the mass-appraisal literature141benchmarks prediction accuracy across model classes \citep{mccluskey2013prediction}.142The rise of machine-learning valuation \citep{mullainathan2017machine} has sharpened the143question of what transparency costs in accuracy terms. \emph{Gap:} few studies report144honestly held-out accuracy for a fully transparent hedonic model at national scale; our145out-of-sample results provide exactly that benchmark.146147\subsection{Canadian evidence}148149Canada has a productive hedonic tradition, largely centred on single metropolitan areas:150the Quebec City program of Des Rosiers and coauthors measures the capitalization of151shopping centres \citep{desrosiers1996shopping} and landscaping quality152\citep{desrosiers2002landscaping}, and \citet{haider2000effects} apply spatial153autoregressive techniques to Toronto transactions. Policy work documents the spatial154dispersion of Canadian prices \citep{cmhc2018canadian}, with Vancouver and Toronto at155the expensive extreme \citep{glaeser2005why}. \emph{Gap:} to our knowledge no previous156study estimates a single hedonic system spanning the Canadian market coast to coast at157the listing level; the national scope, the 1{,}153 absorbed neighbourhood effects, and158the validated out-of-sample accuracy are, jointly, the contribution of this paper.159