% Author: Simon-Pierre Boucher — contact@spboucher.ai % ============================================================================ \section{Related Literature} \label{sec:literature} % ============================================================================ This paper sits at the intersection of six literatures: the hedonic framework and its identification; functional form and distributional heterogeneity in housing hedonics; the treatment of space---submarkets, spatial econometrics and fixed effects; the theory of amenity capitalization; urban spatial structure and the land share of housing value; and property valuation, from price indices to automated valuation models. We review each in turn and state where the present study enters. \subsection{The hedonic framework and its identification} The hedonic method has two intellectual roots. \citet{lancaster1966new} recast consumer theory in terms of the characteristics embodied in goods rather than the goods themselves, while \citet{court1939hedonic} and \citet{griliches1961hedonic} developed the empirical machinery of quality-adjusted price indices. \citet{rosen1974hedonic} unified these ideas in an equilibrium model in which the observed price schedule traces out the envelope of buyers' bid functions and sellers' offer functions; its gradient with respect to a characteristic identifies, at the margin, the implicit price of that characteristic. A subsequent literature clarified exactly what the second stage---recovering preference parameters from the implicit prices---can and cannot deliver. \citet{epple1987hedonic} formalised the simultaneity problem that plagues second-stage demand estimation; \citet{ekeland2004identification} show that nonlinearity of the hedonic price function aids identification of preferences; \citet{bajari2005hedonic} develop a tractable demand estimator with unobserved product characteristics; and \citet{kuminoff2010which} document how sensitive welfare estimates are to specification. We deliberately remain at the first stage and target the implicit-price schedule itself, which is the relevant object for valuation, assessment and index construction, and which is identified under much weaker assumptions than the underlying preferences. \subsection{Functional form and heterogeneity in the implicit prices} A large applied literature studies which attributes belong in a housing hedonic and what functional form to impose. \citet{sirmans2005composition} catalogue the regressors used in decades of published models and document the central role of living area, bathrooms, lot size, age and location; their meta-tabulation also shows that bedroom coefficients are frequently small or negative once floor area is controlled---the pattern we recover for Canada. \citet{malpezzi2003hedonic} surveys functional-form choices and argues that the semi-logarithmic specification---log price on linear (or log) characteristics---is a robust default: it accommodates the right-skew of prices, yields coefficients interpretable as approximate percentage effects \citep{halvorsen1980interpretation}, and mitigates heteroskedasticity. We adopt this specification throughout. A complementary literature relaxes its restrictions in two directions. \citet{mcmillen2010issues} estimate fully nonparametric hedonic surfaces and show that parametric functional forms can mask economically meaningful curvature---a concern we address directly with a quadratic-in-log-area specification that reveals diminishing returns to floor space. Quantile-regression approaches \citep{koenker1978regression,zietz2008determinants} show that implicit prices differ systematically along the price distribution, so that a single conditional-mean coefficient can be an incomplete summary; we document this for the Canadian market in Section~\ref{subsec:quantile}. \emph{Gap:} these strands are almost exclusively metropolitan or regional in scope; national-scale evidence on functional form and distributional heterogeneity, estimated on a single consistent dataset, is rare. \subsection{Space in housing models: submarkets, spatial econometrics, fixed effects} Because housing is immobile, location is intrinsic to its value, and unobserved local amenities induce strong spatial dependence in prices. The empirical importance of this dependence is long established: \citet{dubin1988estimation} showed that ignoring spatially autocorrelated errors distorts hedonic inference, and \citet{basu1998analysis} documented pervasive residual autocorrelation in Dallas transactions. \citet{can1992specification} and \citet{anselin1988spatial} formalised spatial autocorrelation in hedonic errors, while \citet{pace1998spatiotemporal} and, in the Canadian context, \citet{dube2013spatiotemporal} developed spatio-temporal weight-matrix approaches. A parallel strand argues that housing markets are segmented into \emph{submarkets} within which prices are set: \citet{goodman1998housing} formalise market segmentation tests, and \citet{bourassa2003submarkets} show that accounting for submarkets materially improves prediction accuracy---with the practical finding that simple geographic definitions perform as well as statistically derived ones. \citet{bourassa2007spatial} and \citet{case2004modeling} compare these approaches head-to-head and conclude that geographic submarket controls capture most of what parametric spatial models capture, at a fraction of the complexity. Our approach takes this conclusion to its logical end: rather than imposing a spatial weight matrix, we absorb a fixed effect for each of 1{,}153 FSA neighbourhoods, allowing the data to assign an arbitrary price level to each area and identifying structural implicit prices from \emph{within-neighbourhood} variation. This is the housing analogue of the high-dimensional fixed-effects designs pioneered in labour economics by \citet{abowd1999high} and now standard in applied microeconomics, with computation following \citet{guimaraes2010simple} and \citet{correia2017reghdfe} and inference clustered at the neighbourhood level \citep{cameron2015practitioner}. The choice also speaks to a methodological debate: \citet{gibbons2012mostly} caution that parametric spatial-lag models are often hard to interpret causally and that flexible fixed effects are frequently preferable, while \citet{lesage2009introduction} develop the spatial-autoregressive alternative. We side with the fixed-effects approach but validate it directly, testing for residual spatial autocorrelation with Moran's~$I$ \citep{moran1950notes} before and after absorbing the neighbourhood effects. \emph{Gap:} submarket and spatial-fixed-effect studies are overwhelmingly single-city; we are not aware of a prior implementation at the scale of an entire national market with more than a thousand absorbed neighbourhoods. \subsection{Amenity capitalization} Why should a neighbourhood intercept capture economic value at all? The theoretical answer runs from \citet{tiebout1956pure}, in which households sort across jurisdictions according to local public goods, through \citet{oates1969effects}, who first measured the capitalization of taxes and school spending into house prices, to \citet{roback1982wages}, who embedded amenity capitalization in a general spatial equilibrium of wages and rents. Quasi-experimental hedonic studies have since measured the capitalization of specific amenities: school quality \citep{black1999better}, air quality \citep{chay2005does} and local crime risk \citep{linden2008estimates}, while \citet{cheshire1995price} and \citet{albouy2016cities} quantify the total land-value content of local amenities. Our neighbourhood fixed effects deliberately bundle all such capitalized amenities into a single FSA-level premium rather than attempting to disentangle them; the estimated premium is therefore an upper envelope on the value of location that future work can decompose amenity by amenity. \subsection{Urban spatial structure and the land share of housing value} The monocentric tradition of \citet{alonso1964location}, \citet{muth1969cities} and \citet{mills1967aggregative} predicts that land rents---and hence house prices, holding structure constant---decline with distance from employment centres, a prediction whose modern quantitative form is estimated by \citet{ahlfeldt2015density} on uniquely sharp Berlin data. Our distance-to-metro gradient is the reduced-form counterpart of this prediction at national scale, and the agglomeration literature reviewed by \citet{combes2015empirics} supplies the economic mechanism. A related macro-oriented literature measures how much of housing value is land rather than structure: \citet{davis2007price} put the U.S. residential land share near 46\% in aggregate, \citet{knoll2017no} show that land appreciation explains about 80\% of the global house-price increase since World War~II, and \citet{albouy2016cities} document enormous cross-city dispersion in land values. On the supply side, \citet{saiz2010geographic} shows that geographic constraints generate exactly the kind of persistent cross-location price dispersion that our fixed effects absorb, and \citet{gyourko2013superstar} formalise the resulting ``superstar city'' dynamics---with Vancouver and Toronto as textbook candidates \citep{glaeser2005why,cmhc2018canadian}. \emph{Gap:} this literature measures the land/location share with aggregate or city-level data; a listing-level decomposition for an entire national market, in which the location share is estimated directly from within-neighbourhood variation, complements it from the bottom up. \subsection{Valuation: price indices, mass appraisal, and machine learning} The hedonic regression is one of two workhorses of constant-quality house-price measurement, alongside the repeat-sales method of \citet{bailey1963regression} that underlies the \citet{case1989efficiency} index tradition; \citet{hill2013hedonic} surveys and taxonomises the hedonic index literature. The same estimated surface doubles as a valuation engine: \citet{clapp2003semiparametric} develops a location-value surface for automated valuation, \citet{kiel2008location} systematise the ``location, location, location'' decomposition of house-price determinants, and the mass-appraisal literature benchmarks prediction accuracy across model classes \citep{mccluskey2013prediction}. The rise of machine-learning valuation \citep{mullainathan2017machine} has sharpened the question of what transparency costs in accuracy terms. \emph{Gap:} few studies report honestly held-out accuracy for a fully transparent hedonic model at national scale; our out-of-sample results provide exactly that benchmark. \subsection{Canadian evidence} Canada has a productive hedonic tradition, largely centred on single metropolitan areas: the Quebec City program of Des Rosiers and coauthors measures the capitalization of shopping centres \citep{desrosiers1996shopping} and landscaping quality \citep{desrosiers2002landscaping}, and \citet{haider2000effects} apply spatial autoregressive techniques to Toronto transactions. Policy work documents the spatial dispersion of Canadian prices \citep{cmhc2018canadian}, with Vancouver and Toronto at the expensive extreme \citep{glaeser2005why}. \emph{Gap:} to our knowledge no previous study estimates a single hedonic system spanning the Canadian market coast to coast at the listing level; the national scope, the 1{,}153 absorbed neighbourhood effects, and the validated out-of-sample accuracy are, jointly, the contribution of this paper.