% Author: Simon-Pierre Boucher — contact@spboucher.ai % ============================================================================ \section{Related literature} \label{sec:lit} \subsection{The functional-form debate} Hedonic theory is silent about functional form. \citet{rosen1974hedonic} showed that the price surface is an equilibrium envelope of bids and offers, so nothing structural pins down its curvature --- form is an empirical choice. The classical exchange settled into a three-cornered debate: \citet{halvorsen1981choice} advocated data-driven Box--Cox transformations \citep{box1964analysis}; \citet{cassel1985cost} cautioned that flexible transformations optimize in-sample fit while degrading the implicit prices researchers actually use; and \citet{cropper1988choice}, in an influential simulation study, found that simple forms --- linear and semi-log --- are the most robust to the misspecification and omitted variables that characterize real data. Fifty years on, the semi-log remains the profession's default \citep{malpezzi2003hedonic, sirmans2005composition}, but systematic head-to-head evidence on modern samples remains scarce, and the classic studies predate the out-of-sample validation culture. Two further technical strands feed our design: retransformation bias in log models \citep{duan1983smearing}, which we correct with the smearing factor so that every form competes in levels; and forecast-evaluation methodology for house prices \citep{clapp2002predicting, steurer2021metrics}. \subsection{Space and time in hedonic models} A second literature concerns what the covariates cannot capture. House prices are spatially autocorrelated at fine scale \citep{dubin1998spatial, anselin1988spatial}, and the practical remedies range from submarket segmentation \citep{goodman1998housing} to spatial econometrics and simple location fixed effects; comparative studies generally find that flexibly absorbing location buys more predictive accuracy than any parametric spatial process \citep{case2004modeling, bourassa2010predicting, pace1997spatial}. In the Quebec context, \citet{desrosiers2000hedonic} document the importance of access and neighbourhood attributes in Quebec City. On the time dimension, hedonic price-index theory --- the time-dummy versus imputation debate --- is surveyed by \citet{hill2013hedonic} and codified in the RPPI handbook \citep{oecd2013handbook, silver2018house}; a byproduct of our month-FE models is a direct test of how sensitive a constant-quality index is to the underlying form. \subsection{Machine learning and automated valuation} The third strand is the rapid colonization of mass valuation by machine learning. \citet{mullainathan2017machine} frame prediction as ML's comparative advantage, with hedonic pricing as their running example; \citet{athey2019machine} survey the toolkit. Applied comparisons --- random forests \citep{breiman2001random}, gradient boosting \citep{friedman2001greedy, ke2017lightgbm} --- consistently report 20--40\% error reductions over linear hedonics \citep{mayer2019estimation, hong2020machine, kok2017big}, with accuracy degrading in thin rural markets \citep{bogin2020house}; \citet{gencay1996nonlinear} anticipated the pattern with semiparametric estimators. Three gaps motivate our design. First, most comparisons change several ingredients at once (features, sample, validation scheme), so the \emph{sources} of the ML advantage --- curvature? interactions? spatial flexibility? --- are rarely isolated; our axis design holds features constant and varies one choice at a time. Second, almost all published comparisons use random cross-validation, which leaks future price levels into training and flatters any flexible model; we score every model under a forward-in-time split as well. Third, comparisons seldom examine what the flexible models \emph{imply} --- indices, implicit prices --- which is what economists need; our extensions section does exactly that.