spb/wp11_uqo Public
UQO Working Paper No. 11 — Half a million prices, twenty models: a systematic assessment of hedonic specifications.
TeX 54.7%
Python 45.2%
1% Author: Simon-Pierre Boucher — contact@spboucher.ai2% ============================================================================3\section{Beyond the scoreboard: what the models imply}4\label{sec:ext}56Accuracy is not the only output economists take from a hedonic model. This7section asks whether the specification choices that move the scoreboard8also move the two objects hedonic models are most often built to deliver:9constant-quality price indices and implicit attribute prices.1011\subsection{Price indices are form-robust}1213Figure~\ref{fig:index} plots the constant-quality monthly index implied by14the month fixed effects of four linear forms --- semi-log, log-log,15quadratic and spline --- together with a gradient-boosting index obtained16by repricing a fixed 20{,}000-sale reference portfolio at each month17(an imputation \emph{\`a la} \citealp{hill2013hedonic}). All five track18the same cycle --- the 2021--22 boom, the 2022--23 correction, the192024--26 recovery --- and the four linear forms agree to within a few20index points throughout. The choice that dominates the accuracy21scoreboard barely perturbs the index: reassuring news for statistical22agencies, and consistent with the RPPI handbook's pragmatism about form23\citep{oecd2013handbook, silver2018house}.2425\begin{figure}[t]26\centering27\includegraphics[width=\textwidth]{fig_index.png}28\caption{Constant-quality monthly price indices implied by competing29specifications (January 2021 = 100). Linear forms: exponentiated month30fixed effects. Gradient boosting: mean repriced value of a fixed3120{,}000-sale reference portfolio.}32\label{fig:index}33\end{figure}3435\subsection{Implicit prices: the machine agrees with the quadratic}3637Figure~\ref{fig:profiles} compares the ln-price profiles in building age38and floor area implied by the quadratic OLS coefficients with the partial-%39dependence profiles of the boosting model \citep{friedman2001greedy}.40The two agree over the bulk of the data: steep initial depreciation that41flattens with age, and strongly concave returns to floor space. The42visible divergences are where the quadratic \emph{cannot} bend --- the43machine sees a milder gradient for very old stock (a vintage effect) and44a small new-construction premium spike --- but through the ages and areas45where 90\% of sales live, the curves sit within a tenth of a log point. The boosting model's advantage evidently comes from46higher-order interactions and spatial flexibility, \emph{not} from a47different view of the main attribute gradients --- so the interpretable48quadratic form remains a faithful summary of the price surface even where49it loses the prediction race \citep{mullainathan2017machine,50athey2019machine}.5152\begin{figure}[t]53\centering54\includegraphics[width=\textwidth]{fig_profiles.png}55\caption{Implicit ln-price profiles: quadratic OLS coefficients versus56gradient-boosting partial dependence, each normalized to its leftmost57point. Panel A: building age. Panel B: floor area.}58\label{fig:profiles}59\end{figure}6061\subsection{Learning curves: data beat flexibility only at scale}6263Figure~\ref{fig:learning} and Table~\ref{tab:learning} trace accuracy64against training size from 10{,}000 to 400{,}000 sales. The OLS quadratic65is essentially flat from the very first step (18.4\% at $n=10{,}000$,6618.1\% at $n=400{,}000$) --- its bias floor binds almost immediately ---67while the boosting model improves monotonically through the full sample68(17.6\% $\to$ 14.5\%): the machine's advantage grows from under one69error point at $n = 10{,}000$ to 3.6 points at70$n = 400{,}000$. Flexible methods71are not a free lunch in small markets: below a few tens of thousands of72training sales, the machine's edge is modest, one reason rural and thin-%73market AVMs underperform \citep{bogin2020house}.7475\begin{figure}[t]76\centering77\includegraphics[width=0.72\textwidth]{fig_learning.png}78\caption{Learning curves: median absolute error against training-set79size (random holdout fixed), OLS quadratic versus gradient boosting.}80\label{fig:learning}81\end{figure}8283\begin{table}[t]84\centering85\begin{threeparttable}86\caption{Learning curves}87\label{tab:learning}88\small89\input{../results/tables/learning}90\begin{tablenotes}[flushleft]\footnotesize91\item \textit{Notes:} MdAPE (\%) on the fixed random holdout, training on92seeded subsamples of the indicated size.93\end{tablenotes}94\end{threeparttable}95\end{table}9697\subsection{Where the machine wins: segments}9899Table~\ref{tab:segments} and Figure~\ref{fig:segments} decompose the100random-holdout MdAPE by property class and municipality size, and the101pattern identifies the machine's edge with striking precision: it is102spatial resolution. The gap is enormous for condominiums (9.1\% versus10320.4\%) and in the largest cities (11.9\% versus 17.8\%) --- segments104where structure is homogeneous but \emph{micro}-location varies hugely105within a municipality, exactly the variation that municipality dummies106cannot see and raw coordinates can. Once location is resolved, a condo107is the most predictable object in the market. At the other extreme,108cottages defeat both models (32.1\% versus 33.9\%): idiosyncratic,109waterfront-driven stock is hard for everyone, and the machine's spatial110surface cannot rescue attributes it does not observe. The complement of111the assessment-inequity anatomy in our companion paper (UQO WP10) is112instructive: the segments where assessors are most regressive113(heterogeneous, land-heavy stock) are also those where \emph{no}114standard technology, linear or boosted, prices accurately from roll115attributes alone.116117\begin{figure}[t]118\centering119\includegraphics[width=0.8\textwidth]{fig_segments.png}120\caption{Median absolute error by market segment (random holdout):121OLS quadratic versus gradient boosting.}122\label{fig:segments}123\end{figure}124125\begin{table}[t]126\centering127\begin{threeparttable}128\caption{Accuracy by market segment}129\label{tab:segments}130\small131\input{../results/tables/segments}132\begin{tablenotes}[flushleft]\footnotesize133\item \textit{Notes:} MdAPE (\%) on the random holdout, by property class134and by the number of 2021--2026 sales in the municipality.135\end{tablenotes}136\end{threeparttable}137\end{table}138