% Author: Simon-Pierre Boucher — contact@spboucher.ai % ============================================================================ \section{Beyond the scoreboard: what the models imply} \label{sec:ext} Accuracy is not the only output economists take from a hedonic model. This section asks whether the specification choices that move the scoreboard also move the two objects hedonic models are most often built to deliver: constant-quality price indices and implicit attribute prices. \subsection{Price indices are form-robust} Figure~\ref{fig:index} plots the constant-quality monthly index implied by the month fixed effects of four linear forms --- semi-log, log-log, quadratic and spline --- together with a gradient-boosting index obtained by repricing a fixed 20{,}000-sale reference portfolio at each month (an imputation \emph{\`a la} \citealp{hill2013hedonic}). All five track the same cycle --- the 2021--22 boom, the 2022--23 correction, the 2024--26 recovery --- and the four linear forms agree to within a few index points throughout. The choice that dominates the accuracy scoreboard barely perturbs the index: reassuring news for statistical agencies, and consistent with the RPPI handbook's pragmatism about form \citep{oecd2013handbook, silver2018house}. \begin{figure}[t] \centering \includegraphics[width=\textwidth]{fig_index.png} \caption{Constant-quality monthly price indices implied by competing specifications (January 2021 = 100). Linear forms: exponentiated month fixed effects. Gradient boosting: mean repriced value of a fixed 20{,}000-sale reference portfolio.} \label{fig:index} \end{figure} \subsection{Implicit prices: the machine agrees with the quadratic} Figure~\ref{fig:profiles} compares the ln-price profiles in building age and floor area implied by the quadratic OLS coefficients with the partial-% dependence profiles of the boosting model \citep{friedman2001greedy}. The two agree over the bulk of the data: steep initial depreciation that flattens with age, and strongly concave returns to floor space. The visible divergences are where the quadratic \emph{cannot} bend --- the machine sees a milder gradient for very old stock (a vintage effect) and a small new-construction premium spike --- but through the ages and areas where 90\% of sales live, the curves sit within a tenth of a log point. The boosting model's advantage evidently comes from higher-order interactions and spatial flexibility, \emph{not} from a different view of the main attribute gradients --- so the interpretable quadratic form remains a faithful summary of the price surface even where it loses the prediction race \citep{mullainathan2017machine, athey2019machine}. \begin{figure}[t] \centering \includegraphics[width=\textwidth]{fig_profiles.png} \caption{Implicit ln-price profiles: quadratic OLS coefficients versus gradient-boosting partial dependence, each normalized to its leftmost point. Panel A: building age. Panel B: floor area.} \label{fig:profiles} \end{figure} \subsection{Learning curves: data beat flexibility only at scale} Figure~\ref{fig:learning} and Table~\ref{tab:learning} trace accuracy against training size from 10{,}000 to 400{,}000 sales. The OLS quadratic is essentially flat from the very first step (18.4\% at $n=10{,}000$, 18.1\% at $n=400{,}000$) --- its bias floor binds almost immediately --- while the boosting model improves monotonically through the full sample (17.6\% $\to$ 14.5\%): the machine's advantage grows from under one error point at $n = 10{,}000$ to 3.6 points at $n = 400{,}000$. Flexible methods are not a free lunch in small markets: below a few tens of thousands of training sales, the machine's edge is modest, one reason rural and thin-% market AVMs underperform \citep{bogin2020house}. \begin{figure}[t] \centering \includegraphics[width=0.72\textwidth]{fig_learning.png} \caption{Learning curves: median absolute error against training-set size (random holdout fixed), OLS quadratic versus gradient boosting.} \label{fig:learning} \end{figure} \begin{table}[t] \centering \begin{threeparttable} \caption{Learning curves} \label{tab:learning} \small \input{../results/tables/learning} \begin{tablenotes}[flushleft]\footnotesize \item \textit{Notes:} MdAPE (\%) on the fixed random holdout, training on seeded subsamples of the indicated size. \end{tablenotes} \end{threeparttable} \end{table} \subsection{Where the machine wins: segments} Table~\ref{tab:segments} and Figure~\ref{fig:segments} decompose the random-holdout MdAPE by property class and municipality size, and the pattern identifies the machine's edge with striking precision: it is spatial resolution. The gap is enormous for condominiums (9.1\% versus 20.4\%) and in the largest cities (11.9\% versus 17.8\%) --- segments where structure is homogeneous but \emph{micro}-location varies hugely within a municipality, exactly the variation that municipality dummies cannot see and raw coordinates can. Once location is resolved, a condo is the most predictable object in the market. At the other extreme, cottages defeat both models (32.1\% versus 33.9\%): idiosyncratic, waterfront-driven stock is hard for everyone, and the machine's spatial surface cannot rescue attributes it does not observe. The complement of the assessment-inequity anatomy in our companion paper (UQO WP10) is instructive: the segments where assessors are most regressive (heterogeneous, land-heavy stock) are also those where \emph{no} standard technology, linear or boosted, prices accurately from roll attributes alone. \begin{figure}[t] \centering \includegraphics[width=0.8\textwidth]{fig_segments.png} \caption{Median absolute error by market segment (random holdout): OLS quadratic versus gradient boosting.} \label{fig:segments} \end{figure} \begin{table}[t] \centering \begin{threeparttable} \caption{Accuracy by market segment} \label{tab:segments} \small \input{../results/tables/segments} \begin{tablenotes}[flushleft]\footnotesize \item \textit{Notes:} MdAPE (\%) on the random holdout, by property class and by the number of 2021--2026 sales in the municipality. \end{tablenotes} \end{threeparttable} \end{table}