spb/wp2_uqo Public
UQO Working Paper No. 2 — Decoding Real Estate Descriptions: text-based hedonic analysis of housing listings.
TeX 73.8%
Python 26%
1% Author: Simon-Pierre Boucher — contact@spboucher.ai2%3%======================================================================4\section{Results} \label{sec:results}5%======================================================================67\subsection{Model Comparison}89Table~\ref{tab:model_comparison} reports goodness-of-fit statistics for the five model specifications defined in Equations~(\ref{eq:modelA})--(\ref{eq:modelE}). Figure~\ref{fig:model_comparison} visualizes the comparison.1011\input{tables/tab_model_comparison}1213\begin{figure}[!htbp]14\centering15\includegraphics[width=0.8\textwidth]{fig1_model_comparison.pdf}16\caption{Goodness-of-fit comparison across hedonic model specifications.}17\label{fig:model_comparison}18\begin{minipage}{0.9\textwidth}19\footnotesize20Notes: This figure compares $R^2$ and adjusted $R^2$ across the five nested model specifications defined in Equations~(\ref{eq:modelA})--(\ref{eq:modelE}). The dashed line marks the structural baseline (Model A).21\end{minipage}22\end{figure}2324Several results merit discussion. The structural model (A) explains 45.2\% of log-price variation, consistent with the typical range reported in the hedonic literature for single-family homes \citep{sirmans2005composition}. Adding description length alone (Model~B) improves the $R^2$ by 1.2 percentage points, confirming that the quantity of listing text is price-informative---an indirect quality signal.2526Adding the 20 semantic similarity measures without controlling for length (Model~C) yields a substantially larger improvement of 4.4 percentage points. This demonstrates that the \textit{content} of the text matters beyond its mere length. The full model (D), which includes both text length and all semantic features, achieves an $R^2$ of 0.512---an improvement of 6.0 percentage points over the structural baseline.2728The parsimonious model (E), retaining only the 16 individually significant similarity measures, achieves nearly identical fit ($R^2 = 0.512$, BIC = 24,425 vs.\ 24,448 for Model~D). The lower BIC of Model~E suggests it should be preferred on the basis of model parsimony. Four reference dimensions---Renovated, Income/Investment, Heritage/Character, and Pool \& Landscaping---do not contribute independently after controlling for the other variables. The complete coefficient estimates for Model~E are reported in Appendix Table~\ref{tab:parsimonious}.2930Figure~\ref{fig:r2_decomp} decomposes the explained variance in the full model. Structural variables account for 88.3\% of the total $R^2$, text length contributes 2.4\%, and the semantic similarity measures add 9.3\%.3132\begin{figure}[!htbp]33\centering34\includegraphics[width=0.55\textwidth]{fig8_r2_decomposition.pdf}35\caption{Decomposition of explained variance ($R^2$) in the full model (D).}36\label{fig:r2_decomp}37\begin{minipage}{0.9\textwidth}38\footnotesize39Notes: Structural variables account for 88.3\% of the total explained variance, description length contributes 2.4\%, and the 20 semantic similarity measures add 9.3\%.40\end{minipage}41\end{figure}4243\subsection{Joint Significance Test}4445The \fstat{} for the joint significance of all 21 text-based variables (length + 20 similarities) comparing Model~D to Model~A yields:46\begin{equation*}47F = 99.53, \quad p < 0.00148\end{equation*}49This decisively rejects the null hypothesis that listing text contains no price-relevant information beyond structural characteristics. The test statistic is very large, reflecting the substantial improvement in model fit.5051\subsection{Structural Variable Estimates}5253Figure~\ref{fig:structural_coef} presents the coefficient estimates for structural variables in the full model.5455\begin{figure}[!htbp]56\centering57\includegraphics[width=0.8\textwidth]{fig10_structural_coefficients.pdf}58\caption{Structural variable coefficients in the full model (D), with 95\% CI.}59\label{fig:structural_coef}60\begin{minipage}{0.9\textwidth}61\footnotesize62Notes: Coefficients represent the percentage change in price associated with a one-standard-deviation change in each structural variable. Error bars show 95\% confidence intervals based on HC3 robust standard errors.63\end{minipage}64\end{figure}6566Bathrooms have the largest structural effect (+34.5\% per standard deviation), followed by half-bathrooms (+17.1\%), description length (+10.7\%), parking (+9.6\%), and stories (+7.7\%). Bedrooms have a modest positive effect (+3.2\%), and lot size is statistically insignificant---consistent with the noisy, unit-inconsistent measurement of this field in the raw data (Section~\ref{sec:data}).6768\subsection{Semantic Dimension Estimates}6970Table~\ref{tab:full_results} reports the complete coefficient estimates from Model~D. Figure~\ref{fig:coefficient_plot} presents the semantic coefficients graphically.7172\input{tables/tab_full_results}7374\begin{figure}[!htbp]75\centering76\includegraphics[width=0.8\textwidth]{fig2_coefficient_plot.pdf}77\caption{Hedonic price impact of semantic similarity dimensions (Model D). Green bars indicate positive effects; red bars indicate negative effects. Gray bars are not statistically significant at the 5\% level. Error bars show 95\% confidence intervals.}78\label{fig:coefficient_plot}79\end{figure}8081\subsubsection{Positive Price Associations}8283Three semantic dimensions exhibit large, positive, and highly significant associations with listing prices.8485\textbf{Modern/Contemporary} ($+16.4\%$, $p < 0.001$). Listings semantically closer to modern/contemporary reference descriptions---emphasizing contemporary architecture, minimalist design, smart home technology, and energy-efficient construction---are associated with the highest price premium. This is consistent with revealed preferences for modern aesthetics and functional design in the Quebec housing market.8687\textbf{Luxury} ($+14.2\%$, $p < 0.001$). Language describing premium finishes (quartz, granite, hardwood), gourmet kitchens, spa-like bathrooms, and home automation systems is associated with a substantial price premium. This dimension captures quality variation that is invisible to standard structural variables: two properties with the same number of bathrooms can differ enormously in finish quality, and the listing text registers this difference.8889\textbf{Land \& Nature} ($+13.0\%$, $p < 0.001$). Descriptions of wooded lots, professional landscaping, privacy, and access to natural surroundings are associated with a significant positive price gradient, consistent with the well-documented value of outdoor amenity space in residential markets.9091\textbf{Waterfront} ($+4.7\%$, $p < 0.001$) and \textbf{Panoramic View} ($+3.0\%$, $p = 0.01$) capture location-specific amenities---access to water and scenic views---that are not reflected in standard structural variables but are associated with higher prices.9293\subsubsection{Negative Price Associations}9495Several dimensions exhibit negative coefficients that, upon closer examination, reflect market segmentation rather than value destruction per se (see Section~\ref{sec:identification} for the identification discussion).9697\textbf{Family-Friendly} ($-11.6\%$, $p < 0.001$). This is the most strongly negative dimension. Descriptions emphasizing proximity to schools, parks, and family amenities characterize lower-priced suburban markets. After controlling for structural characteristics, this language likely proxies for suburban location rather than indicating that family-oriented features reduce value. This coefficient should be interpreted with the caveat that it may absorb locational price variation in the absence of spatial controls.9899\textbf{New Construction} ($-10.5\%$, $p < 0.001$). The negative coefficient is counterintuitive at first glance, since ``new'' is typically a premium attribute. However, in Quebec's real estate geography, new residential developments are concentrated in peripheral suburban areas (e.g., Mirabel, Mascouche, Lachute) where land costs are lower. The ``new construction'' semantic dimension likely captures this locational sorting rather than an intrinsic quality discount; municipality fixed effects, once available, would disambiguate the two (Section~\ref{sec:discussion}).100101\textbf{Motivated Seller} ($-8.5\%$, $p < 0.001$). This is consistent with the information asymmetry literature \citep{levitt2008information}: language signaling seller urgency (``must sell,'' ``price reduced,'' ``estate sale'') reveals a weak bargaining position and is associated with a price discount relative to comparable properties with non-urgent descriptions.102103\textbf{Needs Renovation} ($-7.9\%$, $p < 0.001$). Descriptions indicating that a property requires work are associated with lower prices, consistent with the cost of deferred maintenance being capitalized into the listing price.104105\subsection{Semantic Profiles Across Price Quintiles}106107Figure~\ref{fig:quintile_heatmap} visualizes the mean cosine similarity for each reference dimension across price quintiles, revealing how the semantic ``fingerprint'' of listing descriptions varies with price level.108109\begin{figure}[!htbp]110\centering111\includegraphics[width=0.85\textwidth]{fig4_quintile_heatmap.pdf}112\caption{Mean cosine similarity by price quintile across 20 reference dimensions.}113\label{fig:quintile_heatmap}114\begin{minipage}{0.9\textwidth}115\footnotesize116Notes: Warmer colors indicate higher similarity. Each cell reports the mean cosine similarity between listings in the given price quintile and the corresponding reference description.117\end{minipage}118\end{figure}119120The gradient reveals that lower-priced properties score higher on nearly all dimensions due to longer, more detailed descriptions. The relative differences across dimensions are nevertheless informative: moving from the bottom to the top quintile, the steepest declines occur for Motivated Seller, Entry-Level, and Premium Location, while Land \& Nature, Modern/Contemporary, and Panoramic View decline the least---precisely the dimensions that carry positive premiums in the regression results.121122\subsection{Inter-Reference Correlations}123124Figure~\ref{fig:corr_matrix} presents the correlation matrix between the 20 similarity dimensions.125126\begin{figure}[!htbp]127\centering128\includegraphics[width=0.75\textwidth]{fig5_correlation_matrix.pdf}129\caption{Pearson correlation matrix between semantic similarity dimensions (lower triangle).}130\label{fig:corr_matrix}131\begin{minipage}{0.9\textwidth}132\footnotesize133Notes: Correlations are computed across all $n = 17{,}087$ observations. The strong positive correlations largely reflect the common influence of description length and verbosity on similarity scores.134\end{minipage}135\end{figure}136137Inter-reference correlations are strong and uniformly positive, ranging from 0.66 to 0.97 with a mean of 0.84. This common variation is driven primarily by description length and verbosity: detailed listings score higher on multiple dimensions simultaneously. The strongest correlations occur between Entry-Level and Motivated Seller ($r = 0.97$)---both characteristic of value-oriented listings---and between Luxury and Renovated ($r = 0.95$), which share upgrade-related vocabulary. This pronounced collinearity motivates two design features of the analysis: the inclusion of description length as a control, which absorbs much of the shared variation, and the multicollinearity diagnostics reported in Section~\ref{sec:robustness}, which quantify its consequences for inference. It also implies that individual semantic coefficients are best read as partial associations within a correlated block of quality signals rather than as isolated effects.138139\subsection{Illustrative Relationships}140141Figure~\ref{fig:scatter} presents scatter plots for four selected dimensions, illustrating the heterogeneity in relationships between semantic similarity and price.142143\begin{figure}[!htbp]144\centering145\includegraphics[width=0.8\textwidth]{fig6_scatter_plots.pdf}146\caption{Bivariate relationships between selected semantic similarities and log-price.}147\label{fig:scatter}148\begin{minipage}{0.9\textwidth}149\footnotesize150Notes: Lines show ordinary least squares linear fit. These bivariate relationships do not control for structural variables or description length.151\end{minipage}152\end{figure}153154\subsection{Residual Diagnostics}155156Figure~\ref{fig:diagnostics} presents residual diagnostics for the full model, including a residual-versus-fitted plot and a normal Q-Q plot.157158\begin{figure}[!htbp]159\centering160\includegraphics[width=0.85\textwidth]{fig11_residual_diagnostics.pdf}161\caption{Residual diagnostics for Model D: (a) residuals vs.\ fitted values; (b) normal Q-Q plot.}162\label{fig:diagnostics}163\begin{minipage}{0.9\textwidth}164\footnotesize165Notes: Panel (a) plots OLS residuals against fitted values, revealing mild heteroskedasticity at the tails. Panel (b) compares the empirical residual distribution to the standard normal, showing heavier-than-normal tails.166\end{minipage}167\end{figure}168169The residual plot (panel a) reveals mild heteroskedasticity, with residual variance increasing at the extremes of the fitted value distribution. This motivates our use of HC3 robust standard errors. The Q-Q plot (panel b) shows approximate normality in the central distribution with heavier-than-normal tails, particularly in the upper tail---consistent with the well-known right-skew of housing prices even after log transformation. The OLS coefficient estimates remain consistent under these conditions, and the HC3 standard errors provide valid inference.170