SPB Git

spb/wp2_uqo Public

UQO Working Paper No. 2 — Decoding Real Estate Descriptions: text-based hedonic analysis of housing listings.

TeX 73.8% Python 26%
21.2 KB · 76 lines latex
Raw Blame History
1% Author: Simon-Pierre Boucher — contact@spboucher.ai2%3%======================================================================4\section{Discussion} \label{sec:discussion}5%======================================================================67\subsection{Interpretation of Key Findings}89Our results establish that free-text property descriptions contain economically significant information that traditional hedonic variables fail to capture. The 6-percentage-point improvement in $R^2$ is substantial given that the baseline model already explains 45\% of price variation with standard structural variables---a level consistent with the hedonic pricing literature \citep{sirmans2005composition}. The semantic features contribute 9.3\% of the total explained variance, suggesting that qualitative housing attributes communicated through listing text represent a meaningful dimension of housing differentiation. To contextualize this magnitude, \citet{sirmans2006empirical} found that adding a full set of neighborhood controls to a structural-only model typically improves $R^2$ by 3--8 percentage points---our semantic features achieve comparable gains from a single text field.1011The positive effects of \textbf{Modern/Contemporary} (+16.4\%) and \textbf{Luxury} (+14.2\%) language align with theoretical expectations and prior empirical work on quality premiums \citep{nowak2017quality, goodwin2014broker}. These dimensions capture aspects of housing quality---design aesthetics, material quality, technological amenities---that are simply not reflected in counts of bedrooms and bathrooms. A house with three bathrooms can range from a modest home with basic fixtures to a luxury property with spa-like en-suites; the listing description captures this variation. The quantile regression analysis (Section~\ref{sec:robustness}) reveals that the luxury premium is amplified in the upper portion of the price distribution, consistent with the notion that luxury is a positional good whose marginal value increases with baseline quality \citep{frank2007falling} and with the general finding of the quantile-hedonic literature that many quality attributes command larger implicit prices in upper market segments \citep{zietz2007determinants}.1213The \textbf{Land \& Nature} premium (+13.0\%) reflects the well-documented value of natural amenities and outdoor space in residential markets \citep{irwin2002effects, tyrvainen2005benefits}. This finding is particularly relevant in the Quebec context: using Quebec City transactions and site-inspection measures, \citet{desrosiers2002landscaping} estimate landscaping premiums ranging from roughly 4\% (hedges, landscaped curbs) to about 12\% (a landscaped patio) of house value. Our text-based measure recovers an amenity gradient of comparable direction and magnitude through the listing narrative alone, without site-level amenity data---evidence that the two measurement strategies tap the same underlying valuation.1415The negative coefficient on \textbf{Family-Friendly} ($-$11.6\%) deserves careful interpretation. This does not imply that schools and parks reduce property values; rather, it reflects the systematic association between family-oriented marketing language and lower-priced suburban markets. After controlling for structural characteristics, family-friendly language serves as a residual proxy for suburban location where prices are lower. This is precisely the equilibrium-sorting logic emphasized by \citet{kuminoff2013new}: households sort across locations on price and amenities jointly, so language that marks a market segment inherits that segment's price level. The finding also echoes \citet{delmelle2021language}, who show that property-advertisement text is a strong predictor of neighborhood socioeconomic composition: listing language is as much about \textit{where} a property is as about \textit{what} it is. This illustrates the importance of interpreting hedonic coefficients as conditional marginal effects rather than causal estimates---a point emphasized by \citet{pakes2003reconsideration} in the broader context of hedonic identification.1617The strong negative effect of \textbf{Motivated Seller} ($-$8.5\%) is consistent with the information asymmetry literature. \citet{levitt2008information} showed that real estate agents selling their own homes achieve prices about 3.7\% higher than when selling clients' homes, partly because they can conceal information about the urgency of the sale; our estimate implies that \textit{explicit} urgency disclosure is associated with a discount more than twice that size, which is coherent with the idea that overt urgency language is a stronger signal of a weak outside option than the subtle cues agents suppress. This result connects to the broader theory of strategic disclosure \citep{grossman1981informational, milgrom1981good}: in unraveling equilibria, transparency about motivation invites adverse inference, as buyers rationally read disclosed urgency as evidence of weak bargaining position. The listing-language literature reports the same qualitative pattern: remarks signaling seller motivation are associated with lower prices and faster sales \citep{haag2000real, goodwin2014broker}.1819The surprising negative coefficient on \textbf{New Construction} ($-$10.5\%) warrants extended discussion. In most housing markets, ``new'' is a premium attribute. However, in Quebec's real estate geography, new residential construction is concentrated in peripheral suburban developments (e.g., Mirabel, Mascouche, Lachute) where land costs are lower. After controlling for structural characteristics, the ``new construction'' semantic dimension captures this locational sorting rather than a quality discount. This interpretation is supported by the quantile regression results, which show a diminished negative effect at higher quantiles---precisely the pattern expected if the coefficient proxies for peripheral location rather than an intrinsic quality penalty.2021\subsection{Relation to Prior Findings}2223Situating the estimates against the closest prior work sharpens both agreements and departures.2425\textit{Agreement on the information content of text.} \citet{nowak2017quality} find that words in agent comments proxy for unobserved quality and that omitting them biases other hedonic coefficients; we corroborate this at the level of model fit (a 6-point $R^2$ gain) and extend it by giving the text features named economic content. \citet{shen2020text} conclude that property descriptions contain substantial information about otherwise unobserved quality; our results agree, and our decomposition (Figure~\ref{fig:r2_decomp}) quantifies the same conclusion in variance terms. The early remark-category studies \citep{haag2000real} and the vernacular studies \citep{goodwin2014broker, goodwin2018connotation} found sign patterns---premia for quality language, discounts for distress language---that match our forest plot dimension by dimension.2627\textit{Agreement on distributional heterogeneity.} The amplification of the Luxury premium at upper quantiles and the attenuation of condition/urgency discounts mirror the findings of \citet{zietz2007determinants}, who document that most attribute prices rise with the conditional price level, and of \citet{mcmillen2008changes}, who shows that coefficient heterogeneity---not just characteristics---drives changes in the price distribution.2829\textit{A departure from the prediction-first literature.} Studies that optimize prediction \citep{shen2020text, baur2023automated} report larger fit gains from unrestricted text representations than our 20 projections deliver; our own PCA benchmark reproduces this ordering (Table~\ref{tab:text_comparison}). The departure is deliberate: we accept a quantified fit penalty in exchange for coefficient-level interpretability, a trade that the prediction literature does not price.3031\textit{A parallel with vision-based studies.} \citet{glaeser2018computer} and \citet{poursaeed2018vision} extract quality from listing photographs and find, as we do with text, that unstructured listing content reveals price-relevant quality invisible to structured fields. That two different modalities produced by the same marketing process yield the same conclusion strengthens the case that structured hedonic variables systematically undermeasure quality---and suggests a natural extension combining both.3233\subsection{The Information Content of Agent Narratives}3435Our findings contribute to a broader literature on information production by market intermediaries. Real estate agents serve a dual role: they match buyers with properties and produce information about property characteristics through listing descriptions. The rhetoric of listings is deliberately crafted \citep{pryce2008rhetoric}, yet the economically significant coefficients on our semantic dimensions suggest that agents' textual descriptions contain genuine informational content---they are not merely ``cheap talk'' or undifferentiated marketing prose.3637This interpretation is supported by two pieces of evidence. First, the signs of most semantic coefficients align with theoretical priors: luxury language commands premiums, renovation-need language is discounted, and urgency language penalizes prices. If listing descriptions were pure noise or uniform boilerplate, the semantic features would exhibit no systematic relationship with prices. Second, the stability of the coefficients across quantile regression, trimmed samples, and bootstrap inference (Section~\ref{sec:robustness}) indicates that the text-price associations are robust features of the data rather than artifacts of particular observations or distributional assumptions.3839At the same time, the correlational nature of our estimates precludes definitive conclusions about the mechanism. The positive coefficient on Luxury language, for instance, could reflect: (a) agents accurately describing genuinely luxurious properties that command higher prices; (b) the language itself influencing buyer perceptions and willingness to pay; or (c) an omitted variable (e.g., neighborhood prestige) correlated with both luxury language and price. Disentangling these mechanisms would require quasi-experimental variation in listing text---an approach feasible with A/B testing data from real estate platforms but beyond the scope of the present study.4041\subsection{Description Length as a Quality Signal}4243The significant positive effect of description length (+10.7\%) merits discussion. This finding may seem paradoxical given that bivariate correlations between similarity scores and price are uniformly negative (reflecting the tendency for lower-priced listings to have longer descriptions). The resolution lies in the multivariate structure: once we control for \textit{what} a description says (via the semantic similarities), \textit{how much} it says becomes a positive signal. Longer descriptions, conditional on content, may indicate agent effort, property complexity, or a genuine abundance of features to describe. This is consistent with \citet{shen2020text}, who find that the informational content of property descriptions proxies for unobserved quality, and with the listing-effort interpretation in \citet{goodwin2014broker}. One caveat is that descriptions in our data are truncated at approximately 700 characters (Section~\ref{sec:data}), so the length variable measures verbosity only up to this bound; the estimated coefficient should be read as the effect of length variation within the observed range.4445\subsection{Comparison with PCA-Based Approaches}4647To contextualize our results, we compare the reference-based approach with the more standard PCA-based method on the same sample (Table~\ref{tab:text_comparison}). Applying PCA to the raw 384-dimensional embedding vectors and retaining 20 principal components---which together capture 49.5\% of the embedding variance---yields a larger fit improvement than the 20 reference similarities: adding the components to the structural baseline raises $R^2$ by 7.0 percentage points, versus 4.4 points for the reference projections. This ordering is unsurprising: the principal components are the linear combinations of the embedding space that maximize retained variance, so any fixed set of 20 projections, including ours, is weakly dominated in pure fit.4849What PCA cannot provide is economic content. A coefficient on ``PC3'' has no natural interpretation as a housing characteristic, cannot be communicated to appraisers or market participants, and is unstable across samples in its meaning even when stable in its fit. The reference-based approach deliberately trades roughly two percentage points of $R^2$ for features that are individually nameable, signable ex ante, and directly usable in the hedonic framework, where inference on implicit prices---not prediction---is the goal. This is the position argued on general grounds by \citet{rudin2019stop}: when a model informs high-stakes decisions, inherent interpretability is preferable to post hoc explanation of a black box \citep{ribeiro2016why, lundberg2017unified}. Researchers whose objective is purely predictive should use the unrestricted embedding space (or the full 384 dimensions with regularization); researchers who need interpretable implicit prices face the trade-off we quantify here.5051\subsection{Practical Implications}5253Our findings have several practical implications. For \textbf{automated valuation models (AVMs)}, incorporating semantic text features could meaningfully improve accuracy. The 6-percentage-point $R^2$ improvement suggests that AVMs relying solely on structured data---still the dominant design in mass appraisal \citep{kok2017big}---leave significant predictive power on the table. Moreover, our reference-based approach produces features that are computationally cheap to generate (requiring only a dot product between pre-computed embeddings) and stable across time.5455For \textbf{real estate appraisers and agents}, the results highlight which qualitative dimensions of listing language are most strongly associated with price variation. Agents crafting listing descriptions can leverage these findings to understand how different linguistic strategies relate to market positioning. However, we caution that the relationship is correlational: changing the words in a listing description is unlikely to change the sale price if the underlying property attributes remain unchanged.5657For \textbf{housing market researchers}, the reference-based framework offers a flexible tool for investigating qualitative dimensions of housing that have been difficult to operationalize. Researchers can define custom reference descriptions tailored to their specific research questions, making the approach portable to different markets, languages, and property types.5859\subsection{Limitations}6061Several limitations should be acknowledged, and we discuss each in turn along with potential avenues for resolution.6263\textbf{Listing prices vs.\ transaction prices.} Our data consists of listing prices, not realized sale prices. List and sale prices differ systematically, and the gap varies with market conditions and marketing time \citep{haurin2010list}; the magnitude of this gap may correlate with our semantic measures (e.g., ``motivated seller'' listings may sell at a larger discount to listing price). If the listing-to-sale price ratio varies systematically with semantic content, our coefficient estimates may be biased. The direction of this bias is ambiguous: luxury language may be associated with a smaller listing premium (if agents of luxury properties price more accurately) or a larger one (if they price aspirationally). The availability of transaction-level data would allow estimation of the listing premium as a function of semantic content and would strengthen the analysis.6465\textbf{Cross-sectional identification.} The cross-sectional nature of our data prevents causal inference. As discussed in Section~\ref{sec:identification}, the relationship between listing language and price may reflect accurate quality description, strategic persuasion, or omitted variable correlation. While our robustness checks (Section~\ref{sec:robustness}) confirm the stability of the estimates across specifications, they do not address the fundamental identification concern. Future work using within-property variation---comparing successive listings for the same property with different agents or different textual strategies---could provide more credible causal identification, and recent work on textual strategies in housing \citep{alfano2022word} points toward research designs that exploit variation in listing composition; platform-level A/B experiments would be the gold standard.6667\textbf{Absence of spatial controls.} Our hedonic model does not include explicit locational variables (municipality fixed effects, distance to CBD, or spatial autoregressive terms), despite the well-established spatial dependence of housing prices \citep{anselin1988spatial, dubin1988estimation, can1992specification}. Some semantic dimensions---particularly Family-Friendly, New Construction, and Quiet \& Peaceful---likely capture locational variation as much as property-level attributes. Including municipality or census-tract fixed effects would absorb location-specific price levels, allowing the semantic coefficients to be interpreted more cleanly as within-location quality signals. However, this comes at the cost of eliminating between-location variation that the semantic features may legitimately capture. A spatial Durbin model \citep{lesage2009introduction} would provide a formal framework for decomposing direct and indirect (spillover) effects of semantic content, and the microdata methods of \citet{dube2014spatial} are directly applicable to our setting.6869\textbf{Language model limitations.} The all-MiniLM-L6-v2 model was primarily trained on English text. While it handles French adequately---the language of most Quebec listings---a model specifically fine-tuned on French real estate text could potentially produce more nuanced embeddings. The model may fail to capture domain-specific nuances: for instance, ``plancher chauffant'' (heated floor) and ``radiant heating'' carry identical meaning but may not embed identically across languages. Cross-lingual embedding models optimized for French, such as CamemBERT \citep{martin2020camembert} or FlauBERT \citep{le2020flaubert}, represent promising alternatives for future work. \citet{conneau2020unsupervised} document material performance losses when multilingual models are applied to specialized monolingual tasks, suggesting that our estimates may understate the true informational content of listing descriptions.7071\textbf{Reference description subjectivity.} Our 20 reference descriptions represent one researcher's operationalization of qualitative housing dimensions. Alternative formulations could yield different similarity scores and potentially different hedonic estimates. While we designed the references following principled criteria (semantic saturation, dimensional specificity, linguistic consistency), the approach lacks a formal optimality criterion. A systematic sensitivity protocol---averaging similarities over multiple paraphrases per dimension, varying reference length, and comparing French-only against bilingual formulations---would quantify the dependence of the estimates on any single phrasing. One could also envision deriving reference descriptions empirically, for example by clustering listing embeddings and using cluster centroids as data-driven references, though this sacrifices the ex ante interpretability that motivated our approach. The full text of all 20 references is reproduced in Appendix Table~\ref{tab:reference_texts} to make this dependence transparent and replicable.7273\textbf{Multicollinearity among similarity features.} The strong correlations among similarity dimensions (0.66--0.97; Figure~\ref{fig:corr_matrix}) introduce substantial multicollinearity: VIFs average 17.2 and exceed the conventional threshold of 10 for 16 of the 20 dimensions (Section~\ref{sec:robustness}). The consequences are visible in the regularized-selection exercise, where the Lasso retains only four block representatives rather than the sixteen individually significant variables. Individual coefficient magnitudes should therefore be interpreted with caution: the ``Luxury'' premium of 14.2\% includes variation that is shared with the Renovated and Bright \& Spacious dimensions, and the data cannot fully apportion credit within such correlated blocks. The block-level conclusions---that semantic content jointly adds 6 percentage points of explanatory power with theoretically sensible signs---are unaffected. Ridge regression or principal component regression on the similarity features would provide bounds on the total effect of correlated quality dimensions, and reducing the number of references (or orthogonalizing them ex ante) is a natural design lever for future applications.7475\textbf{Temporal stability.} Our analysis captures a single cross-section of listings. The relationship between textual content and prices may vary over time as market conditions change, buyer preferences evolve, and listing language conventions shift. During a seller's market, ``motivated seller'' language may carry a larger discount because it is rarer and more diagnostic; during a buyer's market, such language may be more common and less informative. Longitudinal analysis would be needed to assess the stability of semantic implicit prices across market cycles.76