SPB Git

spb/wp2_uqo Public

UQO Working Paper No. 2 — Decoding Real Estate Descriptions: text-based hedonic analysis of housing listings.

TeX 73.8% Python 26%
5.9 KB · 23 lines latex
Raw Blame History
1% Author: Simon-Pierre Boucher — contact@spboucher.ai2%3%======================================================================4\section{Conclusion} \label{sec:conclusion}5%======================================================================67This paper demonstrates that property listing descriptions contain economically meaningful information that can be systematically extracted and integrated into hedonic pricing models. Using sentence embeddings and cosine similarity with 20 interpretable reference descriptions, we show that semantic features from listing text improve the explanatory power of a standard hedonic model by 5.9 percentage points (adjusted $R^2$: 0.452 $\rightarrow$ 0.511) for 17,087 single-family homes in Quebec, Canada. The estimates are stable under bootstrap inference, outlier trimming, and quantile-based heterogeneity analysis, while the accompanying diagnostics candidly document the strong collinearity among semantic dimensions and its implications for variable-by-variable interpretation.89Our approach bridges the hedonic pricing tradition and the text-as-data program in empirical economics \citep{gentzkow2019text, ash2023text} by producing features that satisfy the requirements of both: it generalizes the transparent dictionary method into embedding space while preserving the ex ante interpretability that inference on implicit prices requires \citep{rudin2019stop}. They are semantically rich, capturing deep meaning rather than surface-level keyword overlap, and economically interpretable, with each feature corresponding to a named qualitative dimension carrying a clear coefficient estimate. Descriptions emphasizing luxury finishes (+14.2\%), modern design (+16.4\%), and natural settings (+13.0\%) command significant premiums, while language signaling renovation needs ($-$7.9\%), seller urgency ($-$8.5\%), and family-oriented marketing ($-$11.6\%) is associated with discounts. Importantly, the quantile regression analysis reveals that these effects are heterogeneous across the price distribution: the luxury premium is amplified for high-value properties, while renovation and urgency discounts are attenuated, consistent with quality complementarities and differential buyer price sensitivity.1011The general methodology---embedding domain-specific text, computing cosine similarities against researcher-designed references, and incorporating the resulting features into standard econometric models---is portable to any setting where free-text descriptions accompany structured economic data. Applications beyond real estate include:12\begin{itemize}[noitemsep, topsep=3pt]13\item \textbf{E-commerce}: Product listing text on platforms such as Amazon or eBay contains quality and condition signals that could be incorporated into hedonic price models for used goods, electronics, or collectibles.14\item \textbf{Labor economics}: Job posting descriptions embed information about workplace culture, benefits, and flexibility that are imperfectly captured by structured fields; reference-based similarity could quantify the implicit wage premiums associated with different workplace attributes \citep{marinescu2020opening}.15\item \textbf{Innovation economics}: Patent abstracts describe inventions in varying degrees of technical novelty and commercial potential; semantic features could augment patent valuation models.16\item \textbf{Corporate finance}: Earnings call transcripts and annual report narratives contain forward-looking information about firm prospects; reference-based sentiment dimensions could improve asset pricing models beyond the binary positive/negative sentiment used by \citet{loughran2011liability}.17\item \textbf{Hospitality}: Hotel and Airbnb listing descriptions convey amenity and ambiance information that complements structured data on room type, location, and ratings.18\end{itemize}1920In each case, the reference-based approach can translate qualitative textual information into quantitative features suitable for econometric analysis while preserving the interpretability required for inference on implicit prices.2122Future research should extend this framework along several dimensions. First, applying the method to transaction prices rather than listing prices would provide cleaner identification of text-price relationships and enable estimation of the listing premium as a function of semantic content. Second, temporal analysis using panel data could reveal how the relationship between listing language and prices evolves over market cycles---for instance, whether luxury premiums are amplified in bull markets and attenuated in recessions. Third, spatial econometric extensions incorporating municipality fixed effects and spatial autoregressive structures \citep{lesage2009introduction} could disentangle location-based from property-level semantic effects, addressing the concern that some dimensions (Family-Friendly, New Construction) proxy for suburban location. Fourth, multilingual embedding models specifically optimized for French---such as CamemBERT \citep{martin2020camembert} or domain-adapted variants---together with TF-IDF, keyword, and sentiment baselines, would complete the comparison of text representations begun in Section~\ref{sec:robustness}. Fifth, a systematic sensitivity analysis of the reference descriptions---multi-paraphrase averaging, bilingual variants, and length perturbations---would quantify the robustness of the semantic features to their exact phrasing. Sixth, the framework could be extended to buyer-side text (search queries, saved-listing patterns), connecting to the evidence that housing search is itself segmented along the qualitative dimensions buyers care about \citep{piazzesi2020segmented}, and to listing photographs, whose quality content parallels that of text \citep{glaeser2018computer, poursaeed2018vision}. Finally, causal identification could be strengthened through within-property variation---comparing successive listings for the same property by different agents---or through A/B testing partnerships with real estate platforms.23