PAPER_REVIEW — critical assessment before the scholarly upgrade
Paper: Decoding Real Estate Descriptions (UQO WP No. 2), 47 pp., 59 references.
Assessment date: 2026-08-05. Basis: full read of paper/ + the pipeline in
src/–scripts/ and the verified results in results/.
1. Core contribution — is it clearly stated?
Yes, and it is real: projecting sentence embeddings onto researcher-defined reference descriptions ("concept projection") to obtain named, signed, interpretable text features for hedonic inference, with an honest quantification of the interpretability-vs-fit trade-off against PCA. The three-part contribution statement in the introduction is crisp.
But it is under-positioned. The paper frames itself only against the "text in real estate" niche. It never connects to:
- the text-as-data program in economics (Gentzkow–Kelly–Taddy; Ash–Hansen; Hansen–McMahon–Prat), where the method is naturally described as a continuous, embedding-based generalization of dictionary methods — a framing that both strengthens novelty and inherits a mature methodological literature;
- the structural hedonic identification literature (Ekeland–Heckman–Nesheim, Bajari–Benkard, Kuminoff–Smith–Timmins), which is the right anchor for the paper's careful "conditional association, not causal" stance;
- the parallel literature on unstructured data for housing quality other than text — notably computer vision on listing photos (Glaeser et al.; Poursaeed et al.) — which makes the case that text is one member of a family of quality-revealing unstructured signals.
2. Literature review — gaps
Current coverage is good on: hedonic surveys (Sirmans ×2, Malpezzi), functional form (Halvorsen–Palmquist, Cropper), text-in-real-estate (Nowak–Smith, Shen, Hong, Lam, Demers, Goodwin, Huang), sentence embeddings (SBERT, MiniLM, SimCSE, BERT, anisotropy), and scattered econometrics (Koenker, Breusch–Pagan, HC3, Lasso/ENet, O'Brien).
Missing subfields (each currently has zero citations):
| Strand | Why it matters here | Candidate anchors |
|---|---|---|
| Text-as-data in economics | The paper's method is an instance of this program; dictionary methods are its direct ancestor | Gentzkow–Kelly–Taddy (2019 JEL); Ash–Hansen (2023 ARE); Hansen–McMahon–Prat (2018 QJE); Loughran–McDonald (2016 survey) |
| Hedonic origins & identification | "Implicit prices ≠ willingness to pay" needs its canonical sources; origins misattributed to Lancaster–Rosen alone | Court (1939)/Griliches (1971); Ekeland–Heckman–Nesheim (2004); Bajari–Benkard (2005); Kuminoff–Smith–Timmins (2013) |
| Quantile hedonics | §Robustness runs quantile regressions with no housing-quantile citation | Zietz–Zietz–Sirmans (2008); McMillen (2008) |
| Listing-language & agent remarks | The closest competing literature — absent entirely | Pryce–Oates (2008); Haag–Rutherford–Thomson (2000); Levitt–Syverson already cited but under-used |
| Spatial hedonics | §Limitations discusses spatial controls citing only the LeSage–Pace textbook | Dubin (1988); Can (1992); Anselin (1988) |
| Quebec/Canada housing | A Quebec paper with zero Quebec housing references | Des Rosiers et al.; Dubé–Legros (spatial hedonics, Québec data) |
| Computer vision on listings | Sibling "unstructured data" literature | Glaeser–Kincaid–Naik (2018); Poursaeed et al. (2018) |
| Interpretable ML | The paper's core value proposition, currently anchored only by concept bottlenecks | Rudin (2019); Lundberg–Lee (2017) |
| Disclosure theory | Milgrom (1981) cited alone; the standard pairing is with Grossman (1981) | Grossman (1981) |
| ML econometrics practice | One-sided (Athey; Mullainathan–Spiess) | Varian (2014); Dell (2025, deep learning for economists) |
3. Weak argumentation / unsupported claims
- "Realtor.ca + Centris capture the near-universe of listed properties" — no citation; keep but soften or attribute to platform documentation. Flagged.
- "all-MiniLM-L6-v2 … Spearman 0.82 on STS-B … ~5× faster than BERT-base" — from the model card, uncited. Add the sentence-transformers/model-card source or soften. Flagged.
- Conneau et al. "up to 15% monolingual performance loss" — check against the actual paper's numbers; keep only if defensible. Flagged.
- Urgency discount (−8.5%) is discussed with Levitt–Syverson but never quantitatively compared (their agent-owned premium ≈ 3.7%); an explicit magnitude comparison would sharpen the discussion.
- Quantile findings are not compared to the housing-quantile literature (Zietz et al. find larger implicit prices at upper quantiles for many attributes — our Luxury gradient agrees; that agreement is currently unstated).
- The negative Family-Friendly/New Construction coefficients are interpreted as locational sorting — plausible, but the argument would benefit from the sorting/access-to-amenities literature anchor (Kuminoff et al.).
4. Underdeveloped sections
- Introduction: solid but thin on stakes (AVM/mass-appraisal economic importance; text-as-data momentum). No explicit "this paper speaks to three literatures" paragraph. Roadmap present.
- Related work: organized by method type, which is good, but missing the four strands above; the "research gap" subsection can be sharpened into a two-sided gap (economics side and NLP side).
- Methodology: several choices lack citations — semi-log form is cited, but standardization-for-comparability, the dictionary-method lineage of the reference design, and the zero-shot/anchor analogy are not.
- Discussion: interprets results internally but rarely against prior findings (no numeric comparisons to Nowak–Smith's keyword premiums, Levitt–Syverson's magnitudes, Zietz et al.'s quantile patterns). Limitations are strong (post-restructuring) — keep.
- Conclusion: fine; can absorb one paragraph tying back to text-as-data.
5. Positioning: current vs. target
Current: "an NLP paper for real estate economists" — a niche methods paper competing with Lam et al./Shen on prediction.
Target: an applied-econometrics paper in the text-as-data tradition: interpretable feature construction from unstructured data for classical inference, demonstrated on housing. That positioning (i) inherits the Gentzkow–Kelly–Taddy legitimacy, (ii) makes PCA/embedding opacity the natural foil, (iii) frames the reference descriptions as a continuous dictionary, and (iv) generalizes the contribution beyond real estate (as the conclusion already gestures).
6. Upgrade plan (targets)
- References: 59 → ~85–90, all new entries verified via OpenAlex (DOI when available).
- Length: 47 pp → ~54–56 pp (+15–20%): intro +~1 p, related work +~2.5 pp (two new subsections + restructuring), methodology +~0.5 p (grounding citations), discussion +~2 pp (literature-facing interpretation + magnitude comparisons), conclusion +~0.3 p.
- Untouched: all results, tables, figures, robustness numbers, data section statistics.