SPB Git

spb/wp3_uqo Public

UQO Working Paper No. 3 — Hedonic housing price models for the US: parametric, quantile, and machine-learning approaches.

TeX 77.8% Python 22.1%
9.1 KB · 122 lines markdown
Rendered Raw Blame History
1# PAPER_REVIEW — Critical assessment before the scholarly upgrade (2026-08-05)23Scope: `paper/` (52-page compiled working paper). Assessment based on a full read of4all sections, the analysis code (`src/`, `scripts/`), and the stored results.5**No results, data, or figures are questioned here — this is about framing, literature,6and argumentation only.**78---910## 1. Core contribution — is it clearly stated?1112**What the paper claims:** a methodological + empirical comparison of OLS, quantile13regression, and gradient boosting on one large listing dataset, with a systematic14demonstration of *spatial leakage* in ML evaluation (geographic holdout, ablation,15lat/lon experiment).1617**Assessment:** the contribution IS stated (intro, ¶5–6) and is genuinely interesting —18the spatial-leakage triad is the paper's strongest and most novel element. But:1920- The intro *underplays* it: the three "persistent limitations" framing (linearity,21  mean-only, black-box) is generic and could open any of a hundred ML-hedonics papers.22  The leakage finding — that geographic features actively *harm* generalization — is23  the distinctive result and deserves to lead.24- The contribution statement is not benchmarked against the closest existing work25  (nothing tells the reader what the *delta* is vs. Bourassa et al.'s spatial-ML26  comparisons or vs. the spatial-CV literature imported from ecology).27- "First paper to X" claims are absent (good — none would survive), but the paper28  never says precisely *which combination* is new: national scale + listing data +29  three-way framework comparison + spatial-leakage quantification.3031## 2. Literature review — gaps3233Current: 37 references, 7 short subsections. Solid skeleton (Rosen/Lancaster34foundations, functional form, QR, ML/AVM, SHAP, spatial validation), but thin for a35journal submission in housing/urban economics. Specific gaps:3637| Missing strand | Why it matters here | Representative works to add |38|---|---|---|39| **Pre-Rosen hedonic history** | Court/Griliches are cited but the agricultural origin (Waugh 1928) and the environmental-valuation lineage (Ridker & Henning 1967) anchor the method's breadth | Waugh (1928); Ridker & Henning (1967) |40| **Hedonic identification syntheses** | The paper leans on Ekeland et al. (2004) alone; the modern surveys of what hedonic coefficients can and cannot identify are absent | Kuminoff, Smith & Timmins (2013 JEL); Bishop et al. (2020); Bartik (1987); Epple (1987) |41| **Specification robustness / FE granularity** | The region→state→ZIP3 exercise begs for Kuminoff, Parmeter & Pope (2010), which asks exactly "which hedonic models can we trust" and finds spatial FE crucial | Kuminoff, Parmeter & Pope (2010 JEEM) |42| **Spatial hedonics beyond LeSage-Pace** | Only 3 spatial refs; no housing-specific spatial autocorrelation classics | Dubin (1998); Basu & Thibodeau (1998); Anselin (1988); Pace & Gilley (1997) |43| **Quantile methods depth** | Koenker & Bassett + three applications; missing the accessible survey and the housing-distribution literature | Koenker & Hallock (2001); McMillen (2008); Waltl (2016) |44| **Listing-price / search literature** | The dependent variable IS an asking price; only Knight (2002) cited. The strategic-pricing and loss-aversion literature directly supports the paper's central caveat | Genesove & Mayer (2001); Horowitz (1992); Han & Strange (2016); Haurin (1988) |45| **Capitalization of local public goods** | School rating and property-tax coefficients get counterintuitive signs; the boundary-discontinuity literature is the natural reference point | Black (1999); Bayer, Ferreira & McMillan (2007) |46| **ML-in-economics methodology** | Mullainathan & Spiess is alone; the econometrics-meets-ML canon is absent | Varian (2014); Athey & Imbens (2019); Breiman (2001); Friedman (2001) |47| **AVM evaluation practice** | The AVM framing (deployment, generalization) has its own metrics literature | Steurer, Hill & Pfeifer (2021); Pace & Hayunga (2020) |48| **Explainability debate** | SHAP caveats are well written but un-cited beyond Lundberg; the interpretability-vs-explanation debate strengthens them | Rudin (2019); Ribeiro et al. (2016); Molnar et al. (2020) — verify availability |49| **Spatial CV methods** | Roberts et al. + Meyer & Pebesma only; the blocking-methods and the *dissenting* literature are missing | Valavi et al. (2019); Ploton et al. (2020); Wadoux et al. (2021) — the last one *argues against* spatial CV and must be engaged, not ignored |50| **Climate risk pricing** | One reference (Baldauf et al.); this is now a large literature and the paper lists climate data as future work | Bernstein, Gustafson & Lewis (2019); Murfin & Spiegel (2020) |51| **Walkability premium** | Walk/Bike/Transit scores are regressors but no walkability-capitalization citation exists | Pivo & Fisher (2011) |52| **Moran's I primary source** | The statistic is used but Moran (1950) / Cliff & Ord are not cited | Moran (1950) |5354**Suspect existing entries (to verify in Phase 2):**55- `bourassa2019machine` — dated 2019 but *JRER* 32(2), 139–159 is the **2010** volume.56- `meyer2019importance` — dated 2019 but *MEE* 12(9), 1620–1633 is **2021**; exact57  title/venue need confirmation.58- `chen2020housing` — "Expert Systems with Applications, 145:113142" needs59  author/title/article-number confirmation.60- `mak2010quantile` — plausible but verify volume/pages.6162## 3. Weak argumentation / unsupported claims63641. **"Neighborhood-quality features … carry location-specific scale and meaning that65   may not transfer"** (discussion) — plausible mechanism, no citation, no test.66   Should be flagged as conjecture or supported (e.g., Walk Score's metro-relative67   construction).682. **"Luxury properties are more frequently overpriced, distressed properties may be69   strategically underpriced"** (data section) — cited only to Knight (2002), which70   does not establish both claims; Genesove & Mayer / Han & Strange needed.713. **Counterintuitive-signs subsection** — the school-rating and tax-rate explanations72   are multicollinearity narratives without references to the capitalization73   literature (Oates is cited elsewhere but not connected here; Black 1999 missing).744. **"This is below unity, consistent with diminishing marginal returns to space"**75   fine, but a comparison to the elasticity range in published meta-analyses76   (Sirmans et al. report living-area gradients) would ground it.775. **The spatial-leakage argument never engages the counter-position** — Wadoux et al.78   (2021) argue spatial CV can be *pessimistically* biased under uniform sampling.79   Engaging this strengthens, not weakens, the paper's design (10-state holdout is an80   extrapolation task, where blocking is defensible).816. **Abstract/intro report the ablation-variant geo R² (0.425)** while the headline82   model reaches 0.547 — internally explained (different training composition), but a83   reviewer will push; the discussion should own this more explicitly.8485## 4. Underdeveloped sections8687- **Introduction (31 lines):** no broader stakes (housing = ~$45T US asset class,88  AVM industry, algorithmic valuation policy debates); no explicit "contributions"89  enumeration tied to literature strands; roadmap is one flat sentence.90- **Literature (37 lines):** each subsection is 3–6 lines — closer to an annotated91  list than a review. No synthesis paragraphs, no explicit "gap" argument per strand92  (the Research Gap subsection does some of this but in generic terms).93- **Discussion (44 lines):** four "messages" are well structured but interpret results94  almost entirely *internally* — few comparisons with published magnitudes95  (elasticities, QR patterns vs. Zietz et al., ML gains vs. Bourassa et al.,96  geographic-transfer losses vs. ecology findings).97- **Limitations (37 lines):** good coverage, telegraphic style; several items could98  cite the literature that documents the problem (e.g., listing-vs-transaction:99  Genesove & Mayer; spatial CV design choice: Valavi/Wadoux).100- **Conclusion (19 lines):** adequate; can be sharpened with one paragraph on external101  validity and one on the research agenda, without overselling.102103## 5. Positioning: current vs. recommended104105**Current de facto positioning:** "a careful multi-method comparison with unusually106honest caveats" — reads like a very good methods-audit working paper.107108**Recommended positioning:** "evidence on *when and why* ML predictive advantages in109hedonic valuation are real vs. artifacts of validation design, at national scale" —110i.e., lead with the spatial-leakage contribution, use the three-framework comparison111as the vehicle, and connect explicitly to (a) the hedonic-identification literature112(what the coefficients mean), (b) the spatial-CV literature (what the metrics mean),113and (c) the AVM-deployment literature (why practitioners should care).114115## 6. What must NOT change116117All numbers, tables, figures, estimation choices, and robustness results stay exactly118as they are. The upgrade is: framing, motivation, literature integration, discussion119depth, and citation grounding. Any tension found between the paper's findings and the120added literature is to be *reported* (UPGRADE_REPORT.md) and *discussed* in the text,121never resolved by altering results.122