Scholarly upgrade of the paper (v1.1): literature expansion + rewrite
- Bibliography: 37 -> 69 references, every new entry verified via OpenAlex (exact metadata + DOI). Removed one fabricated reference (chen2020housing) and corrected two mis-dated entries (Bourassa et al. 2010; Meyer & Pebesma 2021). - Rewrote and expanded Introduction, Literature Review (10 thematic subsections), Discussion, and Conclusion; citation-grounded edits in Data, Methodology, Results, and Limitations. - No results, tables, or figures changed. Compiles clean: 61 pages, 0 undefined references, 69/69 entries cited. - Added PAPER_REVIEW.md (critical assessment) and UPGRADE_REPORT.md (per-reference justification, flagged claims, contradicting literature). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Showing 13 changed files with +1,217 and −100
added
PAPER_REVIEW.md
+121 −0
@@ -0,0 +1,121 @@ | ||
| 1 | +# PAPER_REVIEW — Critical assessment before the scholarly upgrade (2026-08-05) | |
| 2 | + | |
| 3 | +Scope: `paper/` (52-page compiled working paper). Assessment based on a full read of | |
| 4 | +all sections, the analysis code (`src/`, `scripts/`), and the stored results. | |
| 5 | +**No results, data, or figures are questioned here — this is about framing, literature, | |
| 6 | +and argumentation only.** | |
| 7 | + | |
| 8 | +--- | |
| 9 | + | |
| 10 | +## 1. Core contribution — is it clearly stated? | |
| 11 | + | |
| 12 | +**What the paper claims:** a methodological + empirical comparison of OLS, quantile | |
| 13 | +regression, and gradient boosting on one large listing dataset, with a systematic | |
| 14 | +demonstration of *spatial leakage* in ML evaluation (geographic holdout, ablation, | |
| 15 | +lat/lon experiment). | |
| 16 | + | |
| 17 | +**Assessment:** the contribution IS stated (intro, ¶5–6) and is genuinely interesting — | |
| 18 | +the spatial-leakage triad is the paper's strongest and most novel element. But: | |
| 19 | + | |
| 20 | +- The intro *underplays* it: the three "persistent limitations" framing (linearity, | |
| 21 | + mean-only, black-box) is generic and could open any of a hundred ML-hedonics papers. | |
| 22 | + The leakage finding — that geographic features actively *harm* generalization — is | |
| 23 | + the distinctive result and deserves to lead. | |
| 24 | +- The contribution statement is not benchmarked against the closest existing work | |
| 25 | + (nothing tells the reader what the *delta* is vs. Bourassa et al.'s spatial-ML | |
| 26 | + comparisons or vs. the spatial-CV literature imported from ecology). | |
| 27 | +- "First paper to X" claims are absent (good — none would survive), but the paper | |
| 28 | + never says precisely *which combination* is new: national scale + listing data + | |
| 29 | + three-way framework comparison + spatial-leakage quantification. | |
| 30 | + | |
| 31 | +## 2. Literature review — gaps | |
| 32 | + | |
| 33 | +Current: 37 references, 7 short subsections. Solid skeleton (Rosen/Lancaster | |
| 34 | +foundations, functional form, QR, ML/AVM, SHAP, spatial validation), but thin for a | |
| 35 | +journal submission in housing/urban economics. Specific gaps: | |
| 36 | + | |
| 37 | +| Missing strand | Why it matters here | Representative works to add | | |
| 38 | +|---|---|---| | |
| 39 | +| **Pre-Rosen hedonic history** | Court/Griliches are cited but the agricultural origin (Waugh 1928) and the environmental-valuation lineage (Ridker & Henning 1967) anchor the method's breadth | Waugh (1928); Ridker & Henning (1967) | | |
| 40 | +| **Hedonic identification syntheses** | The paper leans on Ekeland et al. (2004) alone; the modern surveys of what hedonic coefficients can and cannot identify are absent | Kuminoff, Smith & Timmins (2013 JEL); Bishop et al. (2020); Bartik (1987); Epple (1987) | | |
| 41 | +| **Specification robustness / FE granularity** | The region→state→ZIP3 exercise begs for Kuminoff, Parmeter & Pope (2010), which asks exactly "which hedonic models can we trust" and finds spatial FE crucial | Kuminoff, Parmeter & Pope (2010 JEEM) | | |
| 42 | +| **Spatial hedonics beyond LeSage-Pace** | Only 3 spatial refs; no housing-specific spatial autocorrelation classics | Dubin (1998); Basu & Thibodeau (1998); Anselin (1988); Pace & Gilley (1997) | | |
| 43 | +| **Quantile methods depth** | Koenker & Bassett + three applications; missing the accessible survey and the housing-distribution literature | Koenker & Hallock (2001); McMillen (2008); Waltl (2016) | | |
| 44 | +| **Listing-price / search literature** | The dependent variable IS an asking price; only Knight (2002) cited. The strategic-pricing and loss-aversion literature directly supports the paper's central caveat | Genesove & Mayer (2001); Horowitz (1992); Han & Strange (2016); Haurin (1988) | | |
| 45 | +| **Capitalization of local public goods** | School rating and property-tax coefficients get counterintuitive signs; the boundary-discontinuity literature is the natural reference point | Black (1999); Bayer, Ferreira & McMillan (2007) | | |
| 46 | +| **ML-in-economics methodology** | Mullainathan & Spiess is alone; the econometrics-meets-ML canon is absent | Varian (2014); Athey & Imbens (2019); Breiman (2001); Friedman (2001) | | |
| 47 | +| **AVM evaluation practice** | The AVM framing (deployment, generalization) has its own metrics literature | Steurer, Hill & Pfeifer (2021); Pace & Hayunga (2020) | | |
| 48 | +| **Explainability debate** | SHAP caveats are well written but un-cited beyond Lundberg; the interpretability-vs-explanation debate strengthens them | Rudin (2019); Ribeiro et al. (2016); Molnar et al. (2020) — verify availability | | |
| 49 | +| **Spatial CV methods** | Roberts et al. + Meyer & Pebesma only; the blocking-methods and the *dissenting* literature are missing | Valavi et al. (2019); Ploton et al. (2020); Wadoux et al. (2021) — the last one *argues against* spatial CV and must be engaged, not ignored | | |
| 50 | +| **Climate risk pricing** | One reference (Baldauf et al.); this is now a large literature and the paper lists climate data as future work | Bernstein, Gustafson & Lewis (2019); Murfin & Spiegel (2020) | | |
| 51 | +| **Walkability premium** | Walk/Bike/Transit scores are regressors but no walkability-capitalization citation exists | Pivo & Fisher (2011) | | |
| 52 | +| **Moran's I primary source** | The statistic is used but Moran (1950) / Cliff & Ord are not cited | Moran (1950) | | |
| 53 | + | |
| 54 | +**Suspect existing entries (to verify in Phase 2):** | |
| 55 | +- `bourassa2019machine` — dated 2019 but *JRER* 32(2), 139–159 is the **2010** volume. | |
| 56 | +- `meyer2019importance` — dated 2019 but *MEE* 12(9), 1620–1633 is **2021**; exact | |
| 57 | + title/venue need confirmation. | |
| 58 | +- `chen2020housing` — "Expert Systems with Applications, 145:113142" needs | |
| 59 | + author/title/article-number confirmation. | |
| 60 | +- `mak2010quantile` — plausible but verify volume/pages. | |
| 61 | + | |
| 62 | +## 3. Weak argumentation / unsupported claims | |
| 63 | + | |
| 64 | +1. **"Neighborhood-quality features … carry location-specific scale and meaning that | |
| 65 | + may not transfer"** (discussion) — plausible mechanism, no citation, no test. | |
| 66 | + Should be flagged as conjecture or supported (e.g., Walk Score's metro-relative | |
| 67 | + construction). | |
| 68 | +2. **"Luxury properties are more frequently overpriced, distressed properties may be | |
| 69 | + strategically underpriced"** (data section) — cited only to Knight (2002), which | |
| 70 | + does not establish both claims; Genesove & Mayer / Han & Strange needed. | |
| 71 | +3. **Counterintuitive-signs subsection** — the school-rating and tax-rate explanations | |
| 72 | + are multicollinearity narratives without references to the capitalization | |
| 73 | + literature (Oates is cited elsewhere but not connected here; Black 1999 missing). | |
| 74 | +4. **"This is below unity, consistent with diminishing marginal returns to space"** — | |
| 75 | + fine, but a comparison to the elasticity range in published meta-analyses | |
| 76 | + (Sirmans et al. report living-area gradients) would ground it. | |
| 77 | +5. **The spatial-leakage argument never engages the counter-position** — Wadoux et al. | |
| 78 | + (2021) argue spatial CV can be *pessimistically* biased under uniform sampling. | |
| 79 | + Engaging this strengthens, not weakens, the paper's design (10-state holdout is an | |
| 80 | + extrapolation task, where blocking is defensible). | |
| 81 | +6. **Abstract/intro report the ablation-variant geo R² (0.425)** while the headline | |
| 82 | + model reaches 0.547 — internally explained (different training composition), but a | |
| 83 | + reviewer will push; the discussion should own this more explicitly. | |
| 84 | + | |
| 85 | +## 4. Underdeveloped sections | |
| 86 | + | |
| 87 | +- **Introduction (31 lines):** no broader stakes (housing = ~$45T US asset class, | |
| 88 | + AVM industry, algorithmic valuation policy debates); no explicit "contributions" | |
| 89 | + enumeration tied to literature strands; roadmap is one flat sentence. | |
| 90 | +- **Literature (37 lines):** each subsection is 3–6 lines — closer to an annotated | |
| 91 | + list than a review. No synthesis paragraphs, no explicit "gap" argument per strand | |
| 92 | + (the Research Gap subsection does some of this but in generic terms). | |
| 93 | +- **Discussion (44 lines):** four "messages" are well structured but interpret results | |
| 94 | + almost entirely *internally* — few comparisons with published magnitudes | |
| 95 | + (elasticities, QR patterns vs. Zietz et al., ML gains vs. Bourassa et al., | |
| 96 | + geographic-transfer losses vs. ecology findings). | |
| 97 | +- **Limitations (37 lines):** good coverage, telegraphic style; several items could | |
| 98 | + cite the literature that documents the problem (e.g., listing-vs-transaction: | |
| 99 | + Genesove & Mayer; spatial CV design choice: Valavi/Wadoux). | |
| 100 | +- **Conclusion (19 lines):** adequate; can be sharpened with one paragraph on external | |
| 101 | + validity and one on the research agenda, without overselling. | |
| 102 | + | |
| 103 | +## 5. Positioning: current vs. recommended | |
| 104 | + | |
| 105 | +**Current de facto positioning:** "a careful multi-method comparison with unusually | |
| 106 | +honest caveats" — reads like a very good methods-audit working paper. | |
| 107 | + | |
| 108 | +**Recommended positioning:** "evidence on *when and why* ML predictive advantages in | |
| 109 | +hedonic valuation are real vs. artifacts of validation design, at national scale" — | |
| 110 | +i.e., lead with the spatial-leakage contribution, use the three-framework comparison | |
| 111 | +as the vehicle, and connect explicitly to (a) the hedonic-identification literature | |
| 112 | +(what the coefficients mean), (b) the spatial-CV literature (what the metrics mean), | |
| 113 | +and (c) the AVM-deployment literature (why practitioners should care). | |
| 114 | + | |
| 115 | +## 6. What must NOT change | |
| 116 | + | |
| 117 | +All numbers, tables, figures, estimation choices, and robustness results stay exactly | |
| 118 | +as they are. The upgrade is: framing, motivation, literature integration, discussion | |
| 119 | +depth, and citation grounding. Any tension found between the paper's findings and the | |
| 120 | +added literature is to be *reported* (UPGRADE_REPORT.md) and *discussed* in the text, | |
| 121 | +never resolved by altering results. | |
added
UPGRADE_REPORT.md
+161 −0
@@ -0,0 +1,161 @@ | ||
| 1 | +# UPGRADE_REPORT — Scholarly upgrade of the paper (2026-08-05) | |
| 2 | + | |
| 3 | +Scope executed: Phase 1 (`PAPER_REVIEW.md`), Phase 2 (literature expansion, all | |
| 4 | +references verified via **OpenAlex**), Phase 3 (section rewrites), Phase 4 (this | |
| 5 | +report). **No results, data, figures, or tables were changed.** Paper version | |
| 6 | +bumped 1.0 → 1.1. Compiled: `paper/main.pdf`, **61 pages** (was 52), 0 errors, | |
| 7 | +0 undefined references, **69/69 bibliography entries cited**. | |
| 8 | + | |
| 9 | +Parameters chosen for the bracketed placeholders: field = housing / real estate | |
| 10 | +economics (journal-article standard, e.g., *Journal of Housing Economics* / *Real | |
| 11 | +Estate Economics*); reference target ≈ 60–65 (final: **69**, from 37); length | |
| 12 | +target ≈ +25–35 % prose (achieved: intro ×2.2, literature ×4, discussion ×2.7, | |
| 13 | +conclusion ×2.5 in source lines; total document +9 pages ≈ +17 % because tables | |
| 14 | +and figures are unchanged). | |
| 15 | + | |
| 16 | +--- | |
| 17 | + | |
| 18 | +## 1. New references added (32 net: 33 added, 1 removed) — with justification | |
| 19 | + | |
| 20 | +All entries verified one-by-one against OpenAlex (exact title, authors, venue, | |
| 21 | +volume/pages, DOI). None were added from memory. | |
| 22 | + | |
| 23 | +### Hedonic foundations & identification (7) | |
| 24 | +| Key | Reference | Why added | | |
| 25 | +|---|---|---| | |
| 26 | +| `waugh1928quality` | Waugh (1928), *J. Farm Economics* | Historical origin of hedonic regression | | |
| 27 | +| `ridker1967determinants` | Ridker & Henning (1967), *REStat* | First property-value hedonic; anchors capitalization lineage | | |
| 28 | +| `bartik1987estimation` | Bartik (1987), *JPE* | Second-stage identification problem | | |
| 29 | +| `kuminoff2010which` | Kuminoff, Parmeter & Pope (2010), *JEEM* | Directly supports the FE-granularity exercise (region→state→ZIP3) | | |
| 30 | +| `kuminoff2013new` | Kuminoff, Smith & Timmins (2013), *JEL* | Modern survey of what hedonic estimates identify | | |
| 31 | +| `bishop2020best` | Bishop et al. (2020), *REEP* | Best-practice benchmark the paper now aligns itself with | | |
| 32 | +| `hill2013hedonic` | (existing, re-cited) | Was dropped by rewrite; restored in functional-form review | | |
| 33 | + | |
| 34 | +### Spatial econometrics (4) | |
| 35 | +| Key | Reference | Why added | | |
| 36 | +|---|---|---| | |
| 37 | +| `anselin1988spatial` | Anselin (1988), Kluwer book | Canonical spatial-econometrics reference | | |
| 38 | +| `dubin1998spatial` | Dubin (1998), *J. Housing Econ.* | Why house prices are spatially autocorrelated | | |
| 39 | +| `basu1998analysis` | Basu & Thibodeau (1998), *JREFE* | Housing-specific residual autocorrelation benchmark for Moran's I | | |
| 40 | +| `moran1950notes` | Moran (1950), *Biometrika* | Primary source for the statistic used | | |
| 41 | + | |
| 42 | +### Quantile regression (3) | |
| 43 | +| Key | Reference | Why added | | |
| 44 | +|---|---|---| | |
| 45 | +| `koenker2001quantile` | Koenker & Hallock (2001), *JEP* | Accessible methodological grounding | | |
| 46 | +| `mcmillen2008changes` | McMillen (2008), *JUE* | Coefficients-vs-characteristics evidence supporting QR relevance | | |
| 47 | +| `waltl2019variation` | Waltl (2019), *Real Estate Economics* | Recent comprehensive housing QR (Sydney), closest antecedent | | |
| 48 | + | |
| 49 | +### Listing prices & seller behavior (3) | |
| 50 | +| Key | Reference | Why added | | |
| 51 | +|---|---|---| | |
| 52 | +| `horowitz1992role` | Horowitz (1992), *J. Applied Econometrics* | Theory of list price as commitment device | | |
| 53 | +| `genesove2001loss` | Genesove & Mayer (2001), *QJE* | Loss aversion → systematic asking-price behavior; grounds the central caveat | | |
| 54 | +| `han2016role` | Han & Strange (2016), *JUE* | Directing role of asking price in buyer search | | |
| 55 | + | |
| 56 | +### Capitalization of local public goods (3) | |
| 57 | +| Key | Reference | Why added | | |
| 58 | +|---|---|---| | |
| 59 | +| `black1999better` | Black (1999), *QJE* | Boundary-discontinuity school valuation; disciplines the negative school-rating sign | | |
| 60 | +| `bayer2007unified` | Bayer, Ferreira & McMillan (2007), *JPE* | Sorting framework; same purpose | | |
| 61 | +| `pivo2011walkability` | Pivo & Fisher (2011), *Real Estate Economics* | Walkability premium; grounds Walk Score discussion | | |
| 62 | + | |
| 63 | +### ML in economics & valuation (6) | |
| 64 | +| Key | Reference | Why added | | |
| 65 | +|---|---|---| | |
| 66 | +| `varian2014big` | Varian (2014), *JEP* | ML-for-econometrics canon | | |
| 67 | +| `athey2019machine` | Athey & Imbens (2019), *Annu. Rev. Econ.* | Canonical ŷ-vs-β̂ framing | | |
| 68 | +| `breiman2001random` | Breiman (2001), *Machine Learning* | Random Forest benchmark used in the paper | | |
| 69 | +| `friedman2001greedy` | Friedman (2001), *Annals of Statistics* | Gradient boosting primary source | | |
| 70 | +| `park2015using` | Park & Bae (2015), *ESWA* | Real ML-housing-prediction reference replacing a fabricated one (see §3) | | |
| 71 | +| `steurer2021metrics` | Steurer, Hill & Pfeifer (2021), *J. Property Research* | AVM evaluation-metrics literature; grounds deployment discussion | | |
| 72 | + | |
| 73 | +### Explainability (2) | |
| 74 | +| Key | Reference | Why added | | |
| 75 | +|---|---|---| | |
| 76 | +| `ribeiro2016should` | Ribeiro, Singh & Guestrin (2016), KDD | Surrogate-explanation instability | | |
| 77 | +| `rudin2019stop` | Rudin (2019), *Nature MI* | Post-hoc-explanation caution; strengthens SHAP caveats | | |
| 78 | + | |
| 79 | +### Spatial validation (4) | |
| 80 | +| Key | Reference | Why added | | |
| 81 | +|---|---|---| | |
| 82 | +| `valavi2019blockcv` | Valavi et al. (2019), *MEE* | Standard spatial-blocking methodology | | |
| 83 | +| `ploton2020spatial` | Ploton et al. (2020), *Nature Comms* | Dramatic random-vs-spatial validation gap, parallel to the paper's core result | | |
| 84 | +| `wadoux2021spatial` | Wadoux et al. (2021), *Ecological Modelling* | **Dissenting view** — spatial CV pessimistic for interpolation (see §4) | | |
| 85 | +| `pace2020examining` | Pace & Hayunga (2020), *JREFE* | Trees/forests extract spatial signal from hedonic residuals — closest antecedent to the ablation finding | | |
| 86 | + | |
| 87 | +### Climate risk (2) | |
| 88 | +| Key | Reference | Why added | | |
| 89 | +|---|---|---| | |
| 90 | +| `bernstein2019disaster` | Bernstein, Gustafson & Lewis (2019), *JFE* | Sea-level-rise capitalization; supports "missing climate variables" limitation | | |
| 91 | +| `murfin2020risk` | Murfin & Spiegel (2020), *RFS* | Counterpoint within climate literature (weaker capitalization) | | |
| 92 | + | |
| 93 | +## 2. Corrections to existing entries (verified against OpenAlex) | |
| 94 | + | |
| 95 | +- `bourassa2019machine` → **`bourassa2010predicting`** : year was wrong (2019 → **2010**), pages 139–159 → **139–160**, DOI added. | |
| 96 | +- `meyer2019importance` → **`meyer2021predicting`** : year was wrong (2019 → **2021**), DOI added. | |
| 97 | +- `mak2010quantile` : confirmed; author initials completed; DOI added. | |
| 98 | + | |
| 99 | +## 3. ⚠ Fabricated reference found and removed | |
| 100 | + | |
| 101 | +**`chen2020housing`** — “Chen, Hu & Lin (2020), *Housing price prediction using | |
| 102 | +machine learning: A systematic review*, Expert Systems with Applications, | |
| 103 | +145:113142” — **does not exist**. No such work in OpenAlex; neither candidate DOI | |
| 104 | +resolves; no ESWA article with that article number matches. This is precisely the | |
| 105 | +citation-hallucination pattern the verification pass was designed to catch. | |
| 106 | +Replaced in the text by real literature (`park2015using` for ML housing | |
| 107 | +prediction; `bishop2020best`/`rosen1974hedonic` for the SHAP-is-not-WTP claim). | |
| 108 | + | |
| 109 | +## 4. Sections expanded and how | |
| 110 | + | |
| 111 | +| Section | Before → After (source lines) | Changes | | |
| 112 | +|---|---|---| | |
| 113 | +| Introduction | 31 → ~140 | Broader stakes (AVMs, assessment, underwriting); three-tensions framing; leakage contribution moved to center; four enumerated contributions tied to literature strands; practitioner/researcher implications; full roadmap | | |
| 114 | +| Literature | 37 → ~240 | Rebuilt as a 10-theme structured review (origins/identification, functional form, spatial, QR, **listing prices** (new), **capitalization** (new), ML/AVM, XAI, **spatial validation incl. dissent** (new), research gap) with per-strand gap statements | | |
| 115 | +| Data | +2 paragraphs | Listing-price caveat now grounded (Horowitz; Genesove-Mayer; Han-Strange); neighborhood variables tied to capitalization literature | | |
| 116 | +| Methodology | +5 citation edits | QR, boosting (Friedman), RF (Breiman), HC (White + MacKinnon-White), Moran (1950); new paragraph motivating dual validation incl. Wadoux dissent | | |
| 117 | +| Results | +2 targeted notes | School-rating and Walk Score counterintuitive signs now confronted with Black/Bayer and Pivo-Fisher (no numbers touched) | | |
| 118 | +| Discussion | 44 → ~180 | Each message now interprets against published magnitudes; agreements (Sirmans meta-analysis; Zietz/Mak/Waltl QR patterns; Ploton-style validation collapse; Campbell foreclosure discount) and divergences (school-rating sign vs. boundary designs) stated explicitly; Wadoux scope condition; synthesis subsection | | |
| 119 | +| Limitations | +4 citation-grounded items | Listing wedge, spatial models, validation-design duality, climate omission | | |
| 120 | +| Conclusion | 19 → ~85 | Literature-anchored summary; explicit methodological recommendation (report both validations); non-overselling final framing | | |
| 121 | + | |
| 122 | +## 5. Claims flagged for your verification | |
| 123 | + | |
| 124 | +1. **Black (1999) magnitude** — I state “parents pay approximately 2\% more per | |
| 125 | + 5\% increase in test scores” (literature review). This is the commonly quoted | |
| 126 | + headline of the paper; please confirm you are comfortable with this reading. | |
| 127 | +2. **Campbell, Giglio & Pathak (2011) magnitude** — discussion states a “roughly | |
| 128 | + 27\% forced-sale discount … from Massachusetts transactions.” This is the | |
| 129 | + paper's headline foreclosure discount; confirm the framing. | |
| 130 | +3. **“Walk Score's metro-relative construction” conjecture** — the claim that | |
| 131 | + neighborhood scores carry market-specific scale is now explicitly flagged in | |
| 132 | + the discussion as “plausible rather than established.” If you have a source on | |
| 133 | + Walk Score's construction, it could be cited there. | |
| 134 | +4. **Positioning sentence** — “To our knowledge, this framing has not previously | |
| 135 | + been brought to bear on national-scale hedonic housing models” (end of | |
| 136 | + literature §validation). Standard novelty hedge, but worth your sign-off. | |
| 137 | +5. Waltl is cited as **2019** (print issue of *Real Estate Economics* 47(3)); | |
| 138 | + OpenAlex records the online-first year 2016. Either is defensible; I used the | |
| 139 | + print year. | |
| 140 | + | |
| 141 | +## 6. Literature potentially in tension with the paper's findings | |
| 142 | + | |
| 143 | +- **Wadoux et al. (2021)** argue spatial cross-validation is *pessimistically* | |
| 144 | + biased when the estimand is map accuracy over a sampled region. Rather than | |
| 145 | + ignoring it, the paper now engages it in three places (literature, methodology, | |
| 146 | + limitations) and confines its own claims to the extrapolation setting. This is | |
| 147 | + the most important "contradicting" reference; the engagement strengthens the | |
| 148 | + argument but review it. | |
| 149 | +- **Murfin & Spiegel (2020)** find limited sea-level-rise capitalization, in | |
| 150 | + tension with Bernstein et al. (2019); both are cited to avoid one-sided support | |
| 151 | + for the "climate variables matter" limitation. | |
| 152 | +- **Bayer, Ferreira & McMillan (2007)** imply naive cross-sectional school | |
| 153 | + coefficients confound neighbor characteristics — this *supports* the paper's | |
| 154 | + caution but *contradicts* any residual temptation to interpret the school | |
| 155 | + coefficient; the text now explicitly disclaims that interpretation. | |
| 156 | + | |
| 157 | +## 7. Build status | |
| 158 | + | |
| 159 | +`latexmk` clean build: 61 pages, 0 errors, 0 undefined references/citations, | |
| 160 | +69/69 entries cited, hyperlinks resolving. Version 1.1, dated \today at compile | |
| 161 | +time (freeze before submission if desired). | |
modified
paper/main.pdf
+0 −0
Binary file not shown.
modified
paper/main.tex
+1 −1
@@ -150,7 +150,7 @@ | ||
| 150 | 150 | \newcommand{\WPtitle}{Hedonic Housing Price Models for the United States: A Multi-Method Comparison of Parametric, Quantile, and Machine Learning Approaches} |
| 151 | 151 | \newcommand{\WPsubtitle}{} |
| 152 | 152 | \newcommand{\WPdate}{May 2026} |
| 153 | −\newcommand{\WPversion}{1.0} | |
| 153 | +\newcommand{\WPversion}{1.1} | |
| 154 | 154 | \newcommand{\WPabstract}{% |
| 155 | 155 | This paper compares econometric and machine-learning approaches to hedonic housing valuation using 788,842 active Zillow listings across all 50 U.S.\ states and the District of Columbia. A semi-log OLS model with 62 regressors ($R^2 = 0.634$) provides interpretable listing-price gradients; adding ZIP3 fixed effects raises $R^2$ to 0.725, and Moran's $I = 0.27$ confirms strong residual spatial autocorrelation. Quantile regression reveals distributional heterogeneity, with inter-quantile Wald tests rejecting coefficient equality between $\tau = 0.10$ and $\tau = 0.90$ for 11 of 13 key variables. XGBoost achieves $R^2 = 0.833$ under random validation but only 0.425 under state-level geographic holdout; ablation analysis traces the predictive gain primarily to neighborhood-quality features ($+17.6$ pp) and shows that removing geographic features \emph{improves} geographic holdout $R^2$ to 0.519, revealing spatial overfitting. SHAP importance rankings are stable across models (Spearman $\rho > 0.89$). Robustness checks confirm that the lot-size gradient triples when imputed observations are dropped, while other coefficients remain stable under winsorization and subsampling. Throughout, estimates are interpreted as conditional associations in listing prices, not causal willingness-to-pay parameters.% |
| 156 | 156 | } |
modified
paper/references.bib
+361 −12
@@ -34,14 +34,15 @@ at-sign cannot appear here. | ||
| 34 | 34 | pages = {1256--1295}, |
| 35 | 35 | } |
| 36 | 36 | |
| 37 | −@article{bourassa2019machine, | |
| 37 | +@article{bourassa2010predicting, | |
| 38 | 38 | author = {Bourassa, Steven C. and Cantoni, Eva and Hoesli, Martin}, |
| 39 | 39 | title = {Predicting house prices with spatial dependence: {A} comparison of alternative methods}, |
| 40 | 40 | journal = {Journal of Real Estate Research}, |
| 41 | − year = {2019}, | |
| 41 | + year = {2010}, | |
| 42 | 42 | volume = {32}, |
| 43 | 43 | number = {2}, |
| 44 | − pages = {139--159}, | |
| 44 | + pages = {139--160}, | |
| 45 | + doi = {10.1080/10835547.2010.12091276}, | |
| 45 | 46 | } |
| 46 | 47 | |
| 47 | 48 | @article{campbell2011forced, |
@@ -72,13 +73,15 @@ at-sign cannot appear here. | ||
| 72 | 73 | pages = {785--794}, |
| 73 | 74 | } |
| 74 | 75 | |
| 75 | −@article{chen2020housing, | |
| 76 | − author = {Chen, Jian and Hu, Maggie and Lin, Zhenguo}, | |
| 77 | − title = {Housing price prediction using machine learning: {A} systematic review}, | |
| 76 | +@article{park2015using, | |
| 77 | + author = {Park, Byeonghwa and Bae, Jae Kwon}, | |
| 78 | + title = {Using machine learning algorithms for housing price prediction: {T}he case of {F}airfax {C}ounty, {V}irginia housing data}, | |
| 78 | 79 | journal = {Expert Systems with Applications}, |
| 79 | − year = {2020}, | |
| 80 | − volume = {145}, | |
| 81 | − pages = {113142}, | |
| 80 | + year = {2015}, | |
| 81 | + volume = {42}, | |
| 82 | + number = {6}, | |
| 83 | + pages = {2928--2934}, | |
| 84 | + doi = {10.1016/j.eswa.2014.11.040}, | |
| 82 | 85 | } |
| 83 | 86 | |
| 84 | 87 | @incollection{court1939hedonic, |
@@ -243,13 +246,154 @@ at-sign cannot appear here. | ||
| 243 | 246 | } |
| 244 | 247 | |
| 245 | 248 | @article{mak2010quantile, |
| 246 | − author = {Mak, Stephen and Choy, Lennon and Ho, Winky}, | |
| 249 | + author = {Mak, Stephen and Choy, Lennon H. T. and Ho, Winky K. O.}, | |
| 247 | 250 | title = {Quantile regression estimates of {H}ong {K}ong real estate prices}, |
| 248 | 251 | journal = {Urban Studies}, |
| 249 | 252 | year = {2010}, |
| 250 | 253 | volume = {47}, |
| 251 | 254 | number = {11}, |
| 252 | 255 | pages = {2461--2472}, |
| 256 | + doi = {10.1177/0042098009359032}, | |
| 257 | +} | |
| 258 | + | |
| 259 | +@article{genesove2001loss, | |
| 260 | + author = {Genesove, David and Mayer, Christopher}, | |
| 261 | + title = {Loss aversion and seller behavior: {E}vidence from the housing market}, | |
| 262 | + journal = {Quarterly Journal of Economics}, | |
| 263 | + year = {2001}, | |
| 264 | + volume = {116}, | |
| 265 | + number = {4}, | |
| 266 | + pages = {1233--1260}, | |
| 267 | + doi = {10.1162/003355301753265561}, | |
| 268 | +} | |
| 269 | + | |
| 270 | +@article{horowitz1992role, | |
| 271 | + author = {Horowitz, Joel L.}, | |
| 272 | + title = {The role of the list price in housing markets: {T}heory and an econometric model}, | |
| 273 | + journal = {Journal of Applied Econometrics}, | |
| 274 | + year = {1992}, | |
| 275 | + volume = {7}, | |
| 276 | + number = {2}, | |
| 277 | + pages = {115--129}, | |
| 278 | + doi = {10.1002/jae.3950070202}, | |
| 279 | +} | |
| 280 | + | |
| 281 | +@article{han2016role, | |
| 282 | + author = {Han, Lu and Strange, William C.}, | |
| 283 | + title = {What is the role of the asking price for a house?}, | |
| 284 | + journal = {Journal of Urban Economics}, | |
| 285 | + year = {2016}, | |
| 286 | + volume = {93}, | |
| 287 | + pages = {115--130}, | |
| 288 | + doi = {10.1016/j.jue.2016.03.008}, | |
| 289 | +} | |
| 290 | + | |
| 291 | +@article{black1999better, | |
| 292 | + author = {Black, Sandra E.}, | |
| 293 | + title = {Do better schools matter? {P}arental valuation of elementary education}, | |
| 294 | + journal = {Quarterly Journal of Economics}, | |
| 295 | + year = {1999}, | |
| 296 | + volume = {114}, | |
| 297 | + number = {2}, | |
| 298 | + pages = {577--599}, | |
| 299 | + doi = {10.1162/003355399556070}, | |
| 300 | +} | |
| 301 | + | |
| 302 | +@article{bayer2007unified, | |
| 303 | + author = {Bayer, Patrick and Ferreira, Fernando and McMillan, Robert}, | |
| 304 | + title = {A unified framework for measuring preferences for schools and neighborhoods}, | |
| 305 | + journal = {Journal of Political Economy}, | |
| 306 | + year = {2007}, | |
| 307 | + volume = {115}, | |
| 308 | + number = {4}, | |
| 309 | + pages = {588--638}, | |
| 310 | + doi = {10.1086/522381}, | |
| 311 | +} | |
| 312 | + | |
| 313 | +@article{steurer2021metrics, | |
| 314 | + author = {Steurer, Miriam and Hill, Robert J. and Pfeifer, Norbert}, | |
| 315 | + title = {Metrics for evaluating the performance of machine learning based automated valuation models}, | |
| 316 | + journal = {Journal of Property Research}, | |
| 317 | + year = {2021}, | |
| 318 | + volume = {38}, | |
| 319 | + number = {2}, | |
| 320 | + pages = {99--129}, | |
| 321 | + doi = {10.1080/09599916.2020.1858937}, | |
| 322 | +} | |
| 323 | + | |
| 324 | +@article{pace2020examining, | |
| 325 | + author = {Pace, R. Kelley and Hayunga, Darren}, | |
| 326 | + title = {Examining the information content of residuals from hedonic and spatial models using trees and forests}, | |
| 327 | + journal = {Journal of Real Estate Finance and Economics}, | |
| 328 | + year = {2020}, | |
| 329 | + volume = {60}, | |
| 330 | + number = {1--2}, | |
| 331 | + pages = {170--180}, | |
| 332 | + doi = {10.1007/s11146-019-09724-w}, | |
| 333 | +} | |
| 334 | + | |
| 335 | +@article{valavi2019blockcv, | |
| 336 | + author = {Valavi, Roozbeh and Elith, Jane and Lahoz-Monfort, Jos{\'e} J. and Guillera-Arroita, Gurutzeta}, | |
| 337 | + title = {block{CV}: {A}n {R} package for generating spatially or environmentally separated folds for $k$-fold cross-validation of species distribution models}, | |
| 338 | + journal = {Methods in Ecology and Evolution}, | |
| 339 | + year = {2019}, | |
| 340 | + volume = {10}, | |
| 341 | + number = {2}, | |
| 342 | + pages = {225--232}, | |
| 343 | + doi = {10.1111/2041-210X.13107}, | |
| 344 | +} | |
| 345 | + | |
| 346 | +@article{ploton2020spatial, | |
| 347 | + author = {Ploton, Pierre and Mortier, Fr{\'e}d{\'e}ric and R{\'e}jou-M{\'e}chain, Maxime and Barbier, Nicolas and Picard, Nicolas and Rossi, Vivien and Dormann, Carsten F. and Cornu, Guillaume and Viennois, Ga{\"e}lle and Bayol, Nicolas and Lyapustin, Alexei and Gourlet-Fleury, Sylvie and P{\'e}lissier, Rapha{\"e}l}, | |
| 348 | + title = {Spatial validation reveals poor predictive performance of large-scale ecological mapping models}, | |
| 349 | + journal = {Nature Communications}, | |
| 350 | + year = {2020}, | |
| 351 | + volume = {11}, | |
| 352 | + pages = {4540}, | |
| 353 | + doi = {10.1038/s41467-020-18321-y}, | |
| 354 | +} | |
| 355 | + | |
| 356 | +@article{wadoux2021spatial, | |
| 357 | + author = {Wadoux, Alexandre M. J.-C. and Heuvelink, Gerard B. M. and de Bruin, Sytze and Brus, Dick J.}, | |
| 358 | + title = {Spatial cross-validation is not the right way to evaluate map accuracy}, | |
| 359 | + journal = {Ecological Modelling}, | |
| 360 | + year = {2021}, | |
| 361 | + volume = {457}, | |
| 362 | + pages = {109692}, | |
| 363 | + doi = {10.1016/j.ecolmodel.2021.109692}, | |
| 364 | +} | |
| 365 | + | |
| 366 | +@article{bernstein2019disaster, | |
| 367 | + author = {Bernstein, Asaf and Gustafson, Matthew T. and Lewis, Ryan}, | |
| 368 | + title = {Disaster on the horizon: {T}he price effect of sea level rise}, | |
| 369 | + journal = {Journal of Financial Economics}, | |
| 370 | + year = {2019}, | |
| 371 | + volume = {134}, | |
| 372 | + number = {2}, | |
| 373 | + pages = {253--272}, | |
| 374 | + doi = {10.1016/j.jfineco.2019.03.013}, | |
| 375 | +} | |
| 376 | + | |
| 377 | +@article{murfin2020risk, | |
| 378 | + author = {Murfin, Justin and Spiegel, Matthew}, | |
| 379 | + title = {Is the risk of sea level rise capitalized in residential real estate?}, | |
| 380 | + journal = {Review of Financial Studies}, | |
| 381 | + year = {2020}, | |
| 382 | + volume = {33}, | |
| 383 | + number = {3}, | |
| 384 | + pages = {1217--1255}, | |
| 385 | + doi = {10.1093/rfs/hhz134}, | |
| 386 | +} | |
| 387 | + | |
| 388 | +@article{pivo2011walkability, | |
| 389 | + author = {Pivo, Gary and Fisher, Jeffrey D.}, | |
| 390 | + title = {The walkability premium in commercial real estate investments}, | |
| 391 | + journal = {Real Estate Economics}, | |
| 392 | + year = {2011}, | |
| 393 | + volume = {39}, | |
| 394 | + number = {2}, | |
| 395 | + pages = {185--219}, | |
| 396 | + doi = {10.1111/j.1540-6229.2010.00296.x}, | |
| 253 | 397 | } |
| 254 | 398 | |
| 255 | 399 | @incollection{malpezzi2003hedonic, |
@@ -262,14 +406,15 @@ at-sign cannot appear here. | ||
| 262 | 406 | pages = {67--89}, |
| 263 | 407 | } |
| 264 | 408 | |
| 265 | −@article{meyer2019importance, | |
| 409 | +@article{meyer2021predicting, | |
| 266 | 410 | author = {Meyer, Hanna and Pebesma, Edzer}, |
| 267 | 411 | title = {Predicting into unknown space? {E}stimating the area of applicability of spatial prediction models}, |
| 268 | 412 | journal = {Methods in Ecology and Evolution}, |
| 269 | − year = {2019}, | |
| 413 | + year = {2021}, | |
| 270 | 414 | volume = {12}, |
| 271 | 415 | number = {9}, |
| 272 | 416 | pages = {1620--1633}, |
| 417 | + doi = {10.1111/2041-210X.13650}, | |
| 273 | 418 | } |
| 274 | 419 | |
| 275 | 420 | @article{mullainathan2017machine, |
@@ -362,3 +507,207 @@ at-sign cannot appear here. | ||
| 362 | 507 | number = {4}, |
| 363 | 508 | pages = {317--333}, |
| 364 | 509 | } |
| 510 | + | |
| 511 | +@article{waugh1928quality, | |
| 512 | + author = {Waugh, Frederick V.}, | |
| 513 | + title = {Quality factors influencing vegetable prices}, | |
| 514 | + journal = {Journal of Farm Economics}, | |
| 515 | + year = {1928}, | |
| 516 | + volume = {10}, | |
| 517 | + number = {2}, | |
| 518 | + pages = {185--196}, | |
| 519 | + doi = {10.2307/1230278}, | |
| 520 | +} | |
| 521 | + | |
| 522 | +@article{ridker1967determinants, | |
| 523 | + author = {Ridker, Ronald G. and Henning, John A.}, | |
| 524 | + title = {The determinants of residential property values with special reference to air pollution}, | |
| 525 | + journal = {Review of Economics and Statistics}, | |
| 526 | + year = {1967}, | |
| 527 | + volume = {49}, | |
| 528 | + number = {2}, | |
| 529 | + pages = {246--257}, | |
| 530 | + doi = {10.2307/1928231}, | |
| 531 | +} | |
| 532 | + | |
| 533 | +@article{bartik1987estimation, | |
| 534 | + author = {Bartik, Timothy J.}, | |
| 535 | + title = {The estimation of demand parameters in hedonic price models}, | |
| 536 | + journal = {Journal of Political Economy}, | |
| 537 | + year = {1987}, | |
| 538 | + volume = {95}, | |
| 539 | + number = {1}, | |
| 540 | + pages = {81--88}, | |
| 541 | + doi = {10.1086/261442}, | |
| 542 | +} | |
| 543 | + | |
| 544 | +@article{kuminoff2010which, | |
| 545 | + author = {Kuminoff, Nicolai V. and Parmeter, Christopher F. and Pope, Jaren C.}, | |
| 546 | + title = {Which hedonic models can we trust to recover the marginal willingness to pay for environmental amenities?}, | |
| 547 | + journal = {Journal of Environmental Economics and Management}, | |
| 548 | + year = {2010}, | |
| 549 | + volume = {60}, | |
| 550 | + number = {3}, | |
| 551 | + pages = {145--160}, | |
| 552 | + doi = {10.1016/j.jeem.2010.06.001}, | |
| 553 | +} | |
| 554 | + | |
| 555 | +@article{kuminoff2013new, | |
| 556 | + author = {Kuminoff, Nicolai V. and Smith, V. Kerry and Timmins, Christopher}, | |
| 557 | + title = {The new economics of equilibrium sorting and policy evaluation using housing markets}, | |
| 558 | + journal = {Journal of Economic Literature}, | |
| 559 | + year = {2013}, | |
| 560 | + volume = {51}, | |
| 561 | + number = {4}, | |
| 562 | + pages = {1007--1062}, | |
| 563 | + doi = {10.1257/jel.51.4.1007}, | |
| 564 | +} | |
| 565 | + | |
| 566 | +@article{bishop2020best, | |
| 567 | + author = {Bishop, Kelly C. and Kuminoff, Nicolai V. and Banzhaf, H. Spencer and Boyle, Kevin J. and von Gravenitz, Kathrine and Pope, Jaren C. and Smith, V. Kerry and Timmins, Christopher D.}, | |
| 568 | + title = {Best practices for using hedonic property value models to measure willingness to pay for environmental quality}, | |
| 569 | + journal = {Review of Environmental Economics and Policy}, | |
| 570 | + year = {2020}, | |
| 571 | + volume = {14}, | |
| 572 | + number = {2}, | |
| 573 | + pages = {260--281}, | |
| 574 | + doi = {10.1093/reep/reaa001}, | |
| 575 | +} | |
| 576 | + | |
| 577 | +@article{dubin1998spatial, | |
| 578 | + author = {Dubin, Robin A.}, | |
| 579 | + title = {Spatial autocorrelation: {A} primer}, | |
| 580 | + journal = {Journal of Housing Economics}, | |
| 581 | + year = {1998}, | |
| 582 | + volume = {7}, | |
| 583 | + number = {4}, | |
| 584 | + pages = {304--327}, | |
| 585 | + doi = {10.1006/jhec.1998.0236}, | |
| 586 | +} | |
| 587 | + | |
| 588 | +@article{basu1998analysis, | |
| 589 | + author = {Basu, Sabyasachi and Thibodeau, Thomas G.}, | |
| 590 | + title = {Analysis of spatial autocorrelation in house prices}, | |
| 591 | + journal = {Journal of Real Estate Finance and Economics}, | |
| 592 | + year = {1998}, | |
| 593 | + volume = {17}, | |
| 594 | + number = {1}, | |
| 595 | + pages = {61--85}, | |
| 596 | + doi = {10.1023/A:1007703229507}, | |
| 597 | +} | |
| 598 | + | |
| 599 | +@book{anselin1988spatial, | |
| 600 | + author = {Anselin, Luc}, | |
| 601 | + title = {Spatial Econometrics: Methods and Models}, | |
| 602 | + publisher = {Kluwer Academic Publishers}, | |
| 603 | + address = {Dordrecht}, | |
| 604 | + year = {1988}, | |
| 605 | + doi = {10.1007/978-94-015-7799-1}, | |
| 606 | +} | |
| 607 | + | |
| 608 | +@article{moran1950notes, | |
| 609 | + author = {Moran, P. A. P.}, | |
| 610 | + title = {Notes on continuous stochastic phenomena}, | |
| 611 | + journal = {Biometrika}, | |
| 612 | + year = {1950}, | |
| 613 | + volume = {37}, | |
| 614 | + number = {1--2}, | |
| 615 | + pages = {17--23}, | |
| 616 | + doi = {10.1093/biomet/37.1-2.17}, | |
| 617 | +} | |
| 618 | + | |
| 619 | +@article{koenker2001quantile, | |
| 620 | + author = {Koenker, Roger and Hallock, Kevin F.}, | |
| 621 | + title = {Quantile regression}, | |
| 622 | + journal = {Journal of Economic Perspectives}, | |
| 623 | + year = {2001}, | |
| 624 | + volume = {15}, | |
| 625 | + number = {4}, | |
| 626 | + pages = {143--156}, | |
| 627 | + doi = {10.1257/jep.15.4.143}, | |
| 628 | +} | |
| 629 | + | |
| 630 | +@article{mcmillen2008changes, | |
| 631 | + author = {McMillen, Daniel P.}, | |
| 632 | + title = {Changes in the distribution of house prices over time: {S}tructural characteristics, neighborhood, or coefficients?}, | |
| 633 | + journal = {Journal of Urban Economics}, | |
| 634 | + year = {2008}, | |
| 635 | + volume = {64}, | |
| 636 | + number = {3}, | |
| 637 | + pages = {573--589}, | |
| 638 | + doi = {10.1016/j.jue.2008.06.002}, | |
| 639 | +} | |
| 640 | + | |
| 641 | +@article{waltl2019variation, | |
| 642 | + author = {Waltl, Sofie R.}, | |
| 643 | + title = {Variation across price segments and locations: {A} comprehensive quantile regression analysis of the {S}ydney housing market}, | |
| 644 | + journal = {Real Estate Economics}, | |
| 645 | + year = {2019}, | |
| 646 | + volume = {47}, | |
| 647 | + number = {3}, | |
| 648 | + pages = {723--756}, | |
| 649 | + doi = {10.1111/1540-6229.12177}, | |
| 650 | +} | |
| 651 | + | |
| 652 | +@article{varian2014big, | |
| 653 | + author = {Varian, Hal R.}, | |
| 654 | + title = {Big data: {N}ew tricks for econometrics}, | |
| 655 | + journal = {Journal of Economic Perspectives}, | |
| 656 | + year = {2014}, | |
| 657 | + volume = {28}, | |
| 658 | + number = {2}, | |
| 659 | + pages = {3--28}, | |
| 660 | + doi = {10.1257/jep.28.2.3}, | |
| 661 | +} | |
| 662 | + | |
| 663 | +@article{athey2019machine, | |
| 664 | + author = {Athey, Susan and Imbens, Guido W.}, | |
| 665 | + title = {Machine learning methods that economists should know about}, | |
| 666 | + journal = {Annual Review of Economics}, | |
| 667 | + year = {2019}, | |
| 668 | + volume = {11}, | |
| 669 | + pages = {685--725}, | |
| 670 | + doi = {10.1146/annurev-economics-080217-053433}, | |
| 671 | +} | |
| 672 | + | |
| 673 | +@article{breiman2001random, | |
| 674 | + author = {Breiman, Leo}, | |
| 675 | + title = {Random forests}, | |
| 676 | + journal = {Machine Learning}, | |
| 677 | + year = {2001}, | |
| 678 | + volume = {45}, | |
| 679 | + number = {1}, | |
| 680 | + pages = {5--32}, | |
| 681 | + doi = {10.1023/A:1010933404324}, | |
| 682 | +} | |
| 683 | + | |
| 684 | +@article{friedman2001greedy, | |
| 685 | + author = {Friedman, Jerome H.}, | |
| 686 | + title = {Greedy function approximation: {A} gradient boosting machine}, | |
| 687 | + journal = {Annals of Statistics}, | |
| 688 | + year = {2001}, | |
| 689 | + volume = {29}, | |
| 690 | + number = {5}, | |
| 691 | + pages = {1189--1232}, | |
| 692 | + doi = {10.1214/aos/1013203451}, | |
| 693 | +} | |
| 694 | + | |
| 695 | +@article{rudin2019stop, | |
| 696 | + author = {Rudin, Cynthia}, | |
| 697 | + title = {Stop explaining black box machine learning models for high stakes decisions and use interpretable models instead}, | |
| 698 | + journal = {Nature Machine Intelligence}, | |
| 699 | + year = {2019}, | |
| 700 | + volume = {1}, | |
| 701 | + number = {5}, | |
| 702 | + pages = {206--215}, | |
| 703 | + doi = {10.1038/s42256-019-0048-x}, | |
| 704 | +} | |
| 705 | + | |
| 706 | +@inproceedings{ribeiro2016should, | |
| 707 | + author = {Ribeiro, Marco T{\'u}lio and Singh, Sameer and Guestrin, Carlos}, | |
| 708 | + title = {``{W}hy should {I} trust you?'': {E}xplaining the predictions of any classifier}, | |
| 709 | + booktitle = {Proceedings of the 22nd {ACM} {SIGKDD} International Conference on Knowledge Discovery and Data Mining}, | |
| 710 | + year = {2016}, | |
| 711 | + pages = {1135--1144}, | |
| 712 | + doi = {10.1145/2939672.2939778}, | |
| 713 | +} | |
modified
paper/sections/conclusion.tex
+67 −7
@@ -6,16 +6,76 @@ | ||
| 6 | 6 | \section{Conclusion} |
| 7 | 7 | \label{sec:conclusion} |
| 8 | 8 | |
| 9 | −This paper has compared three approaches to hedonic housing price analysis---OLS, quantile regression, and gradient-boosted machine learning with SHAP interpretation---using a large Zillow listing sample of 788,842 U.S.\ properties. Each framework answers a different question, and the robustness extensions in this version document the boundaries of each framework's conclusions. | |
| 9 | +This paper has compared three approaches to hedonic housing price analysis---OLS, | |
| 10 | +quantile regression, and gradient-boosted machine learning with SHAP | |
| 11 | +interpretation---using a national sample of 788,842 U.S.\ Zillow listings. Each | |
| 12 | +framework answers a different question, and the robustness program documents the | |
| 13 | +boundaries of each framework's conclusions. | |
| 10 | 14 | |
| 11 | −OLS provides interpretable conditional mean associations: a price-to-area elasticity of 0.63, a 22.6\% bathroom listing-price gradient, a 33.4\% foreclosure discount, and regional listing-price differentials of 49\%--63\% between coastal and interior markets. The progression from region FE ($R^2 = 0.634$) to ZIP3 FE ($R^2 = 0.725$) demonstrates that fine-grained geographic controls capture substantial variation, but the imputation sensitivity analysis reveals that the lot-size gradient is attenuated threefold by state-median imputation. | |
| 15 | +OLS provides interpretable conditional mean associations whose signs and | |
| 16 | +magnitudes align with the meta-analytic record of hedonic housing studies | |
| 17 | +\citep{sirmans2005composition}: a price-to-area elasticity of 0.63, a 22.6\% | |
| 18 | +bathroom listing-price gradient, a 33.4\% foreclosure discount, and regional | |
| 19 | +listing-price differentials of 49\%--63\% between coastal and interior markets. | |
| 20 | +The progression from region fixed effects ($R^2 = 0.634$) to ZIP3 fixed effects | |
| 21 | +($R^2 = 0.725$) demonstrates---in line with the specification advice of | |
| 22 | +\citet{kuminoff2010which}---that fine-grained geographic controls capture | |
| 23 | +substantial variation, while the imputation sensitivity analysis shows the | |
| 24 | +lot-size gradient is attenuated threefold by state-median imputation. | |
| 12 | 25 | |
| 13 | −Quantile regression reveals that mean effects mask substantial distributional heterogeneity, now formally confirmed by inter-quantile Wald tests that reject coefficient equality for 11 of 13 variables. The garage listing-price gradient is ten times larger at the 10th percentile than the 90th ($z = 28.92$); the pool gradient is insignificant below the median but reaches 9.9\% at $\tau = 0.90$. Subsample stability analysis identifies age as the one coefficient that is genuinely unstable across geographic subsamples. | |
| 26 | +Quantile regression reveals that mean effects mask substantial distributional | |
| 27 | +heterogeneity, formally confirmed by inter-quantile Wald tests that reject | |
| 28 | +coefficient equality for 11 of 13 variables. The garage listing-price gradient is | |
| 29 | +ten times larger at the 10th percentile than the 90th ($z = 28.92$); the pool | |
| 30 | +gradient is insignificant below the median but reaches 9.9\% at $\tau = 0.90$. | |
| 31 | +Subsample stability analysis identifies age as the one coefficient that is | |
| 32 | +genuinely unstable across geographic subsamples. | |
| 14 | 33 | |
| 15 | −XGBoost captures non-linearities and interactions that improve prediction from $R^2 = 0.630$ (OLS) to $R^2 = 0.833$ under random validation. However, this figure substantially overstates generalization ability. The ablation analysis reveals that neighborhood features drive the largest predictive gain ($+17.6$ pp) but that the full model with region dummies actively harms geographic holdout ($-7.6$ pp). The most striking result is that removing all geographic features \emph{improves} geographic holdout $R^2$ from 0.425 to 0.519, while adding latitude and longitude does the opposite ($+3.7$ pp random, $-5.5$ pp geographic). Geographic features help the model memorize location-specific prices but do not improve its ability to generalize attribute-price relationships to new markets. | |
| 34 | +XGBoost captures non-linearities and interactions that improve prediction from | |
| 35 | +$R^2 = 0.630$ (OLS) to $R^2 = 0.833$ under random validation. That figure, | |
| 36 | +however, substantially overstates generalization to new markets. The ablation | |
| 37 | +analysis attributes the largest predictive gain to neighborhood features | |
| 38 | +($+17.6$ pp) and shows that the full model with region dummies actively harms | |
| 39 | +geographic holdout performance ($-7.6$ pp); removing all geographic features | |
| 40 | +\emph{improves} geographic holdout $R^2$ from 0.425 to 0.519, while adding | |
| 41 | +latitude and longitude does the opposite ($+3.7$ pp random, $-5.5$ pp | |
| 42 | +geographic). Geographic features help the model memorize location-specific price | |
| 43 | +levels but do not improve---and can degrade---its ability to transfer | |
| 44 | +attribute-price relationships to unseen markets. | |
| 16 | 45 | |
| 17 | −SHAP values identify the features that contribute most to predictions---living area, bathrooms, school quality, lot size, region---and this ranking is stable across three tree-based models (Spearman $\rho = 0.89$--$0.99$). But these remain predictive decompositions, not structural listing-price gradients. Moran's $I = 0.27$ on OLS residuals---stable across three subsamples (std $= 0.008$)---confirms that regional controls leave substantial spatial dependence unaddressed. | |
| 46 | +SHAP values identify the features contributing most to predictions---living | |
| 47 | +area, bathrooms, school quality, lot size, region---and this ranking is stable | |
| 48 | +across three tree-based models (Spearman $\rho = 0.89$--$0.99$). These remain | |
| 49 | +predictive decompositions rather than implicit prices, a distinction the | |
| 50 | +explainability literature itself insists upon \citep{rudin2019stop}. Moran's | |
| 51 | +$I = 0.27$ on OLS residuals---stable across subsamples---confirms that regional | |
| 52 | +controls leave substantial spatial dependence unaddressed. | |
| 18 | 53 | |
| 19 | −The central contribution is methodological: we demonstrate that model evaluation depends critically on validation design. Random splits overstate generalization, geographic features can harm geographic holdout, SHAP values are stable across models but remain non-structural, and imputation can attenuate key coefficients. Future research should integrate spatial econometric methods, use transaction prices, employ geographic holdout validation as standard practice, incorporate climate risk data, and benchmark against generalized additive models to decompose the ML predictive gain into non-linearity versus spatial partitioning components. | |
| 54 | +The central contribution is methodological: model evaluation in spatially | |
| 55 | +dependent housing data hinges on validation design, and the random-versus-blocked | |
| 56 | +distinction developed in ecology \citep{roberts2017cross, ploton2020spatial, | |
| 57 | +meyer2021predicting, wadoux2021spatial} transfers directly---and consequentially | |
| 58 | +---to hedonic economics. Random splits answer the interpolation question; held-out | |
| 59 | +regions answer the extrapolation question; and the two rank both models and | |
| 60 | +feature sets differently. Studies comparing machine learning to hedonic | |
| 61 | +regression on random splits alone will systematically overstate the practical ML | |
| 62 | +advantage for any application involving new markets. | |
| 20 | 63 | |
| 21 | −The central lesson is not that machine learning replaces hedonic econometrics, but that prediction, distributional heterogeneity, and economic interpretation answer different questions and should be combined carefully in modern housing valuation research. | |
| 64 | +Several extensions follow naturally. Spatial econometric estimation | |
| 65 | +\citep{lesage2009introduction} would model the dependence we only diagnose; | |
| 66 | +replication on transaction prices would quantify the listing-price wedge that the | |
| 67 | +seller-behavior literature predicts \citep{genesove2001loss, han2016role}; | |
| 68 | +climate-risk variables \citep{bernstein2019disaster, murfin2020risk} are a | |
| 69 | +first-order omission our residual diagnostics likely reflect; generalized | |
| 70 | +additive models would decompose the ML gain into non-linearity versus spatial | |
| 71 | +partitioning; and finer geographic blocking would map the continuum between our | |
| 72 | +two validation extremes. We would regard geographic holdout validation, reported | |
| 73 | +alongside random validation, as a reasonable default for future ML-hedonic | |
| 74 | +comparisons. | |
| 75 | + | |
| 76 | +The broader lesson is not that machine learning replaces hedonic econometrics, | |
| 77 | +nor the reverse. Prediction, distributional heterogeneity, and economic | |
| 78 | +interpretation are different questions; each of the tools examined here is strong | |
| 79 | +on exactly one of them, and honest housing-market analysis will continue to | |
| 80 | +require all three---each validated against the question it is actually meant to | |
| 81 | +answer. | |
modified
paper/sections/data.tex
+2 −2
@@ -14,7 +14,7 @@ The dataset is sourced from Zillow, the largest online real estate marketplace i | ||
| 14 | 14 | |
| 15 | 15 | \subsection{Listing Prices versus Transaction Prices} |
| 16 | 16 | |
| 17 | −A critical limitation is that our dependent variable is the \textit{listing} (asking) price, not the realized transaction price. Listing prices reflect seller expectations and strategic pricing behavior, not necessarily market-clearing values. List-to-sale price ratios vary by market condition, property type, and price tier: luxury properties are more frequently overpriced, distressed properties may be strategically underpriced, and regional norms for overbidding versus negotiation differ substantially \citep{knight2002listing}. Throughout this paper, we interpret the estimated coefficients as \textbf{listing-price capitalization gradients}---conditional associations between attributes and asking prices---rather than transaction-price implicit prices. If listing premiums correlate systematically with property attributes (e.g., if waterfront properties are more frequently overpriced), the estimated gradients will reflect this pricing behavior in addition to underlying valuation differences. | |
| 17 | +A critical limitation is that our dependent variable is the \textit{listing} (asking) price, not the realized transaction price. Listing prices are strategic objects: theory models them as commitment devices in seller search \citep{horowitz1992role} and as instruments that direct buyer attention \citep{han2016role}, and empirically they embed seller behavior---overpricing lengthens time-on-market and lowers eventual sale prices \citep{knight2002listing}, while loss-averse sellers systematically set higher asking prices \citep{genesove2001loss}. List-to-sale price ratios therefore vary by market condition, property type, and price tier: luxury properties are more frequently overpriced, distressed properties may be strategically underpriced, and regional norms for overbidding versus negotiation differ substantially. Throughout this paper, we interpret the estimated coefficients as \textbf{listing-price capitalization gradients}---conditional associations between attributes and asking prices---rather than transaction-price implicit prices. If listing premiums correlate systematically with property attributes (e.g., if waterfront properties are more frequently overpriced), the estimated gradients will reflect this pricing behavior in addition to underlying valuation differences. | |
| 18 | 18 | |
| 19 | 19 | \subsection{Sample Construction} |
| 20 | 20 | |
@@ -64,7 +64,7 @@ Binary indicators for swimming pool (39.8\% prevalence), spa (6.1\%), basement ( | ||
| 64 | 64 | |
| 65 | 65 | \subsubsection{Neighborhood and Location Variables (7 Variables)} |
| 66 | 66 | |
| 67 | −Walk Score (0--100), Bike Score, Transit Score, average GreatSchools rating (1--10), school count, distance to nearest school, and local property tax rate. | |
| 67 | +Walk Score (0--100), Bike Score, Transit Score, average GreatSchools rating (1--10), school count, distance to nearest school, and local property tax rate. Walk Score is the accessibility measure for which a capitalization premium has been documented in the literature \citep{pivo2011walkability}; school ratings and tax rates proxy the local public goods whose capitalization into house prices is the subject of a large identification literature \citep{oates1969effects, black1999better, bayer2007unified}. | |
| 68 | 68 | |
| 69 | 69 | \textbf{Data quality note.} Bike Score has a maximum value of 248 in our data, exceeding the expected 0--100 range. Only 38 observations (0.005\%) exceed 100, and these are retained without truncation. Results are robust to capping Bike Score at 100. |
| 70 | 70 | |
modified
paper/sections/discussion.tex
+163 −26
@@ -6,41 +6,178 @@ | ||
| 6 | 6 | \section{Discussion} |
| 7 | 7 | \label{sec:discussion} |
| 8 | 8 | |
| 9 | −The results support four main messages, each associated with a specific modeling framework. | |
| 9 | +The results support four main messages, each associated with a specific modeling | |
| 10 | +framework. For each, we interpret our estimates against the prior literature, | |
| 11 | +noting where they agree with published findings, where they diverge, and what the | |
| 12 | +divergences imply. | |
| 10 | 13 | |
| 11 | 14 | \subsection{Message 1: OLS Remains Useful for Interpretable Associations} |
| 12 | 15 | |
| 13 | −The semi-log OLS model explains 63.4\% of the variation in log listing prices using 62 regressors, comparable to the typical hedonic $R^2$ in the literature \citep{sirmans2005composition}. Its primary value lies in interpretability: the coefficients provide conditional mean associations in a transparent functional form. Regularized linear models (Ridge, Lasso, Elastic Net) achieve essentially identical performance ($R^2 = 0.628$--$0.630$), confirming that the OLS specification is not overfit and that the 20-percentage-point gap relative to XGBoost reflects genuine non-linearities and interactions, not parametric overfitting. | |
| 14 | − | |
| 15 | −The progression from region FE ($R^2 = 0.634$) to state FE ($0.678$) to ZIP3 FE ($0.725$) reveals that a large portion of the unexplained variation is spatial in nature. The ZIP3 FE model closes roughly half the gap between baseline OLS and XGBoost, suggesting that fine-grained geographic controls---rather than non-linear functional forms---account for much of the ML advantage. However, ZIP3 FE introduces 886 parameters and sacrifices the parsimony that makes OLS attractive for economic interpretation. | |
| 16 | − | |
| 17 | −The imputation sensitivity analysis reveals that the baseline lot-size elasticity (0.02) is substantially attenuated by state-median imputation; the complete-case estimate (0.06) is more plausible but applies to a selected subsample. This finding underscores the importance of data quality audits in hedonic research. | |
| 16 | +The semi-log OLS model explains 63.4\% of the variation in log listing prices | |
| 17 | +using 62 regressors---within the range typical of large cross-sectional hedonic | |
| 18 | +studies \citep{sirmans2005composition, malpezzi2003hedonic}---and its headline | |
| 19 | +gradients align well with the meta-analytic record. The living-area elasticity of | |
| 20 | +0.63 is consistent with the concave size-price relationship documented across the | |
| 21 | +125 studies synthesized by \citet{sirmans2005composition}; the positive bathroom | |
| 22 | +and garage gradients and the negative age gradient likewise match the modal signs | |
| 23 | +in that synthesis. The negative conditional bedroom gradient---often treated as an | |
| 24 | +anomaly---is in fact the standard result once total living area is held constant, | |
| 25 | +reflecting a room-size trade-off rather than a distaste for bedrooms. The 33.4\% | |
| 26 | +foreclosure discount is larger than the roughly 27\% forced-sale discount | |
| 27 | +estimated from Massachusetts transactions by \citet{campbell2011forced}, as | |
| 28 | +expected given that our estimate reflects asking-price positioning of distressed | |
| 29 | +listings rather than realized sale prices. Regularized | |
| 30 | +linear benchmarks (Ridge, Lasso, Elastic Net) achieve essentially identical | |
| 31 | +performance ($R^2 = 0.628$--$0.630$), confirming that the parametric specification | |
| 32 | +is not overfit and that the 20-percentage-point gap relative to XGBoost reflects | |
| 33 | +genuine non-linearities and interactions, not estimation noise. | |
| 34 | + | |
| 35 | +The fixed-effects progression carries a sharper lesson. Moving from four Census | |
| 36 | +regions ($R^2 = 0.634$) to state fixed effects ($0.678$) to 886 ZIP3 fixed effects | |
| 37 | +($0.725$) shows that a large share of "unexplained" variation is spatial, echoing | |
| 38 | +the simulation-based advice of \citet{kuminoff2010which} that spatial fixed | |
| 39 | +effects are the single most consequential specification choice in hedonic work, | |
| 40 | +and the best-practice guidance of \citet{bishop2020best}. The ZIP3 model closes | |
| 41 | +roughly half the OLS-to-XGBoost gap, implying that fine geographic partitioning | |
| 42 | +---not flexible functional form---accounts for much of what tree ensembles add. | |
| 43 | +This dovetails with \citet{pace2020examining}, who find that machine-learning | |
| 44 | +methods extract predictive content from the residuals of hedonic models precisely | |
| 45 | +because those residuals retain spatial structure. The residual Moran's $I$ of | |
| 46 | +0.2745 after regional controls---comparable in magnitude to the residual | |
| 47 | +autocorrelation documented in metropolitan samples by \citet{basu1998analysis} | |
| 48 | +---confirms that even our richest interpretable specification leaves the spatial | |
| 49 | +error structure that the spatial econometrics literature has long modeled | |
| 50 | +explicitly \citep{anselin1988spatial, lesage2009introduction, dubin1998spatial}. | |
| 51 | + | |
| 52 | +Two coefficient-level findings deserve confrontation with the capitalization | |
| 53 | +literature. The negative school-rating coefficient contradicts the positive | |
| 54 | +valuation of school quality identified by boundary-discontinuity designs | |
| 55 | +\citep{black1999better, bayer2007unified}. We read this divergence not as evidence | |
| 56 | +against school-quality capitalization but as a textbook illustration of why those | |
| 57 | +designs exist: in a national cross-section, school ratings are collinear with | |
| 58 | +property-tax rates (which enter separately), regional price levels, and unobserved | |
| 59 | +neighborhood composition, so the partial coefficient is uninformative about the | |
| 60 | +structural valuation. Analogously, the negative Walk Score coefficient---despite a | |
| 61 | +documented walkability premium \citep{pivo2011walkability}---reflects conditioning | |
| 62 | +simultaneously on bike and transit accessibility; the three scores are highly | |
| 63 | +correlated, and their premia cannot be attributed variable-by-variable. Both cases | |
| 64 | +reinforce the paper's interpretive stance: hedonic coefficients are conditional | |
| 65 | +associations whose structural content depends on identification that a national | |
| 66 | +listing cross-section does not provide \citep{kuminoff2013new}. | |
| 67 | + | |
| 68 | +Finally, the imputation sensitivity analysis shows the baseline lot-size gradient | |
| 69 | +(0.02) is attenuated roughly threefold by state-median imputation, with the | |
| 70 | +complete-case estimate (0.06) more plausible though subject to selection. The | |
| 71 | +broader point---data-quality audits are not optional in scraped-data hedonics---is | |
| 72 | +a practical extension of the measurement cautions in \citet{bishop2020best}. | |
| 18 | 73 | |
| 19 | 74 | \subsection{Message 2: Quantile Regression Adds Distributional Insight} |
| 20 | 75 | |
| 21 | −Quantile regression reveals that OLS coefficients mask substantial distributional heterogeneity. The garage gradient ratio (10.3:1 between $\tau = 0.10$ and $\tau = 0.90$) and the pool gradient reversal (insignificant at lower quantiles, strongly positive at upper quantiles) demonstrate that the same attribute can play fundamentally different roles across market segments. Inter-quantile Wald tests confirm that these differences are statistically significant for 11 of 13 tested variables, with garage ($z = 28.92$) being the most significantly heterogeneous. | |
| 22 | − | |
| 23 | −This has implications for housing policy: interventions that improve basic amenities (garages, heating systems) may generate the largest listing-price differentials at the lower end, while luxury amenity provision primarily differentiates upper-market properties. | |
| 24 | − | |
| 25 | −The subsample stability analysis shows that most quantile coefficients are highly robust (CV $< 8\%$), but age stands out as unstable (CV = 78.6\%, sign stability = 80\%), reflecting genuine geographic heterogeneity in how property age relates to listing prices. | |
| 76 | +Quantile regression reveals that mean effects mask economically large | |
| 77 | +distributional heterogeneity, and the shape of that heterogeneity is consistent | |
| 78 | +with the housing literature. The declining living-area gradient (0.358 at | |
| 79 | +$\tau = 0.10$ versus 0.251 at $\tau = 0.90$) mirrors the pattern in | |
| 80 | +\citet{zietz2008determinants}, where square footage is valued relatively more in | |
| 81 | +lower-priced homes. The rising pool and lot-size gradients toward the top of the | |
| 82 | +distribution are consistent with amenity-driven luxury pricing documented at upper | |
| 83 | +quantiles in Hong Kong \citep{mak2010quantile} and Sydney \citep{waltl2019variation}. | |
| 84 | +The garage gradient---ten times larger at the bottom decile than the top | |
| 85 | +($z = 28.9$)---extends this logic: basic functional amenities differentiate | |
| 86 | +lower-priced properties where they are scarce, and become nearly universal (hence | |
| 87 | +unpriced) upstream. Such quantile-dependent pricing is one empirical face of | |
| 88 | +housing-market segmentation \citep{goodman1998andrew}: submarkets defined by | |
| 89 | +price tier value the same attribute bundle differently. Our inter-quantile Wald tests formalize what earlier housing | |
| 90 | +QR studies typically showed graphically: coefficient equality is rejected for 11 | |
| 91 | +of 13 attributes, so the conditional listing-price distribution is differentially | |
| 92 | +stretched and compressed by attributes, not merely shifted. This is the | |
| 93 | +cross-sectional analogue of the finding in \citet{mcmillen2008changes} that house | |
| 94 | +price distributions move through coefficients rather than characteristics. | |
| 95 | + | |
| 96 | +For policy, the pattern implies that improvements to basic amenities are | |
| 97 | +associated with the largest proportional listing-price differences at the lower | |
| 98 | +end of the market, while luxury amenities differentiate the top---though we | |
| 99 | +emphasize, again, that these are associations in asking prices, filtered through | |
| 100 | +the seller-behavior mechanisms of \citet{genesove2001loss} and \citet{han2016role}, | |
| 101 | +not causal renovation returns. | |
| 102 | + | |
| 103 | +The subsample stability analysis shows most quantile coefficients are highly | |
| 104 | +robust (CV $< 8\%$, 100\% sign stability), with age the notable exception | |
| 105 | +(CV = 78.6\%). Geographically heterogeneous vintage effects---historic premia in | |
| 106 | +some markets, obsolescence discounts in others---are the natural reading, and they | |
| 107 | +caution against interpreting any single national age gradient too literally. | |
| 26 | 108 | |
| 27 | 109 | \subsection{Message 3: ML Performance Depends Critically on Validation Design} |
| 28 | 110 | |
| 29 | −The XGBoost model achieves substantially better predictions than OLS under random validation ($R^2 = 0.833$ vs.\ $0.630$), capturing non-linearities and interactions that the parametric specification cannot. However, the geographic holdout results provide an important corrective: under state-level holdout, XGBoost achieves only $R^2 = 0.425$--$0.547$ depending on specification. | |
| 30 | − | |
| 31 | −The ablation analysis pinpoints what drives the ML gain. Neighborhood features contribute $+17.6$ pp in random $R^2$---the largest single gain---but only $+3.5$ pp in geographic holdout $R^2$. This disparity suggests that neighborhood variables (school ratings, Walk Score, etc.) are highly informative within observed markets but carry location-specific scale and meaning that may not transfer. The full model with region dummies actually \emph{reduces} geographic holdout $R^2$ by 7.6 pp relative to the market-status-only specification. | |
| 32 | − | |
| 33 | −The most striking finding is that removing all geographic features \emph{improves} geographic holdout $R^2$ from 0.425 to 0.519. This is the opposite of what intuition might suggest and has direct implications for AVM deployment: in contexts requiring geographic generalization, simpler models without explicit geography may outperform richer ones. | |
| 111 | +XGBoost improves on OLS by 20 percentage points of $R^2$ under random validation | |
| 112 | +(0.833 vs.\ 0.630)---squarely within the 15--25\% error-reduction range reported in | |
| 113 | +the ML-valuation literature \citep{bourassa2010predicting, park2015using, | |
| 114 | +chen2016xgboost}. Had we stopped there, the paper would read as one more | |
| 115 | +confirmation that boosting beats hedonics. The geographic holdout overturns that | |
| 116 | +reading: under state-level holdout XGBoost attains only $R^2 = 0.425$--$0.547$ | |
| 117 | +depending on specification. The 29--41-point collapse is strikingly similar in | |
| 118 | +kind to what \citet{ploton2020spatial} document for ecological mapping models, | |
| 119 | +and it validates, in a housing context, the warnings of \citet{roberts2017cross} | |
| 120 | +and \citet{meyer2021predicting} about evaluating spatial predictions on randomly | |
| 121 | +held-out data. | |
| 122 | + | |
| 123 | +The ablation analysis identifies \emph{what} fails to transfer. Neighborhood | |
| 124 | +features contribute the largest random-validation gain ($+17.6$ pp) but only | |
| 125 | +$+3.5$ pp under geographic holdout---consistent with our conjecture that scores | |
| 126 | +like Walk Score and school ratings carry market-specific scaling; we flag this | |
| 127 | +mechanism as plausible rather than established, since we do not test it directly. | |
| 128 | +More striking, the features most obviously "about" geography are actively harmful | |
| 129 | +out-of-market: the full model with region dummies loses 7.6 pp of geographic | |
| 130 | +holdout $R^2$ relative to the market-status stage, and adding raw coordinates | |
| 131 | +---the most flexible location encoding---produces the largest interpolation gain | |
| 132 | +($+3.7$ pp random) alongside a large extrapolation loss. Region dummies and | |
| 133 | +coordinates let trees partition price levels by place, which is exactly what | |
| 134 | +\citet{pace2020examining} show residual-based ML exploits, and exactly what cannot | |
| 135 | +transfer to states never seen in training. | |
| 136 | + | |
| 137 | +We stress the scope of this conclusion, in light of the debate opened by | |
| 138 | +\citet{wadoux2021spatial}: for \emph{interpolation}---valuing a property in a | |
| 139 | +market represented in training data, the typical AVM production setting---random | |
| 140 | +validation is informative and the geographic features earn their keep. Our claim | |
| 141 | +concerns \emph{extrapolation} to unrepresented markets, where blocked validation | |
| 142 | +is the appropriate benchmark \citep{valavi2019blockcv, meyer2021predicting} and | |
| 143 | +where the AVM-evaluation literature already counsels reporting more than headline | |
| 144 | +accuracy \citep{steurer2021metrics}. The practical rule for deployment follows: | |
| 145 | +match the validation protocol to the deployment question, and if the question | |
| 146 | +involves new markets, prefer the leaner, geography-free specification---it costs | |
| 147 | +0.3 pp of interpolation accuracy and buys 9.4 pp of extrapolation accuracy. | |
| 34 | 148 | |
| 35 | 149 | \subsection{Message 4: SHAP Interprets Prediction, Not Economics} |
| 36 | 150 | |
| 37 | −SHAP values provide useful prediction-level decompositions that identify which features drive the XGBoost model's predictions for individual properties. The cross-model stability analysis ($\rho = 0.89$--$0.99$ across three tree models) addresses the common criticism that SHAP rankings are model-specific: while individual SHAP values differ, the ranking of feature importance is remarkably consistent, with the same six features dominating all three models. | |
| 38 | − | |
| 39 | −However, SHAP values should not be equated with hedonic listing-price gradients: | |
| 40 | −\begin{itemize}[nosep] | |
| 41 | − \item SHAP values decompose predictions, not the data-generating process. A feature with a large SHAP value may be predictively important because it proxies for unobserved variables, not because it has a large association with listing prices in the structural sense. | |
| 42 | − \item Feature correlation distributes SHAP contributions among correlated features in ways that may not reflect economic importance. | |
| 43 | − \item SHAP importance rankings differ from OLS coefficient rankings (e.g., school rating ranks 3rd by SHAP but has a counterintuitive negative OLS sign), reflecting the different questions each framework answers. | |
| 44 | −\end{itemize} | |
| 45 | − | |
| 46 | −In the terminology of \citet{mullainathan2017machine}, SHAP is a tool for interpreting predictions. The hedonic gradient $\partial P / \partial z_k$ is a tool for understanding market structure. These serve different purposes and should not be conflated. | |
| 151 | +SHAP values identify living area, bathrooms, school quality, lot size, and region | |
| 152 | +as the dominant predictive contributors, and this ranking is remarkably stable | |
| 153 | +across three tree ensembles ($\rho = 0.89$--$0.99$). The stability result addresses | |
| 154 | +a genuine concern in the explainability literature---that importance rankings are | |
| 155 | +artifacts of a particular fitted model \citep{ribeiro2016should}---and parallels | |
| 156 | +the motivation for model-agnostic explanation methods. But stability is not | |
| 157 | +structure. Three cautions from the literature apply directly. First, SHAP | |
| 158 | +decomposes predictions, not the data-generating process; a feature may earn a | |
| 159 | +large SHAP value by proxying unobservables \citep{lundberg2017unified, | |
| 160 | +lundberg2020local}. Second, correlated features share contributions in ways that | |
| 161 | +defeat attribute-level economic interpretation---our school-rating case is again | |
| 162 | +illustrative, ranking 3rd by SHAP while its OLS sign is negative and its credible | |
| 163 | +structural valuation \citep{black1999better, bayer2007unified} is positive. | |
| 164 | +Third, as \citet{rudin2019stop} argues, post-hoc explanation of a black box is not | |
| 165 | +a substitute for an interpretable model when stakes are high; in our framework, | |
| 166 | +the interpretable model (OLS/QR) and the black box answer different questions, | |
| 167 | +and SHAP does not convert the latter into the former. In Rosen's terms: SHAP | |
| 168 | +values are not implicit prices, and the stability we document should raise | |
| 169 | +confidence in SHAP as a description of \emph{this prediction technology}, not as | |
| 170 | +a measurement of \emph{market valuation}. | |
| 171 | + | |
| 172 | +\subsection{Synthesis} | |
| 173 | + | |
| 174 | +Across the four messages, a single theme recurs: each framework is reliable | |
| 175 | +precisely within the question it was built to answer. OLS with rich spatial | |
| 176 | +controls yields stable, literature-consistent conditional associations; quantile | |
| 177 | +regression reveals formally significant distributional structure; boosting | |
| 178 | +delivers real interpolation gains whose extrapolation content must be established | |
| 179 | +by design, not assumed; and SHAP describes the predictive machine without | |
| 180 | +licensing economic claims. The methodological corollary---that validation design | |
| 181 | +and feature choice interact, and that random-split comparisons overstate ML | |
| 182 | +advantages for out-of-market questions---is, we believe, the paper's most | |
| 183 | +transferable lesson for both the hedonic and the AVM literatures. | |
modified
paper/sections/introduction.tex
+109 −15
@@ -6,28 +6,122 @@ | ||
| 6 | 6 | \section{Introduction} |
| 7 | 7 | \label{sec:introduction} |
| 8 | 8 | |
| 9 | −Housing is the dominant asset class in the typical American household's portfolio and a central object of study in urban economics, household finance, and public policy. The hedonic pricing framework, formalized by \citet{rosen1974hedonic} and rooted in the consumer theory of \citet{lancaster1966new}, provides the canonical approach to decomposing observed housing prices into the implicit valuations of constituent characteristics. Despite nearly five decades of applied hedonic research, three limitations persist. | |
| 9 | +Housing is the dominant asset in the typical American household's portfolio, the | |
| 10 | +collateral underpinning the largest class of household debt, and a central object | |
| 11 | +of study in urban economics, household finance, and public policy. How housing | |
| 12 | +attributes map into prices matters far beyond academia: property-tax assessment, | |
| 13 | +mortgage underwriting, and the automated valuation models (AVMs) that increasingly | |
| 14 | +mediate transactions all rest on some version of that mapping. The hedonic pricing | |
| 15 | +framework---formalized by \citet{rosen1974hedonic} on the consumer-theoretic | |
| 16 | +foundations of \citet{lancaster1966new}, with empirical roots stretching back to | |
| 17 | +\citet{waugh1928quality} and \citet{court1939hedonic}---provides the canonical | |
| 18 | +approach: decompose observed prices into the implicit valuations of constituent | |
| 19 | +characteristics. Yet after five decades of applied hedonic research, three | |
| 20 | +methodological tensions remain unresolved, and the arrival of machine-learning | |
| 21 | +valuation has sharpened rather than settled them. | |
| 10 | 22 | |
| 11 | −First, the standard OLS hedonic model imposes a linear relationship between attributes and log-prices, an assumption that may poorly approximate the complex, non-linear, and interactive relationships governing housing markets. Second, by estimating conditional mean effects, OLS constrains implicit prices to be constant across the price distribution---a restriction that quantile regression studies have shown to be empirically untenable \citep{zietz2008determinants, liao2012hedonic}. Third, while machine learning methods have demonstrated substantial predictive gains in property valuation, their ``black box'' nature has limited adoption in settings where economic interpretation matters \citep{mullainathan2017machine}. | |
| 23 | +First, the standard OLS hedonic model imposes a (log-)linear relationship between | |
| 24 | +attributes and prices, an assumption chosen largely for robustness under attribute | |
| 25 | +omission \citep{cropper1988choice} but one that may poorly approximate the | |
| 26 | +non-linear, interactive structure of housing markets. Second, by estimating | |
| 27 | +conditional mean effects, OLS constrains implicit prices to be constant across the | |
| 28 | +price distribution---a restriction the quantile-regression literature has shown to | |
| 29 | +be empirically untenable \citep{zietz2008determinants, mcmillen2008changes, | |
| 30 | +liao2012hedonic}. Third, while gradient-boosted models deliver large predictive | |
| 31 | +gains in property valuation \citep{bourassa2010predicting, kok2017big}, their | |
| 32 | +black-box character limits adoption where economic interpretation matters | |
| 33 | +\citep{mullainathan2017machine, athey2019machine}---and, as this paper documents, | |
| 34 | +their headline accuracy can be a partial illusion created by the way they are | |
| 35 | +evaluated. | |
| 12 | 36 | |
| 13 | −This paper compares three approaches to hedonic valuation on a single large dataset: | |
| 37 | +This paper confronts the three tensions on a single dataset of 788,842 active | |
| 38 | +Zillow listings covering all 50 U.S. states and the District of Columbia, with 62 | |
| 39 | +regressors spanning structural attributes, lot characteristics, amenities, | |
| 40 | +neighborhood quality, market status, and six interaction terms. We estimate three | |
| 41 | +families of models, each answering a distinct question: | |
| 14 | 42 | \begin{enumerate}[nosep] |
| 15 | − \item \textbf{Semi-log OLS} with HC3 robust standard errors for interpretable conditional mean associations; | |
| 16 | − \item \textbf{Quantile regression} at five quantiles for distribution-specific listing-price gradients; | |
| 17 | − \item \textbf{Gradient-boosted models} (XGBoost, LightGBM) with SHAP-based interpretation for non-linear prediction. | |
| 43 | + \item \textbf{Semi-log OLS} with HC3 robust standard errors---and progressively | |
| 44 | + finer geographic fixed effects---for interpretable conditional mean | |
| 45 | + associations; | |
| 46 | + \item \textbf{Quantile regression} at five quantiles, with formal | |
| 47 | + inter-quantile Wald tests, for distribution-specific listing-price gradients; | |
| 48 | + \item \textbf{Gradient-boosted ensembles} (XGBoost, LightGBM) with SHAP-based | |
| 49 | + interpretation \citep{lundberg2017unified} for non-linear prediction. | |
| 18 | 50 | \end{enumerate} |
| 19 | 51 | |
| 20 | −The dataset comprises 788,842 residential properties listed for sale on Zillow across all 50 U.S.\ states and the District of Columbia. We specify 62 regressors capturing structural attributes, lot characteristics, amenities, neighborhood quality, market status, and six strategically designed interaction terms. | |
| 52 | +The core methodological contribution is a systematic examination of | |
| 53 | +\textbf{spatial leakage} in hedonic model evaluation. Housing data are spatially | |
| 54 | +dependent: nearby properties share unobserved local price determinants | |
| 55 | +\citep{dubin1998spatial, basu1998analysis}. Standard random train-test splits | |
| 56 | +therefore allow information to leak from training to test sets through spatial | |
| 57 | +proximity, inflating out-of-sample performance---a phenomenon documented in | |
| 58 | +ecology and geostatistics \citep{roberts2017cross, ploton2020spatial, | |
| 59 | +meyer2021predicting} but rarely confronted in housing economics. We quantify its | |
| 60 | +consequences through three complementary designs: (i) a \emph{geographic holdout} | |
| 61 | +in which ten entire states (438,315 listings) are withheld from training; (ii) a | |
| 62 | +six-stage \emph{ablation cascade} that traces the XGBoost predictive gain to | |
| 63 | +specific feature groups under both validation schemes and reveals that region | |
| 64 | +dummies actively \emph{harm} geographic generalization; and (iii) a | |
| 65 | +\emph{coordinate experiment} in which adding raw latitude and longitude boosts | |
| 66 | +random-split $R^2$ by 3.7 percentage points while \emph{reducing} geographic | |
| 67 | +holdout $R^2$ by 5.5 percentage points. Together these designs show that the gap | |
| 68 | +between random and geographic validation is not a matter of degree: the two | |
| 69 | +protocols measure qualitatively different model capabilities---interpolation | |
| 70 | +within observed markets versus extrapolation to new ones---and they rank feature | |
| 71 | +sets differently. | |
| 21 | 72 | |
| 22 | −A core methodological contribution of this paper is the systematic examination of \textbf{spatial leakage} in hedonic model evaluation. Standard random train-test splits allow geographically proximate properties---which share unobserved local price determinants---to appear in both training and test sets, artificially inflating out-of-sample performance metrics. We document the consequences through three complementary analyses: (i) a geographic holdout design in which 10 entire states are withheld from training; (ii) an ablation study that traces the XGBoost predictive gain to specific feature groups and reveals that region dummies \emph{harm} geographic generalization; and (iii) an experiment adding raw latitude and longitude coordinates, which boosts random $R^2$ by 3.7 percentage points but \emph{reduces} geographic holdout $R^2$ by 5.5 percentage points. These findings demonstrate that the gap between random and geographic validation is not merely a matter of degree but reflects qualitatively different assessments of model capability. | |
| 23 | − | |
| 24 | −Our contribution is methodological and empirical, not causal. The analysis is not designed to estimate structural willingness-to-pay parameters. Rather, it compares how different modeling frameworks summarize and predict listing-price variation, and it documents the gains and losses from moving along the interpretability-prediction frontier. We emphasize four findings: | |
| 73 | +Our contribution is methodological and empirical, not causal. Following the | |
| 74 | +identification literature \citep{bartik1987estimation, ekeland2004identification, | |
| 75 | +kuminoff2013new, bishop2020best}, we interpret all estimates as conditional | |
| 76 | +associations in \emph{listing} prices---capitalization gradients in asking | |
| 77 | +prices---rather than structural willingness-to-pay parameters, and we engage the | |
| 78 | +listing-price microstructure literature \citep{horowitz1992role, genesove2001loss, | |
| 79 | +han2016role} when drawing that distinction. Within that discipline, the paper | |
| 80 | +makes four contributions: | |
| 25 | 81 | |
| 26 | 82 | \begin{enumerate}[nosep] |
| 27 | − \item OLS with progressively finer geographic fixed effects (region $\rightarrow$ state $\rightarrow$ ZIP3) traces how spatial granularity drives explanatory power from $R^2 = 0.634$ to $0.725$. | |
| 28 | − \item Quantile regression documents monotonic variation in several attribute gradients; inter-quantile Wald tests reject coefficient equality for 11 of 13 variables between the 10th and 90th percentiles. | |
| 29 | − \item XGBoost achieves $R^2 = 0.833$ under random validation but only $R^2 = 0.425$ under geographic holdout; removing geographic features \emph{improves} geographic holdout to $R^2 = 0.519$. | |
| 30 | − \item SHAP importance rankings are stable across three tree-based models (Spearman $\rho > 0.89$), lending credibility to the ranking even though individual SHAP values remain model-specific. | |
| 83 | + \item \textbf{Geographic granularity in OLS.} Progressively finer spatial | |
| 84 | + fixed effects (region $\rightarrow$ state $\rightarrow$ ZIP3) raise $R^2$ from | |
| 85 | + 0.634 to 0.725, providing national-scale, prediction-oriented evidence for the | |
| 86 | + specification advice of \citet{kuminoff2010which}; residual Moran's $I$ of | |
| 87 | + 0.27 shows that even ZIP3 controls leave substantial spatial structure | |
| 88 | + unabsorbed. | |
| 89 | + \item \textbf{Formal distributional heterogeneity.} Inter-quantile Wald tests | |
| 90 | + reject coefficient equality between the 10th and 90th conditional percentiles | |
| 91 | + for 11 of 13 key attributes; the pattern (garage and living-area gradients | |
| 92 | + concentrated at the bottom of the distribution, pool and lot-size gradients at | |
| 93 | + the top) is stable across ten independent subsamples. | |
| 94 | + \item \textbf{Anatomy of the ML advantage.} XGBoost's $R^2 = 0.833$ under | |
| 95 | + random validation falls to 0.425--0.547 under state-level holdout; removing | |
| 96 | + all geographic features \emph{improves} geographic holdout $R^2$ to 0.519. | |
| 97 | + The ablation traces the random-validation gain primarily to | |
| 98 | + neighborhood-quality variables ($+17.6$ pp) and shows their contribution | |
| 99 | + largely fails to transfer across states. | |
| 100 | + \item \textbf{Cross-model stability of SHAP.} Importance rankings correlate at | |
| 101 | + $\rho = 0.89$--$0.99$ across XGBoost, LightGBM, and random forests, supporting | |
| 102 | + SHAP as a description of predictive structure while our framing---informed by | |
| 103 | + the explainability debate \citep{rudin2019stop}---keeps it distinct from | |
| 104 | + Rosen's implicit prices. | |
| 31 | 105 | \end{enumerate} |
| 32 | 106 | |
| 33 | −The remainder of this paper is organized as follows. Section~\ref{sec:literature} reviews the literature. Section~\ref{sec:data} describes the data, sample construction, and variable coding. Section~\ref{sec:methodology} presents the methodology. Section~\ref{sec:results} reports estimation results. Section~\ref{sec:robustness} presents robustness checks. Section~\ref{sec:discussion} discusses implications. Section~\ref{sec:limitations} enumerates limitations. Section~\ref{sec:conclusion} concludes. | |
| 107 | +For practitioners, the message is direct: in AVM deployment contexts that require | |
| 108 | +generalization to unfamiliar markets, validation design is not a technicality but | |
| 109 | +the difference between a model that appears excellent and one that actually | |
| 110 | +transfers; and geographic features, however helpful in-sample, can be | |
| 111 | +counterproductive out-of-market \citep{steurer2021metrics}. For researchers, the | |
| 112 | +results argue that random-split accuracy comparisons between ML and hedonic | |
| 113 | +models---now common in the literature---systematically overstate the ML advantage | |
| 114 | +whenever the deployment question involves new locations. | |
| 115 | + | |
| 116 | +The remainder of the paper is organized as follows. | |
| 117 | +Section~\ref{sec:literature} situates the study in the hedonic, quantile, | |
| 118 | +machine-learning, and spatial-validation literatures. | |
| 119 | +Section~\ref{sec:data} describes the data, sample construction, variable coding, | |
| 120 | +and the listing-price caveat. Section~\ref{sec:methodology} presents the | |
| 121 | +econometric and machine-learning methodology, the three validation designs, and | |
| 122 | +the interpretation framework. Section~\ref{sec:results} reports estimation | |
| 123 | +results across the three model families. Section~\ref{sec:robustness} presents | |
| 124 | +imputation, winsorization, and subsample-stability checks. | |
| 125 | +Section~\ref{sec:discussion} interprets the findings against the literature, | |
| 126 | +Section~\ref{sec:limitations} enumerates limitations, and | |
| 127 | +Section~\ref{sec:conclusion} concludes. | |
modified
paper/sections/limitations.tex
+4 −4
@@ -9,11 +9,11 @@ | ||
| 9 | 9 | We summarize the principal limitations in a structured format. |
| 10 | 10 | |
| 11 | 11 | \begin{enumerate}[nosep] |
| 12 | − \item \textbf{Listing prices, not transaction prices.} The dependent variable is the asking price, which may differ from the realized sale price due to strategic pricing, negotiation, and market conditions. Listing-price gradients are not necessarily equivalent to transaction-price implicit prices. | |
| 12 | + \item \textbf{Listing prices, not transaction prices.} The dependent variable is the asking price, which may differ from the realized sale price due to strategic pricing, negotiation, and market conditions \citep{horowitz1992role, genesove2001loss, han2016role}. Listing-price gradients are not necessarily equivalent to transaction-price implicit prices, and the wedge is plausibly correlated with attributes. | |
| 13 | 13 | |
| 14 | 14 | \item \textbf{No causal identification.} All estimates are conditional associations. Without exogenous variation, we cannot distinguish the causal effects of attributes from sorting, supply constraints, and omitted variables \citep{ekeland2004identification, roberts2013endogeneity}. |
| 15 | 15 | |
| 16 | − \item \textbf{Spatial dependence not modeled.} Moran's $I = 0.27$ confirms substantial spatial autocorrelation in OLS residuals. The paper does not estimate spatial lag, spatial error, or geographically weighted regression models, which may affect both efficiency and consistency of estimates. | |
| 16 | + \item \textbf{Spatial dependence not modeled.} Moran's $I = 0.27$ confirms substantial spatial autocorrelation in OLS residuals. The paper does not estimate spatial lag, spatial error, or geographically weighted regression models \citep{anselin1988spatial, lesage2009introduction}, which may affect both efficiency and consistency of estimates. | |
| 17 | 17 | |
| 18 | 18 | \item \textbf{Random validation overstates ML performance.} The random train-test split allows spatial leakage. Under geographic holdout, XGBoost $R^2$ drops from 0.833 to 0.425--0.547 depending on specification. |
| 19 | 19 | |
@@ -21,7 +21,7 @@ We summarize the principal limitations in a structured format. | ||
| 21 | 21 | |
| 22 | 22 | \item \textbf{Missing data and imputation.} Year built (19.4\%), lot size (16.3\%), and transit score (70.1\%) have substantial missingness. State-median imputation preserves geographic variation but attenuates the lot-size coefficient by a factor of three. |
| 23 | 23 | |
| 24 | − \item \textbf{Climate risk variables entirely missing.} Despite their growing importance, flood, fire, heat, wind, and air risk factors could not be incorporated \citep{baldauf2020does}. | |
| 24 | + \item \textbf{Climate risk variables entirely missing.} Despite their growing importance, flood, fire, heat, wind, and air risk factors could not be incorporated. The climate-capitalization literature finds economically significant price effects of exposure \citep{baldauf2020does, bernstein2019disaster, murfin2020risk}, so their omission plausibly contributes to the residual spatial autocorrelation we document. | |
| 25 | 25 | |
| 26 | 26 | \item \textbf{SHAP values are not structural implicit prices.} SHAP provides prediction decompositions, not marginal willingness-to-pay estimates. Cross-model stability supports the ranking but not the economic interpretation. |
| 27 | 27 | |
@@ -33,7 +33,7 @@ We summarize the principal limitations in a structured format. | ||
| 33 | 33 | |
| 34 | 34 | \item \textbf{No hyperparameter tuning via cross-validation.} XGBoost and LightGBM hyperparameters were set based on common defaults rather than optimized via nested cross-validation. Performance could potentially improve with systematic tuning. |
| 35 | 35 | |
| 36 | − \item \textbf{Geographic holdout design is conservative.} Holding out 10 entire states is a stringent test; finer geographic blocking (county-level or MSA-level) might yield intermediate performance estimates that are more relevant for some AVM applications. | |
| 36 | + \item \textbf{Geographic holdout design is conservative.} Holding out 10 entire states is a stringent test; finer geographic blocking (county-level or MSA-level) might yield intermediate performance estimates that are more relevant for some AVM applications \citep{valavi2019blockcv}. Conversely, for pure interpolation objectives, spatially blocked validation can be pessimistically biased \citep{wadoux2021spatial}; our random-split results remain the relevant benchmark for that use case. | |
| 37 | 37 | |
| 38 | 38 | \item \textbf{Single cross-section.} The data represent a snapshot of active listings at one point in time. Temporal variation in hedonic gradients---due to market cycles, interest rate changes, or policy shifts---cannot be examined. |
| 39 | 39 | \end{enumerate} |
modified
paper/sections/literature.tex
+213 −25
@@ -6,34 +6,222 @@ | ||
| 6 | 6 | \section{Literature Review} |
| 7 | 7 | \label{sec:literature} |
| 8 | 8 | |
| 9 | −\subsection{Hedonic Pricing and the Interpretation of Implicit Prices} | |
| 10 | − | |
| 11 | −The theoretical foundations of hedonic pricing were established by \citet{lancaster1966new} and \citet{rosen1974hedonic}. In Rosen's framework, the market price $P$ of a differentiated good is a function of its characteristics $\mathbf{z}$, and the partial derivative $\partial P / \partial z_k$ yields the implicit price of attribute $z_k$. \citet{sirmans2005composition} provide a meta-analysis of 125 hedonic housing studies, documenting consistent positive associations for living area, bathrooms, and garages, and negative associations for property age. | |
| 12 | − | |
| 13 | −An important but often underappreciated distinction is that hedonic price gradients estimated from cross-sectional regressions are conditional associations, not structural demand parameters. \citet{rosen1974hedonic} himself emphasized that identifying demand and supply functions from the hedonic price schedule requires a second stage with instruments---a requirement rarely satisfied in applied work \citep{ekeland2004identification}. This paper follows the conventional first-stage hedonic approach and interprets coefficients as conditional associations rather than causal willingness-to-pay estimates. | |
| 14 | − | |
| 15 | −\subsection{Functional Form, Omitted Variables, and Spatial Dependence} | |
| 16 | − | |
| 17 | −The semi-log specification has become the standard functional form following the Monte Carlo evidence of \citet{cropper1988choice}, who showed it outperforms more complex forms under attribute omission. Nonetheless, non-linearities remain a concern. Researchers have explored Box-Cox transformations \citep{halvorsen1981choice}, semiparametric specifications \citep{anglin1996semiparametric}, and generalized additive models. | |
| 18 | − | |
| 19 | −Spatial dependence poses a fundamental challenge. \citet{can1992specification} demonstrated that ignoring spatial autocorrelation biases hedonic estimates. \citet{lesage2009introduction} formalized the spatial lag, spatial error, and spatial Durbin specifications. The failure to model spatial dependence---which we document in this paper through a Moran's $I$ test---remains a significant limitation of many hedonic studies, including ours. | |
| 9 | +This paper sits at the intersection of four literatures: the hedonic pricing | |
| 10 | +tradition in housing economics, the econometrics of distributional heterogeneity, | |
| 11 | +the rapidly growing machine-learning strand of property valuation, and the | |
| 12 | +methodological literature on validating predictive models under spatial dependence. | |
| 13 | +We review each in turn, emphasizing in every case how it bears on the design choices | |
| 14 | +of this study and where the gap addressed by this paper lies. | |
| 15 | + | |
| 16 | +\subsection{Hedonic Pricing: Origins, Theory, and the Interpretation of Implicit Prices} | |
| 17 | + | |
| 18 | +The practice of regressing prices on product characteristics predates its | |
| 19 | +theoretical justification by several decades. \citet{waugh1928quality} priced | |
| 20 | +quality attributes of vegetables, \citet{court1939hedonic} constructed hedonic | |
| 21 | +price indexes for automobiles, and \citet{griliches1961hedonic} revived the method | |
| 22 | +for quality-adjusted price measurement. In housing, \citet{ridker1967determinants} | |
| 23 | +provided the first regression-based estimates of how a local disamenity---air | |
| 24 | +pollution---is capitalized into residential property values, inaugurating the | |
| 25 | +property-value approach to valuing local public goods. The theoretical foundation | |
| 26 | +arrived with \citet{lancaster1966new}, who recast goods as bundles of | |
| 27 | +characteristics, and \citet{rosen1974hedonic}, who showed that in a competitive | |
| 28 | +market the equilibrium price schedule $P(\mathbf{z})$ traces out a double envelope | |
| 29 | +of bid and offer functions, so that the gradient $\partial P/\partial z_k$ equals | |
| 30 | +the marginal implicit price of attribute $z_k$. \citet{sirmans2005composition} | |
| 31 | +synthesize 125 empirical applications, documenting robust positive gradients for | |
| 32 | +living area, bathrooms, and garages and negative gradients for property age---the | |
| 33 | +same qualitative pattern our conditional-mean estimates reproduce at national scale. | |
| 34 | + | |
| 35 | +A central and often underappreciated distinction is that first-stage hedonic | |
| 36 | +gradients are equilibrium price phenomena, not preference parameters. Recovering | |
| 37 | +willingness-to-pay functions requires a second stage whose identification demands | |
| 38 | +exclusion restrictions rarely available in practice \citep{bartik1987estimation, | |
| 39 | +ekeland2004identification}. Modern syntheses formalize when hedonic estimates can | |
| 40 | +be trusted: \citet{kuminoff2013new} survey the equilibrium-sorting framework that | |
| 41 | +now underpins structural interpretation of housing-market regressions, and | |
| 42 | +\citet{bishop2020best} distill best practices for credible hedonic estimation, | |
| 43 | +emphasizing fine spatial controls, robustness to specification, and transparent | |
| 44 | +treatment of data limitations. Particularly relevant to our fixed-effects | |
| 45 | +progression, \citet{kuminoff2010which} show in large-scale simulations that | |
| 46 | +specifications with rich spatial fixed effects dramatically outperform sparse | |
| 47 | +cross-sectional specifications in recovering marginal willingness to pay when | |
| 48 | +omitted spatial variables correlate with amenities. Our region~$\rightarrow$ | |
| 49 | +state~$\rightarrow$ ZIP3 exercise provides complementary, purely predictive | |
| 50 | +evidence on the same point. Throughout, we follow this literature's counsel and | |
| 51 | +interpret coefficients as conditional associations---listing-price capitalization | |
| 52 | +gradients---rather than structural demand parameters. | |
| 53 | + | |
| 54 | +\subsection{Functional Form and Specification} | |
| 55 | + | |
| 56 | +The semi-log specification became the workhorse of applied hedonics after | |
| 57 | +\citet{cropper1988choice} showed in Monte Carlo experiments that simple forms | |
| 58 | +outperform flexible ones when some attributes are omitted or measured with | |
| 59 | +error---precisely the situation of scraped listing data; \citet{hill2013hedonic} | |
| 60 | +surveys the subsequent evolution of hedonic specification practice in residential | |
| 61 | +applications. Earlier work had already | |
| 62 | +cautioned against atheoretic flexibility: \citet{halvorsen1981choice} demonstrated | |
| 63 | +the sensitivity of Box-Cox hedonic estimates, and \citet{anglin1996semiparametric} | |
| 64 | +documented gains from semiparametric estimation of the price function. This | |
| 65 | +literature frames one of our central questions: how much of the predictive | |
| 66 | +advantage of gradient-boosted trees reflects genuine non-linearity that the | |
| 67 | +semi-log form misses, and how much reflects something else---in our case, spatial | |
| 68 | +memorization. | |
| 69 | + | |
| 70 | +\subsection{Spatial Dependence in Housing Markets} | |
| 71 | + | |
| 72 | +Housing data are spatial data. \citet{can1992specification} showed that ignoring | |
| 73 | +spatial autocorrelation biases hedonic coefficient estimates; | |
| 74 | +\citet{dubin1998spatial} provides an accessible treatment of why residential | |
| 75 | +prices are spatially correlated (shared unobserved neighborhood attributes, | |
| 76 | +market-mediated spillovers); and \citet{basu1998analysis} document strong | |
| 77 | +spatial autocorrelation in Dallas transaction prices even after conditioning on | |
| 78 | +rich attribute sets. The formal apparatus of spatial econometrics---spatial lag, | |
| 79 | +spatial error, and spatial Durbin models---is developed in \citet{anselin1988spatial} | |
| 80 | +and \citet{lesage2009introduction}, with \citet{anselin2010thirty} providing a | |
| 81 | +retrospective. We deliberately stop short of estimating spatial econometric models: | |
| 82 | +our contribution is diagnostic. Using Moran's~$I$ \citep{moran1950notes} on OLS | |
| 83 | +residuals, we quantify the spatial dependence that remains after 62 attributes and | |
| 84 | +regional controls, and we show that it is a stable, structural feature of the data | |
| 85 | +rather than a sampling artifact. \citet{pace2020examining} pursue a complementary | |
| 86 | +strategy, using trees and forests to examine what information hedonic and spatial | |
| 87 | +model residuals still contain; their finding---that machine-learning methods | |
| 88 | +extract signal from residual spatial structure---anticipates our result that | |
| 89 | +tree-based models exploit location rather than only attribute non-linearity. | |
| 20 | 90 | |
| 21 | 91 | \subsection{Quantile Regression and Distributional Heterogeneity} |
| 22 | 92 | |
| 23 | −\citet{zietz2008determinants} pioneered quantile regression in hedonic housing models, showing that implicit prices vary across the conditional price distribution. \citet{liao2012hedonic} documented similar heterogeneity in Australian data, and \citet{mak2010quantile} showed that the view premium in Hong Kong is substantially larger at upper quantiles. These findings motivate our use of quantile regression to examine whether listing-price gradients differ between lower-end and higher-end properties. | |
| 24 | − | |
| 25 | −\subsection{Machine Learning in Automated Valuation Models} | |
| 26 | − | |
| 27 | −\citet{mullainathan2017machine} distinguish between prediction and inference tasks in economics, arguing that machine learning is most naturally suited to prediction. In real estate, \citet{bourassa2019machine} compare tree-based methods to hedonic models, finding 15--25\% reductions in prediction error. \citet{kok2017big} demonstrate the value of large-scale data in property valuation. The key tension is between predictive accuracy and economic interpretability: machine learning models can capture complex non-linearities and interactions, but their coefficients lack the direct economic interpretation of parametric hedonic models. | |
| 28 | − | |
| 29 | −\subsection{Explainable AI and SHAP in Economic Applications} | |
| 30 | − | |
| 31 | −SHAP values \citep{lundberg2017unified}, based on Shapley values from cooperative game theory \citep{shapley1953value}, provide a principled framework for interpreting complex model predictions. TreeSHAP \citep{lundberg2020local} enables efficient computation for tree-based models. However, SHAP values are prediction-level decompositions, not causal estimates. They are model-specific, sensitive to feature correlation, and do not satisfy the conditions required for interpreting them as marginal willingness-to-pay \citep{chen2020housing}. We use SHAP values to describe the predictive structure of our XGBoost model while maintaining this distinction. | |
| 32 | − | |
| 33 | −\subsection{Spatial Validation and Leakage} | |
| 34 | − | |
| 35 | −A growing literature emphasizes that standard random train-test splits can overstate model performance in spatially structured data. \citet{roberts2017cross} demonstrate that spatial and temporal blocking in cross-validation produces substantially lower---but more realistic---performance estimates in ecological models, and this insight applies directly to hedonic pricing. \citet{meyer2019importance} show that spatial validation is critical when predictors exhibit spatial autocorrelation, as in housing data with geographic controls. Our geographic holdout and ablation designs contribute to this emerging literature by documenting how geographic features simultaneously boost random-split performance and degrade geographic generalization. | |
| 93 | +Quantile regression \citep{koenker1978regression, koenker2001quantile} replaces the | |
| 94 | +conditional mean with a family of conditional quantiles, allowing implicit prices | |
| 95 | +to vary across the price distribution. \citet{zietz2008determinants} introduced the | |
| 96 | +approach to housing, finding that square footage and other attributes are priced | |
| 97 | +differently at different quantiles; \citet{mak2010quantile} document quantile-varying | |
| 98 | +gradients in Hong Kong; and \citet{liao2012hedonic} combine quantile regression with | |
| 99 | +spatial methods. \citet{mcmillen2008changes} shows that changes in the | |
| 100 | +\emph{distribution} of house prices over time are driven largely by changes in | |
| 101 | +coefficients rather than in characteristics---evidence that quantile-specific | |
| 102 | +pricing is economically meaningful, not a statistical curiosity. More recently, | |
| 103 | +\citet{waltl2019variation} provides a comprehensive quantile analysis of the Sydney | |
| 104 | +market, documenting systematic variation across both price segments and locations. | |
| 105 | +Our contribution to this strand is scale and formality: we estimate five quantiles | |
| 106 | +on a national sample, verify coefficient stability across ten independent | |
| 107 | +subsamples, and subject the tail differences to formal inter-quantile Wald tests, | |
| 108 | +rejecting coefficient equality for 11 of 13 key attributes. | |
| 109 | + | |
| 110 | +\subsection{Listing Prices, Search, and Seller Behavior} | |
| 111 | + | |
| 112 | +Because our dependent variable is an asking price, the microstructure of listing | |
| 113 | +behavior matters for interpretation. Theory and evidence establish that list prices | |
| 114 | +are strategic objects: \citet{horowitz1992role} models the list price as a | |
| 115 | +commitment device in seller search; \citet{knight2002listing} shows that | |
| 116 | +overpricing lengthens time-on-market and reduces eventual sale prices; | |
| 117 | +\citet{genesove2001loss} demonstrate that loss aversion leads sellers---especially | |
| 118 | +those facing nominal losses---to set systematically higher asking prices; and | |
| 119 | +\citet{han2016role} show that the asking price plays a directing role in buyer | |
| 120 | +search, so that its information content varies across market segments. Two | |
| 121 | +implications follow for our estimates. First, listing-price gradients need not | |
| 122 | +equal transaction-price gradients, and the wedge is likely correlated with | |
| 123 | +attributes (luxury segments and distressed properties exhibit different list-to-sale | |
| 124 | +gaps). Second, this wedge affects all three modeling frameworks identically, so | |
| 125 | +\emph{comparisons across frameworks}---our primary object of interest---are | |
| 126 | +unaffected even where levels must be interpreted cautiously. | |
| 127 | + | |
| 128 | +\subsection{Capitalization of Local Public Goods and Amenities} | |
| 129 | + | |
| 130 | +Several of our regressors proxy local public goods, for which a mature | |
| 131 | +capitalization literature exists. \citet{oates1969effects} initiated the study of | |
| 132 | +property-tax and public-spending capitalization; \citet{black1999better} used | |
| 133 | +school-attendance boundaries to isolate the value of school quality, finding that | |
| 134 | +parents pay approximately 2\% more per 5\% increase in test scores; and | |
| 135 | +\citet{bayer2007unified} embed boundary discontinuities in an equilibrium sorting | |
| 136 | +model, showing that naive cross-sectional estimates confound school quality with | |
| 137 | +neighbor characteristics. This literature disciplines our reading of the | |
| 138 | +counterintuitive negative school-rating coefficient in the OLS results: without | |
| 139 | +boundary-style identification, school ratings in a national cross-section absorb | |
| 140 | +correlated neighborhood and fiscal variation (the property-tax rate enters | |
| 141 | +separately), and the coefficient should not be read as the value of school quality. | |
| 142 | +Similarly, \citet{pivo2011walkability} document a walkability premium using Walk | |
| 143 | +Score---the same measure we employ---while our specification, which conditions | |
| 144 | +simultaneously on bike and transit accessibility, illustrates how collinear | |
| 145 | +accessibility measures split the premium in ways that resist attribute-by-attribute | |
| 146 | +interpretation. Finally, the growing climate-capitalization literature | |
| 147 | +\citep{baldauf2020does, bernstein2019disaster, murfin2020risk} identifies a class | |
| 148 | +of price-relevant risk variables that are entirely missing from our data---a | |
| 149 | +limitation we return to in Section~\ref{sec:limitations}. | |
| 150 | + | |
| 151 | +\subsection{Machine Learning in Property Valuation} | |
| 152 | + | |
| 153 | +The econometrics profession has converged on a division of labor in which machine | |
| 154 | +learning excels at prediction ($\hat{y}$) problems while classical methods target | |
| 155 | +parameter ($\hat{\beta}$) problems \citep{mullainathan2017machine, varian2014big, | |
| 156 | +athey2019machine}. In real estate, this maps onto automated valuation models | |
| 157 | +(AVMs): \citet{bourassa2010predicting} compare methods for exploiting spatial | |
| 158 | +dependence in prediction; \citet{park2015using} document early machine-learning | |
| 159 | +gains in county-level housing data; \citet{kok2017big} describe the shift from | |
| 160 | +manual appraisal to big-data valuation; and \citet{steurer2021metrics} catalogue | |
| 161 | +the metrics appropriate for evaluating AVM performance, emphasizing that headline | |
| 162 | +$R^2$ figures conceal economically relevant tail behavior. The tree-ensemble | |
| 163 | +methods we deploy---random forests \citep{breiman2001random}, gradient boosting | |
| 164 | +\citep{friedman2001greedy}, and their modern implementations XGBoost | |
| 165 | +\citep{chen2016xgboost} and LightGBM \citep{ke2017lightgbm}---dominate tabular | |
| 166 | +prediction tasks of this kind. Against this backdrop, our contribution is to ask | |
| 167 | +not \emph{whether} boosting beats OLS (it does, by 20 percentage points of $R^2$ | |
| 168 | +under random validation) but \emph{what that gap is made of}: the ablation design | |
| 169 | +decomposes it into feature-group contributions, and the geographic holdout reveals | |
| 170 | +that a substantial share reflects spatial memorization rather than transferable | |
| 171 | +attribute-price structure. | |
| 172 | + | |
| 173 | +\subsection{Explainability: SHAP and Its Limits} | |
| 174 | + | |
| 175 | +SHAP values \citep{lundberg2017unified}, rooted in the cooperative-game solution | |
| 176 | +concept of \citet{shapley1953value} and computable efficiently for trees via | |
| 177 | +TreeSHAP \citep{lundberg2020local}, have become the de facto standard for | |
| 178 | +interpreting ensemble predictions. Yet the explainability literature itself urges | |
| 179 | +caution: local surrogate explanations can be unstable \citep{ribeiro2016should}, | |
| 180 | +and \citet{rudin2019stop} argues that post-hoc explanations of black-box models | |
| 181 | +should not be conflated with intrinsically interpretable modeling, particularly in | |
| 182 | +high-stakes settings. In the hedonic context the danger is specific: SHAP values | |
| 183 | +are prediction decompositions, sensitive to feature correlation, and do not satisfy | |
| 184 | +the equilibrium conditions under which Rosen's gradient equals a marginal implicit | |
| 185 | +price \citep{rosen1974hedonic, bishop2020best}. We therefore use SHAP for what it | |
| 186 | +can do---describe the predictive structure of the fitted model---and we probe the | |
| 187 | +robustness of that description by comparing importance rankings across three | |
| 188 | +different tree ensembles, finding rank correlations of 0.89--0.99. | |
| 189 | + | |
| 190 | +\subsection{Validation Under Spatial Dependence} | |
| 191 | + | |
| 192 | +A methodological literature largely developed in ecology and geostatistics warns | |
| 193 | +that random cross-validation overstates predictive skill whenever observations are | |
| 194 | +spatially dependent, because information leaks from training to test folds through | |
| 195 | +spatial proximity. \citet{roberts2017cross} systematize blocking strategies for | |
| 196 | +structured data; \citet{valavi2019blockcv} provide the standard software | |
| 197 | +implementation of spatial blocking; \citet{ploton2020spatial} demonstrate, | |
| 198 | +strikingly, that large-scale ecological mapping models with excellent random-CV | |
| 199 | +scores lose most of their skill under spatial validation; and | |
| 200 | +\citet{meyer2021predicting} formalize the ``area of applicability'' of spatial | |
| 201 | +prediction models. The lesson is not uncontested: \citet{wadoux2021spatial} show | |
| 202 | +that when the goal is map accuracy over a sampled region---an interpolation | |
| 203 | +problem---spatial cross-validation can be \emph{pessimistically} biased, and | |
| 204 | +design-based random validation is appropriate. This debate sharpens rather than | |
| 205 | +undermines our design: predicting prices in entirely unobserved states is an | |
| 206 | +\emph{extrapolation} task, for which held-out-region validation is the relevant | |
| 207 | +benchmark, while our random split answers the interpolation question. Reporting | |
| 208 | +both, and showing that feature sets rank differently under each, is precisely what | |
| 209 | +the debate prescribes. To our knowledge, this framing has not previously been | |
| 210 | +brought to bear on national-scale hedonic housing models. | |
| 36 | 211 | |
| 37 | 212 | \subsection{Research Gap} |
| 38 | 213 | |
| 39 | −Few papers compare OLS, quantile regression, and gradient-boosted models on the same large-scale dataset while carefully distinguishing prediction from inference. Machine learning papers in real estate often focus on accuracy metrics without addressing economic interpretation; hedonic papers often focus on coefficient interpretation without assessing predictive performance. Moreover, the interaction between geographic feature inclusion and validation design has received insufficient attention. This paper bridges both literatures, documenting what each framework reveals---and what it cannot---while providing systematic evidence on spatial leakage and feature ablation. | |
| 214 | +Three gaps emerge from this review. First, the hedonic and AVM literatures rarely | |
| 215 | +meet: machine-learning papers report accuracy without engaging identification and | |
| 216 | +interpretation, while econometric papers report coefficients without assessing | |
| 217 | +predictive generalization. Comparative studies on a single large dataset that treat | |
| 218 | +OLS, quantile regression, and boosting as complementary lenses---each answering a | |
| 219 | +different question---remain scarce. Second, the spatial-validation insights of | |
| 220 | +ecology have barely penetrated housing economics, despite housing being a | |
| 221 | +canonically spatial asset; the interaction between geographic \emph{features} and | |
| 222 | +validation \emph{design} (our ablation and coordinate experiments) is essentially | |
| 223 | +unexplored at national scale. Third, the literature offers little guidance on how | |
| 224 | +explanation tools like SHAP behave across model families in housing applications. | |
| 225 | +This paper addresses all three, at the scale of 788,842 listings spanning every | |
| 226 | +U.S. state, while maintaining the interpretive discipline the identification | |
| 227 | +literature demands. | |
modified
paper/sections/methodology.tex
+13 −6
@@ -46,12 +46,12 @@ where $P_i$ is the listing price, $x_{ik}$ are continuous and binary attributes, | ||
| 46 | 46 | |
| 47 | 47 | Binary (0/1) and dummy variables are not standardized. For these, the percentage listing-price effect is $(e^{\hat{\beta}_k} - 1) \times 100\%$. |
| 48 | 48 | |
| 49 | −Standard errors are computed using the HC3 estimator \citep{mackinnon1985some}. | |
| 49 | +Standard errors are computed using the HC3 estimator \citep{white1980heteroskedasticity, mackinnon1985some}. | |
| 50 | 50 | |
| 51 | 51 | \subsection{Quantile Regression} |
| 52 | 52 | \label{sec:qr_methodology} |
| 53 | 53 | |
| 54 | −Quantile regression \citep{koenker1978regression} models the $\tau$-th conditional quantile: | |
| 54 | +Quantile regression \citep{koenker1978regression, koenker2001quantile} models the $\tau$-th conditional quantile: | |
| 55 | 55 | \begin{equation} |
| 56 | 56 | Q_{\tau}(\ln P_i | \mathbf{x}_i) = \mathbf{x}_i' \boldsymbol{\beta}(\tau), \quad \tau \in (0, 1) |
| 57 | 57 | \end{equation} |
@@ -70,7 +70,7 @@ Rejection indicates that the attribute's association with listing prices differs | ||
| 70 | 70 | |
| 71 | 71 | \subsubsection{XGBoost} |
| 72 | 72 | |
| 73 | −XGBoost \citep{chen2016xgboost} fits an additive ensemble of regression trees by sequentially minimizing a regularized loss function: | |
| 73 | +XGBoost \citep{chen2016xgboost} implements gradient boosting \citep{friedman2001greedy}, fitting an additive ensemble of regression trees by sequentially minimizing a regularized loss function: | |
| 74 | 74 | \begin{equation} |
| 75 | 75 | \mathcal{L}^{(t)} = \sum_{i=1}^{n} l(y_i, \hat{y}_i^{(t-1)} + f_t(\mathbf{x}_i)) + \Omega(f_t) |
| 76 | 76 | \end{equation} |
@@ -82,14 +82,21 @@ LightGBM \citep{ke2017lightgbm} uses Gradient-based One-Side Sampling (GOSS) and | ||
| 82 | 82 | |
| 83 | 83 | \subsubsection{Additional Benchmarks} |
| 84 | 84 | |
| 85 | −We also report results for Ridge regression ($\alpha = 1$), Lasso ($\alpha = 0.001$), Elastic Net, Random Forest (500 trees, depth 20), and OLS with state fixed effects. | |
| 85 | +We also report results for Ridge regression ($\alpha = 1$), Lasso ($\alpha = 0.001$), Elastic Net, Random Forest \citep[500 trees, depth 20;][]{breiman2001random}, and OLS with state fixed effects. | |
| 86 | 86 | |
| 87 | 87 | \textbf{Ablation design.} To understand which feature groups drive the ML predictive gain over OLS, we estimate XGBoost models using progressively richer feature sets: (i) structural attributes only (8 features); (ii) $+$ lot characteristics (9); (iii) $+$ amenity indicators (20); (iv) $+$ neighborhood variables (27); (v) $+$ market status (32); and (vi) the full model including interactions, categorical controls, and region dummies (62). At each stage, we evaluate both random and geographic holdout $R^2$ to identify which feature groups contribute to genuine predictive generalization versus spatial memorization. |
| 88 | 88 | |
| 89 | 89 | \subsection{Validation Designs} |
| 90 | 90 | \label{sec:validation_designs} |
| 91 | 91 | |
| 92 | −We employ three validation strategies: | |
| 92 | +Because housing prices are spatially dependent, the choice of validation protocol | |
| 93 | +is substantive rather than technical: random splits assess interpolation within | |
| 94 | +observed markets, while spatially blocked designs assess extrapolation to new | |
| 95 | +markets \citep{roberts2017cross, valavi2019blockcv, meyer2021predicting}. Blocked | |
| 96 | +validation is not universally preferable---for interpolation objectives it can be | |
| 97 | +pessimistically biased \citep{wadoux2021spatial}---so we report both protocols and | |
| 98 | +interpret each against its own deployment question. We employ three validation | |
| 99 | +strategies: | |
| 93 | 100 | |
| 94 | 101 | \begin{enumerate}[nosep] |
| 95 | 102 | \item \textbf{Random 80/20 split} (seed = 42): 631,073 training, 157,769 test. This is the standard approach but permits spatial leakage---nearby properties from the same neighborhood can appear in both train and test sets. |
@@ -122,7 +129,7 @@ We use SHAP to describe which features contribute most to predictions, not to es | ||
| 122 | 129 | \subsection{Spatial Autocorrelation Diagnostics} |
| 123 | 130 | \label{sec:moran_method} |
| 124 | 131 | |
| 125 | −We compute Moran's $I$ on OLS residuals using a row-standardized KNN spatial weight matrix ($k = 8$) on random subsamples of 5,000 observations. To assess the stability of this diagnostic, we repeat the computation across three independent subsamples and report the mean, standard deviation, and significance of the resulting $I$ statistics. Moran's $I$ is defined as: | |
| 132 | +We compute Moran's $I$ \citep{moran1950notes} on OLS residuals using a row-standardized KNN spatial weight matrix ($k = 8$) on random subsamples of 5,000 observations. To assess the stability of this diagnostic, we repeat the computation across three independent subsamples and report the mean, standard deviation, and significance of the resulting $I$ statistics. Moran's $I$ is defined as: | |
| 126 | 133 | \begin{equation} |
| 127 | 134 | I = \frac{N}{\sum_{i}\sum_{j} w_{ij}} \cdot \frac{\sum_{i}\sum_{j} w_{ij}(e_i - \bar{e})(e_j - \bar{e})}{\sum_{i}(e_i - \bar{e})^2} |
| 128 | 135 | \label{eq:moran} |
modified
paper/sections/results.tex
+2 −2
@@ -128,9 +128,9 @@ Several coefficients warrant discussion. | ||
| 128 | 128 | |
| 129 | 129 | \textbf{Negative waterfront coefficient ($-0.135$, standardized).} The standalone waterfront effect is evaluated at the mean of the standardized interaction term $\text{waterfront} \times \ln(\text{sqft})$. Because the interaction is strongly positive ($+0.137$), the net waterfront effect becomes positive for larger properties. This artifact of the interacted specification does not imply that waterfront reduces price on average. |
| 130 | 130 | |
| 131 | −\textbf{Negative school rating ($-0.054$ per rating point, unstandardized).} This likely reflects confounding with property tax rates and regional effects. In higher-tax jurisdictions, school quality is partially capitalized through the tax rate, which enters separately. Conditional on tax rate, region, and other controls, the residual school rating variation may capture unobserved factors that correlate negatively with prices. This coefficient should be interpreted with caution. | |
| 131 | +\textbf{Negative school rating ($-0.054$ per rating point, unstandardized).} This likely reflects confounding with property tax rates and regional effects. In higher-tax jurisdictions, school quality is partially capitalized through the tax rate, which enters separately. Conditional on tax rate, region, and other controls, the residual school rating variation may capture unobserved factors that correlate negatively with prices. The boundary-discontinuity literature, which isolates school quality from neighborhood composition, consistently finds positive valuations \citep{black1999better, bayer2007unified}; the divergence illustrates why cross-sectional partial coefficients on school quality should not be interpreted structurally. | |
| 132 | 132 | |
| 133 | −\textbf{Negative Walk Score ($-0.0004$ per point, unstandardized).} Walk Score is highly correlated with Bike Score ($r \approx 0.56$) and Transit Score ($r \approx 0.62$). In the presence of all three accessibility measures, the Walk Score coefficient may reflect residual urban density effects---after controlling for biking and transit access, the remaining variation in walkability may capture denser, smaller-lot neighborhoods. | |
| 133 | +\textbf{Negative Walk Score ($-0.0004$ per point, unstandardized).} Walk Score is highly correlated with Bike Score ($r \approx 0.56$) and Transit Score ($r \approx 0.62$). In the presence of all three accessibility measures, the Walk Score coefficient may reflect residual urban density effects---after controlling for biking and transit access, the remaining variation in walkability may capture denser, smaller-lot neighborhoods. A walkability premium is well documented when accessibility enters alone \citep{pivo2011walkability}; with three collinear accessibility scores, the premium is split across coefficients and cannot be attributed variable-by-variable. | |
| 134 | 134 | |
| 135 | 135 | \textbf{Negative basement ($-0.042$).} Basements are predominantly found in older properties in colder regions. Conditional on age, region, and size, the basement indicator may proxy for older construction quality or layout features valued less in modern markets. |
| 136 | 136 | |
| 137 | 137 | |