SPB Git

spb/wp5_uqo Public

UQO Working Paper No. 5 — Airbnb, residential rents and housing market pressure.

TeX 53.4% Python 46.5%
10.3 KB

# AUDIT — Original project immo-wp5-spb-20260519

Audit of ~/Desktop/UQO/UQO_WP/immo-wp5-spb-20260519 performed on 2026-08-05, before any restructuring. The original folder is left untouched; the clean repository is rebuilt at ~/Desktop/wp5_uqo.

# 1. Project summary

Hedonic and spatial econometric analysis of the relationship between Airbnb activity and residential rents in Quebec, Canada (UQO Working Paper No. 5, Simon-Pierre Boucher, May 2026). ~5,000 Airbnb listings + 8,356 Realtor.ca rental listings, spatial buffer merge (Haversine, 250 m–2 km), OLS/hedonic models, spatial models (SAR/SEM via spreg), quantile regressions, ML robustness (LASSO, Elastic Net, RF, GBM, SHAP).

# 2. Scripts (scripts/, 10 files, ~2,930 lines)

Script Role Inputs Outputs
01_load_inspect_data.py Inspect raw data, write data dictionary log airbnb.csv, rent.json (missing, see §5.1) outputs/logs/data_inspection.txt
02_clean_airbnb.py Standardize cities, winsorize price p1/p99, log_price, impute rating (median) & reviews (0), is_entire_home airbnb.csv (missing) data_clean/airbnb_clean.parquet
03_clean_rent.py Filter Single Family + Monthly, parse rent/city/borough/lat/lon/beds/baths/size, winsorize p1/p99, log_rent rent.json (missing) data_clean/rent_clean.parquet
04_merge_data.py Spatial buffer merge (Haversine 250 m/500 m/1 km/2 km, chunked), city/borough aggregation, combined analysis file the 2 clean parquets merged_spatial.parquet, merged_neighborhood.parquet, merged_analysis.parquet
05_descriptive_analysis.py Summary-stat tables, correlation matrix, 7 descriptive figures clean + merged parquets 3 tables, 7 figures
06_hedonic_models.py Models 1a–1e (rent), 2a–2c (Airbnb pricing), 3 fwd/rev (city level), OLS HC1 merged_analysis, airbnb_clean 3 .tex tables
07_spatial_models.py Model 4: SAR (GM_Lag) + SEM (GM_Error) via spreg, KNN(5) weights; buffer robustness 250 m–2 km merged_analysis 2 tables + coefficient_buffer_comparison.pdf
08_quantile_models.py Model 5: quantile regressions τ∈{.10,.25,.50,.75,.90} + fine grid plot merged_analysis 1 table + quantile_coefficients.pdf
09_ml_robustness.py Model 6: OLS/LASSO/ENet/RF/GBM (seed 42, 80/20 split), SHAP merged_analysis ml_comparison.{tex,csv} + 3 figures
10_generate_tables_figures.py Outlier/winsorization + subsample robustness tables, hexbin heatmap, coefficient-robustness plot, quintile bar chart merged_analysis 2 tables + 3 figures

All paths are already relative to the project root via Path(__file__).resolve().parent.parent — good. Scripts are standalone monoliths with heavy duplication (see §5.4).

# 3. Data

  • data_clean/ (5 parquet files, all present and readable): airbnb_clean.parquet (4,950 rows expected), rent_clean.parquet (8,303 rows), merged_spatial.parquet, merged_neighborhood.parquet, merged_analysis.parquet (8,258 rows × 42 cols).
  • Raw data (airbnb.csv, rent.json) are absent from the folder and from the whole disk (searched ~/Desktop/UQO and Spotlight). See §5.1.

# 4. Figure inventory

Figures generated by the code (15) vs figures present in outputs/figures/ (10):

Figure Generated by In outputs/figures/ Referenced by paper
dist_airbnb_price.pdf 05 yes yes
dist_rent.pdf 05 yes yes
airbnb_by_city.pdf 05 yes no
rent_by_city.pdf 05 no no
scatter_airbnb_rent.pdf 05 yes no
map_airbnb.pdf 05 no yes → broken
map_rent.pdf 05 no yes → broken
coefficient_buffer_comparison.pdf 07 yes yes (×2)
quantile_coefficients.pdf 08 no yes → broken
ml_predicted_vs_actual.pdf 09 yes no
feature_importance.pdf 09 yes yes
shap_summary.pdf 09 yes yes
rent_airbnb_heatmap.pdf 10 yes no
coefficient_robustness.pdf 10 yes yes
rent_by_airbnb_bins.pdf 10 no no

The 5 missing PDFs were evidently deleted (or the run predates those functions); the code that produces them is intact. Consequence: wp5/main.tex fails to compile (fatal error, no PDF); only wp5/main_web.tex compiles because it replaces missing figures with placeholder boxes. Re-running the pipeline regenerates all 15 figures and fixes the paper build.

wp5/figures/ contains a stale copy of the same 10 PDFs (duplicate of outputs/figures/).

# 5. Problems found (nothing was changed; fixes happen in the new repo)

# 5.1 Raw data missing (blocking for steps 01–03 only)

airbnb.csv and rent.json are referenced by scripts 01–03 and by the original README but do not exist anywhere on disk. The pipeline is therefore reproducible from data_clean/ onward (steps 04–10 + paper), which covers every number and figure in the paper. Steps 01–03 are kept in the new repo and fail with an explicit message if the raw files are absent. → Requires user review: locate/restore the raw files if full from-scratch reproduction is needed.

# 5.2 Two divergent copies of the paper

  • paper/ — same sections, figures referenced as ../outputs/figures/…, no title page, compiled OK.
  • wp5/ — adds UQO title page (uq_logo.jpg), Makefile, .latexmkrc, \graphicspath{{figures/}}, plus main_web.tex (variant with a "Figure indisponible" fallback macro).
  • Only real content differences: figure paths in 03_data.tex, 05_results.tex, 06_robustness.tex.
  • wp5/ is the canonical/most recent version (June 2026) → used as the base for the rewrite.

# 5.3 Figure caption vs content mismatch

03_data.tex captions the two distribution histograms as "(log scale)" but scripts plot levels (CAD). Caption corrected in the rewrite (editorial fix; figures unchanged).

# 5.4 Code-quality issues (fixed by refactor, results preserved)

  • Significance-star helper duplicated in 4 scripts; LaTeX regression-table builder duplicated in 4 scripts; matplotlib style blocks duplicated in 3 scripts; city-FE OLS helper duplicated within script 10.
  • Scripts 05 requests columns size_sqft, guests_count, amenities_count, num_images, quality_score that don't exist in the parquets (rent_clean has size_interior_sqft); the helper silently skips them, so the published summary tables lack those rows. Kept as-is to reproduce identical tables; noted for future work.
  • Scripts 06–08 are top-level scripts (no main()); 02/03 likewise.
  • scripts/__pycache__/ committed; .DS_Store files scattered.

# 5.5 Dead / duplicate / unused files

  • paper/ entire directory superseded by wp5/ (kept only as reference in the original folder).
  • wp5/main_web.tex + main_web.* build artifacts: web variant with placeholder hack — obsolete once figures are regenerated.
  • LaTeX build artifacts (.aux, .log, .out, .bbl, .blg, .fls, .fdb_latexmk, .toc, .synctex.gz).
  • scripts/__pycache__/, .DS_Store (×6).
  • outputs/figures/ vs wp5/figures/: duplicated PDFs.

# 5.6 Environment

Python 3.14.4 (Homebrew) with pandas 3.0.2, numpy 2.4.4, statsmodels 0.14.6, scikit-learn 1.6.1, scipy 1.17.1, matplotlib 3.10.9, pyarrow 24.0.0, shap 0.48.0, libpysal 4.14.1, spreg 1.9.0. TeX Live (pdflatex, latexmk, bibtex) available. Note: with libpysal+spreg installed, script 07 takes the SAR/SEM branch — consistent with the published spatial_models.tex.

# 6. Paper (wp5/)

  • main.tex (184 lines): clean preamble (newtx, booktabs, natbib/apalike, fancyhdr, hyperref), metadata macros, inputs 8 sections + title page + 3 appendices; 24 BibTeX entries, all cited.
  • Sections: 01 introduction, 02 literature, 03 data, 04 methodology (models 1–6 + identification), 05 results, 06 robustness, 07 discussion, 08 conclusion; appendices: data, methods, robustness.
  • Prose already in good academic English; rewrite pass = flow/consistency polish, caption fix (§5.3), figure/table path normalization to the new layout, and making every figure resolvable.
  • Tables are \input from ../outputs/tables/*.tex → becomes ../results/tables/ in the new repo.

# 7. Verification plan (Phase 2b)

  1. Copy the 5 parquets to data/processed/ (bit-identical, checksummed).
  2. Re-run refactored steps 04→10; compare regenerated merged_*.parquet (DataFrame equality) and every .tex/.csv table against the originals (diff). Figures: confirm regeneration and spot-check values; PDF bytes differ (timestamps) so tables are the numeric ground truth.
  3. Any discrepancy is recorded below — never silently corrected.

# 8. Discrepancies found during verification

Full pipeline (steps 04–10) re-run in the new repo on 2026-08-05 and compared to the originals:

  • Merged parquets (3/3): identicalmerged_spatial, merged_neighborhood, merged_analysis regenerated from the committed clean parquets are value-identical to the originals (pandas.testing.assert_frame_equal).
  • Tables: 14/16 byte-identical after aligning the generators with the published fragments (the original outputs/tables/*.tex had been hand-stripped of their \begin{table}…\end{table} wrappers after generation; the new scripts emit those fragments directly — see CHANGES.md). Residual differences, both preserved and NOT corrected:
    1. quantile_regression.tex — in the τ=0.10 column only, three control coefficients differ in the 4th decimal (bathrooms 0.2410→0.2411, bt_House 0.0607→0.0608, bt_Row/Townhouse 0.1302→0.1310, SE 0.0469→0.0468). The Airbnb coefficient, all other quantiles, and the OLS column are identical. Cause: IRLS convergence jitter of statsmodels.QuantReg at the extreme quantile; no significance level or claim in the paper changes.
    2. ml_comparison.csv — two values differ at the 16th significant digit (~1e-16 relative; float representation). The 4-decimal ml_comparison.tex is byte-identical.
  • Figures: all 15 regenerated, including the 5 missing from the original (map_airbnb, map_rent, quantile_coefficients, rent_by_city, rent_by_airbnb_bins), which un-breaks the paper build. PDF bytes differ from the surviving originals (embedded timestamps/IDs); the numeric ground truth is the table set above, plus the underlying data verified identical.