AUDIT — immo-wp2-spb-20260519
Audit of the original project at ~/Desktop/UQO/UQO_WP/immo-wp2-spb-20260519
(354 MB), performed 2026-08-05 before restructuring into ~/Desktop/wp2_uqo.
The original folder is left untouched and serves as the backup.
1. What the project is
Working Paper No. 2 (UQO): "Decoding Real Estate Descriptions: Semantic
Embeddings and Hedonic Pricing of Residential Properties in Quebec"
(Simon-Pierre Boucher). The analysis embeds the free-text PublicRemarks of
17,087 Quebec single-family house listings with all-MiniLM-L6-v2
(sentence-transformers), computes cosine similarities to 20 researcher-defined
French reference descriptions ("luxury", "needs renovation", …), and adds
those 20 similarity features to a log-price hedonic OLS model (models A–E,
HC3 errors), plus robustness checks (VIF, Breusch–Pagan, quantile regression,
winsorization, 1,000-rep bootstrap, Lasso/Elastic Net).
Headline numbers (paper, model D vs A): adjusted R² 0.452 → 0.511, F = 99.53 (p < 0.001), n = 17,087.
2. Data
| File | Size | Role |
|---|---|---|
louka.db |
268 MB | Raw data. SQLite; table properties (id, mls_number, category, price_value, bedrooms, bathrooms, data JSON blob). Counts: house 17,089, condo 8,586, land 8,506, rent 8,240, plex 4,058. Sample = houses with price > 0 and remarks ≥ 20 chars → 17,087. |
hedonic_maison_results.csv |
17,087 rows | Processed: houses + 20 similarity features (main analysis dataset). |
hedonic_results.csv |
29,731 rows | Processed (earlier PCA approach, house/condo/plex). |
hedonic_cosine_results.csv |
29,731 rows | Processed (earlier 16-reference approach, all categories). |
embeddings_maisons.npy |
25 MB | 17,087 × 384 embeddings (houses; used by the paper). |
embeddings_remarks.npy |
44 MB | 29,731 × 384 embeddings (all categories; earlier approach). |
sim_matrix_maisons.npy |
1.3 MB | 17,087 × 20 similarity matrix (paper). |
similarity_matrix.npy |
1.8 MB | 29,731 × 16 (earlier approach). |
reference_embeddings.npz |
43 KB | 16 reference embeddings (earlier approach). |
⚠️ Notes:
louka.dbalso contains auserstable with plaintext usernames/passwords (application leftover). It is not used by any script. Not copied onward.- All scripts hardcode
DB_PATH = "/Users/simon-pierreboucher/Desktop/d/louka.db"and write outputs to~/Desktop/d/— a folder that no longer exists. The project copy oflouka.dbis what remains; outputs were copied into the project at some point. robustness_results.csv(written byhedonic_maison.pystep 12) was saved to~/Desktop/d/and is not present anywhere — regenerated in the new repo.
3. Scripts (chronological / by role)
| Script | Role | Outputs | Status |
|---|---|---|---|
hedonic_semantic.py (348 l.) |
Exploration #1: PCA(20) of embeddings, house/condo/plex, models 1–4 | hedonic_results.csv, embeddings_remarks.npy |
Superseded; not used by the paper |
hedonic_cosine_references.py (447 l.) |
Exploration #2: 16 reference similarities, all categories, models A–C | hedonic_cosine_results.csv, reference_embeddings.npz, similarity_matrix.npy |
Superseded; not used by the paper |
hedonic_maison.py (818 l.) |
Main analysis: 20 references, houses only, models A–E + all robustness checks | hedonic_maison_results.csv, embeddings_maisons.npy, sim_matrix_maisons.npy, robustness_results.csv (lost) |
Paper source (tables 1–8) |
paper_figures.py (578 l.) |
Figures v1 (figs 1–8) + table1–4_*.csv |
figures/ |
Superseded by v2 for figures; still the only producer of the 4 table CSVs |
paper_figures_v2.py (386 l.) |
Figures v2 (final): figs 1–11, tighter layout | figures/fig1…fig11 |
Paper source (all figures) |
Duplication: the DB-extraction block (~70 lines) and the OLS-fit helper are
copy-pasted in all 5 scripts, with small drifts (half_baths and lat/lon only
in some; 16 vs 20 references; French vs English reference names). The 20
reference descriptions are pasted 3× (French keys in hedonic_maison.py,
English display keys in both figure scripts — same French texts).
4. Figures
figures/ and wp2/figures/ (identical copies) contain fig1, fig4–fig11
(PDF, some PNG) + table1–4_*.csv.
Missing: fig2_coefficient_plot.pdf and fig3_similarity_distributions.pdf,
although both papers reference them (sections/results.tex, sections/data.tex).
wp2/main_web.tex silently masks this with a custom \includegraphics
fallback that prints a "Figure indisponible" box — the compiled
wp2/main_web.pdf therefore ships with two placeholder boxes instead of
Figures 2 (semantic coefficient forest plot) and 3 (similarity distributions).
Both are regenerated by paper_figures_v2.py. This is the main defect fixed
by the restructuring.
Figure → script map (all from paper_figures_v2.py):
fig1 model comparison; fig2 semantic coefficient forest plot (was missing);
fig3 similarity boxplots (was missing); fig4 price-quintile heatmap;
fig5 similarity correlation matrix; fig6 scatter plots (4 dims);
fig7 methodology diagram; fig8 R² decomposition; fig9 price distributions;
fig10 structural coefficients; fig11 residual diagnostics.
5. LaTeX
Two generations coexist:
- Root (obsolete, May 20): monolithic
paper.tex(112 KB, first draft), thenmain.tex+preamble.tex+titlepage.tex+sections/+tables/appendix/+references.bib.
wp2/(canonical, May 20 – June 13):main.tex(self-contained preamble, UQO WP style),main_web.tex(same + figure-fallback hack, source of the compiledmain_web.pdf),sections/(9 files incl. titlepage),tables/(8 hand-written booktabs tables),appendix/appendix.tex,references.bib(551 lines, ~40 entries),Makefile,uq_logo.jpg.
diff -r confirms root sections/, tables/, appendix/, references.bib
are byte-identical to the wp2/ ones (root just lacks titlepage.tex).
So the root LaTeX tree is a pure duplicate → dropped.
Numbers in tables/*.tex were hand-transcribed from hedonic_maison.py
console output (no automated table generation).
6. Dead / duplicate / unused files
paper.tex, rootmain.tex,preamble.tex,titlepage.tex, rootsections/,tables/,appendix/, rootfigures/— superseded duplicates.- Build artifacts:
*.aux, *.log, *.out, *.toc, *.bbl, *.blg, *.fls, *.fdb_latexmk, *.synctex.gzat root and inwp2/. .DS_Storefiles.- Earlier-approach outputs (
hedonic_results.csv,hedonic_cosine_results.csv,embeddings_remarks.npy,similarity_matrix.npy,reference_embeddings.npz) are kept indata/processed/legacy/for traceability but are not part of the pipeline.
7. Discrepancies & reproduction status
- fig2/fig3 missing in the original (see §4) — regenerated here.
robustness_results.csvlost in the original — regenerated here.- Reproduction check (new pipeline vs original artifacts) — results recorded
in
CHANGES.md§Verification:- similarity features vs
hedonic_maison_results.csv(17,087 × 20), - model A–E statistics vs
tables/tab_model_comparison.tex(R² 0.4525 / 0.4647 / 0.4962 / 0.5122 / 0.5118), - regenerated figures vs shipped fig1, fig4–fig11.
- similarity features vs
- Bootstrap (
np.random.seed(42)) and Lasso/ENet CV (random_state=42) are seeded, so robustness numbers are reproducible on the same package versions.