# METHODOLOGY AUDIT — reconstruction exacte du modèle actuel (§100.3) Baseline figé : commit `e632529`, modèle 3.0.0, 29 tests verts, forecast et backtest reproduits localement le 2026-08-30 (graines fixes). ## Chaîne actuelle (observé → publié) 1. **Observation** : sondage Wikipédia (sections nationales) → renormalisation → correction house effect (biais final 21 j, shrink 0,5, borné ±3 pp) → alr → covariance multinomiale (n_eff = n/deff par mode, plafond 2500) + erreur excédentaire diag (±1,5 pp; ×1,25 maison inconnue). Partielles : pseudo-observations nationales (swing local ×0,6, n≈350, ±3 pp). 2. **État latent** : marche aléatoire alr 5D, q estimé par MV (grille 12 pts, borné), Kalman exact + RTS. 3. **Forecast** : moyenne martingale; variance += q·jours·1,6·m_volatilité (m ∈ [1;1,5], attention wiki+social) + 0,12² systémique. 4. **Fondamentaux** : ridge (vote sortant ~ vote préc. + satisfaction) sur 8 élections, pénalité 1,5 pp/mandat; blend précision w(t)=min(0,8,(j/540)^1,3). 5. **Médias** : nudge borné ±0,35 pp (tanh sentiment 14 j, seuil volume). 6. **Circonscriptions** : alr(base 2022 locale) + swing national alr + ε_région (0,14) + ε_circ (0,20) − retrait sortant (0,06); softmax. 7. **Simulation** : 25 000 tirages corrélés (état→national→régional→local); sièges, P(1er), P(maj), pivots (8k). 8. **Publication** : P(1er) = 0,88·modèle + 0,12·Polymarket. ## Écarts vs la cible « ultra avancée » (§ du cahier des charges) | § | Cible | État actuel | Écart | |---|---|---|---| | 2 | état latent joint (démo × région) | national seulement + couches d'erreur | majeur (Phases C/D) | | 3-4 | RW challengé (LLT, mean-rev, Student-t) | RW gaussien | à tester par replay | | 6-10 | modèle de mesure complet (mode/population/LV/cov inter-partis) | house+mode deff+excès | partiel → `national/pollster_error.py` | | 11 | terrain = moyenne sur la période | assigné à la fin de terrain | mineur, plan replay | | 14-16 | MRP dynamique | absent (plan POPULATION_FEATURES) | majeur (Phase D) | | 17-19 | participation probabiliste | scénario manuel (simulateur) | majeur (Phase E) | | 20-24 | hiérarchie/facteurs spatiaux appris | régions administratives + ε indépendants | majeur (Phase C) | | 27-29 | effets candidats estimés | pénalité retrait fixe | Phase E | | 30 | modèle de partielle décomposé | shrink fixe 0,6 | amélioration ciblée | | 31-33 | fondamentaux bayésiens LOEO | ridge LOO, fuite en replay (corrigée) | ✔ corrigé | | 34-36 | marchés = observation d'événement, poids appris | blend fixe 12 % | dette assumée (pas d'historique) | | 37-45 | social = attention/volatilité | ✔ conforme (bornée, jamais directionnelle) | ✔ | | 46-48 | distribution prédictive du prochain sondage | absent | backlog court terme | | 49-51 | stacking / champion-challenger | ablation implémentée; stacking à venir | Phase F | | 52-58 | replay multi-élections + scoring propre | 2022 seulement → **implémenté : 2007-2022 (6 élections), CRPS/log score** | ✔ ce commit | | 59-61 | hyperparamètres estimés | voir STATISTICAL_DEBT.md | en cours (σ_industrie empirique) | | 62-66 | simulation copules/queues | gaussien corrélé | à tester | | 89-93 | modèle du soir d'élection | absent | Phase G (avant le 5 oct.) |