Modèle 1.1.0 : recensement 2021 (StatCan ADA → 127 circonscriptions), swing par segments langue/région, calibration nationale, sondages de référence 2022
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
15 changed files +794 −23
added
engine/data/derived/riding_demographics_2021.SOURCE.txt
+3 −0
@@ -0,0 +1,3 @@ | ||
| 1 | +Statistique Canada, Recensement de la population 2021, Profil du recensement (98-401-X2021012, ADA) + limites des ADA 2021 ; agrégation surfacique QC26 vers la carte électorale 2026 | |
| 2 | +computed_at=2026-09-06T05:33:24.708846+00:00 | |
| 3 | +ADA=1175 | |
added
engine/data/derived/riding_demographics_2021.csv
+128 −0
@@ -0,0 +1,128 @@ | ||
| 1 | +riding_code,population_2021,density,median_age,pct_65_plus,pct_french_mt,pct_english_mt,pct_other_mt,pct_plop_french,pct_immigrants,pct_visible_minority,median_household_income,pct_owner,pct_bachelor_plus,unemployment_rate,ada_count | |
| 2 | +117,59877,65.2,50.1,26.9,92.97,4.84,1.06,94.27,1.85,1.39,61100.7,74.39,11.56,6.5,20 | |
| 3 | +119,69357,1174.5,42.6,22.74,84.32,7.49,5.79,88.64,7.17,7.83,62714.8,52.26,22.16,5.7,16 | |
| 4 | +121,66763,1913.5,42.2,24.17,83.35,3.37,10.55,91.63,11.76,13.03,55492.0,35.82,37.79,7.8,18 | |
| 5 | +127,54813,373.9,54.3,29.89,83.05,12.62,2.01,85.04,5.08,1.79,72900.0,73.63,27.5,6.7,18 | |
| 6 | +131,56240,50.3,43.8,19.71,94.34,3.06,1.35,95.99,2.41,1.52,73336.9,75.63,13.24,5.3,28 | |
| 7 | +137,80331,491.5,41.5,18.12,91.74,3.75,3.01,94.55,4.29,3.88,83355.4,74.05,26.58,4.4,25 | |
| 8 | +139,61070,918.9,41.9,18.8,94.26,1.18,3.45,97.48,3.75,4.57,69336.2,62.35,15.27,5.4,21 | |
| 9 | +141,62826,630.0,47.1,24.97,95.27,1.23,2.37,97.72,3.28,3.42,66088.8,62.12,16.36,5.3,31 | |
| 10 | +147,77182,880.6,47.3,25.77,96.0,0.75,2.51,98.24,2.55,3.19,64761.0,65.22,15.86,4.4,23 | |
| 11 | +151,49305,35.7,47.5,24.71,97.24,0.81,1.3,98.62,1.63,1.18,66585.7,75.14,16.2,4.8,20 | |
| 12 | +157,64821,120.1,49.4,26.33,79.17,15.53,2.34,81.27,4.75,1.88,72814.2,68.66,24.27,5.4,15 | |
| 13 | +159,72314,1061.7,47.6,25.15,91.86,2.18,4.35,95.99,5.16,5.27,66039.2,55.52,16.31,5.5,14 | |
| 14 | +161,64739,487.4,42.6,19.23,93.77,2.65,2.24,95.9,3.47,2.53,78527.2,71.8,14.95,5.1,23 | |
| 15 | +167,76657,1650.5,44.8,21.65,91.93,2.65,3.37,95.36,4.73,4.44,80892.9,58.5,19.49,6.0,18 | |
| 16 | +171,55823,29.5,44.9,20.73,80.34,14.46,2.37,82.49,3.63,1.86,74679.0,77.75,11.41,5.7,17 | |
| 17 | +179,68337,744.5,45.5,22.94,91.7,3.93,2.15,94.02,3.23,3.19,69240.8,60.41,13.41,6.2,12 | |
| 18 | +181,76066,242.3,42.7,16.96,65.12,23.14,7.53,67.74,9.27,7.81,100223.6,84.04,25.33,5.8,13 | |
| 19 | +187,76138,1893.2,42.2,17.06,49.8,24.73,18.87,55.39,20.26,22.17,92829.4,76.61,30.89,7.0,12 | |
| 20 | +191,71424,1517.6,42.0,18.26,62.01,19.85,13.44,69.63,17.72,19.53,84979.4,74.24,21.77,7.3,14 | |
| 21 | +197,57004,1855.7,41.4,16.13,82.04,4.65,10.2,88.43,10.87,11.35,93281.3,74.47,21.93,5.5,11 | |
| 22 | +199,64791,1234.3,41.9,16.31,75.02,6.42,14.69,82.16,15.14,15.97,105402.6,75.79,40.16,5.6,17 | |
| 23 | +201,80349,2807.4,43.8,21.29,39.2,11.45,42.22,53.06,40.72,50.45,87852.6,71.35,46.89,9.5,15 | |
| 24 | +207,70831,1181.2,41.0,15.1,87.61,3.78,6.11,92.26,8.22,8.2,103824.0,78.97,31.2,4.6,20 | |
| 25 | +211,74747,1986.9,42.2,18.23,70.9,6.11,18.53,82.62,21.24,24.8,90027.7,76.5,30.35,6.5,18 | |
| 26 | +217,73974,3719.1,43.7,21.86,60.84,12.89,20.7,72.01,23.68,25.89,74918.9,52.78,38.3,8.2,18 | |
| 27 | +219,64727,4757.8,40.5,18.43,74.4,3.52,17.59,87.32,21.36,27.13,60149.4,33.78,30.26,9.1,15 | |
| 28 | +221,72574,3682.3,43.1,20.54,76.96,2.94,16.09,89.08,19.77,23.17,76799.9,50.81,36.43,7.6,16 | |
| 29 | +227,66775,831.7,47.4,24.22,85.66,4.57,7.06,90.71,9.75,8.15,116165.7,84.73,52.41,5.1,15 | |
| 30 | +231,69700,599.9,42.2,16.21,93.14,1.59,3.54,96.69,5.23,5.31,102135.3,79.39,27.07,4.7,14 | |
| 31 | +237,71701,912.8,44.1,20.85,92.57,2.74,2.86,95.43,4.89,3.45,96153.9,77.59,29.1,5.0,19 | |
| 32 | +239,71549,1229.9,45.8,24.14,92.16,0.89,5.58,97.32,6.43,8.05,67928.8,53.48,17.1,6.0,20 | |
| 33 | +241,65121,778.3,49.0,26.25,96.15,1.07,1.71,98.19,2.82,2.74,69471.0,67.74,13.16,7.1,21 | |
| 34 | +247,69406,8526.7,40.5,17.64,56.01,15.86,22.42,64.21,24.42,24.27,73646.3,37.35,52.55,9.5,13 | |
| 35 | +251,80452,5953.6,41.8,20.2,31.72,25.35,35.81,42.71,35.54,45.14,64183.9,39.19,32.38,11.1,14 | |
| 36 | +257,64922,3675.4,43.0,19.23,41.98,28.86,22.5,49.29,24.84,30.54,70287.8,47.29,37.03,9.2,18 | |
| 37 | +259,62063,1684.8,47.7,24.65,21.3,47.68,24.51,24.17,25.52,22.5,110007.3,75.52,55.6,8.2,16 | |
| 38 | +261,80169,2359.1,44.1,18.1,24.78,31.09,35.33,31.67,35.66,38.83,100494.9,76.09,44.79,8.9,17 | |
| 39 | +267,75174,3409.1,42.7,19.73,18.8,33.3,38.54,26.95,41.27,45.94,92134.2,71.64,42.45,10.1,16 | |
| 40 | +271,94189,4642.8,39.4,16.13,28.61,14.63,47.31,45.71,48.31,56.59,79582.8,49.86,48.46,11.0,25 | |
| 41 | +277,84880,9328.6,40.9,20.77,19.94,32.17,38.31,27.56,45.72,44.55,74444.0,35.11,53.57,11.6,24 | |
| 42 | +279,72786,8575.2,40.1,17.38,23.6,36.48,31.85,30.81,34.87,38.1,70685.7,37.12,57.55,10.6,18 | |
| 43 | +281,86296,7756.5,36.4,13.18,48.96,20.09,23.97,58.26,26.23,31.84,65631.0,34.69,50.82,10.1,20 | |
| 44 | +287,62926,10583.9,35.6,13.92,60.88,12.16,21.15,70.64,26.08,27.61,57828.8,27.56,58.96,11.0,18 | |
| 45 | +291,87978,12918.4,36.5,16.83,26.29,27.57,37.64,32.44,32.56,43.38,69579.2,29.91,72.49,12.6,21 | |
| 46 | +297,104357,8131.6,36.8,16.56,37.89,16.09,37.32,50.29,38.06,41.15,81779.1,34.41,62.73,11.2,31 | |
| 47 | +299,82425,7453.0,40.9,18.94,32.08,8.09,50.43,54.58,51.27,53.54,64715.6,38.21,43.22,12.2,19 | |
| 48 | +301,65691,6755.3,41.9,20.15,63.87,3.69,26.7,83.79,28.28,32.48,63619.9,36.03,45.89,8.7,15 | |
| 49 | +307,74423,14226.6,36.1,13.14,47.26,8.06,37.49,60.18,31.93,40.12,58569.0,27.09,47.73,10.8,19 | |
| 50 | +311,61347,10884.1,36.4,11.44,75.11,6.57,13.88,84.77,18.3,15.28,62888.9,30.43,57.43,8.0,17 | |
| 51 | +313,67818,12928.7,34.3,10.54,63.14,13.84,17.62,71.13,22.41,15.99,66470.4,31.21,65.07,10.1,17 | |
| 52 | +317,58580,9224.1,36.1,11.97,76.37,4.67,14.69,86.58,19.02,22.07,58551.3,28.53,44.33,9.2,14 | |
| 53 | +319,77301,9292.6,41.9,19.39,69.38,3.61,22.11,85.88,24.21,25.18,59665.8,32.03,45.06,8.7,17 | |
| 54 | +321,69941,8917.2,37.5,14.32,40.44,4.12,47.01,73.39,44.12,59.6,55238.7,28.47,27.65,13.0,19 | |
| 55 | +327,81233,8389.8,40.6,18.61,46.03,4.72,41.54,78.66,39.76,58.23,57553.5,28.96,17.31,11.6,17 | |
| 56 | +331,79197,6764.5,42.4,20.82,29.88,8.27,52.61,62.66,46.94,47.1,66517.3,34.92,29.57,10.9,18 | |
| 57 | +337,74680,6241.9,41.3,18.37,71.68,3.2,20.42,87.53,22.71,28.17,64559.2,38.65,31.26,9.8,15 | |
| 58 | +339,57720,3426.1,44.0,19.25,39.03,11.82,41.62,61.89,34.87,38.32,83951.0,66.87,22.01,8.1,15 | |
| 59 | +341,54216,3515.4,44.8,19.19,79.32,3.12,13.78,91.43,16.46,22.16,70618.5,54.73,18.8,7.9,16 | |
| 60 | +347,77402,4325.3,42.4,21.21,57.27,5.06,31.19,76.66,33.11,37.42,69575.3,44.09,32.39,9.4,14 | |
| 61 | +351,73429,4629.2,45.8,24.47,32.57,9.98,48.08,50.0,45.02,39.06,68166.4,52.34,28.87,12.7,13 | |
| 62 | +357,81641,2545.4,41.9,14.83,46.59,10.61,35.2,60.53,31.73,31.61,101326.0,79.15,33.9,8.6,13 | |
| 63 | +359,78797,2585.4,42.2,15.92,59.98,6.18,27.53,75.71,27.82,30.11,99449.0,76.67,35.16,7.4,12 | |
| 64 | +569,57409,2473.8,44.0,17.72,58.49,8.28,27.11,75.25,23.56,26.49,96527.3,77.84,30.75,6.8,12 | |
| 65 | +571,65396,2004.2,43.3,18.08,59.33,7.1,27.55,77.1,26.1,30.3,97522.7,77.44,31.14,6.9,14 | |
| 66 | +577,67073,1910.2,44.1,20.75,77.92,5.82,12.6,84.52,10.51,9.51,86600.5,60.73,27.74,7.2,15 | |
| 67 | +581,63017,2480.1,44.3,20.17,81.06,6.37,9.08,87.54,10.75,11.13,83249.7,69.91,22.06,6.2,15 | |
| 68 | +587,72233,836.7,38.9,13.07,88.14,3.49,5.88,92.77,7.36,7.61,93699.4,76.95,20.25,5.2,16 | |
| 69 | +589,47868,70.2,52.5,27.05,81.3,13.35,2.53,83.42,4.46,1.53,64921.0,73.34,14.92,8.9,16 | |
| 70 | +591,55062,685.1,39.5,14.32,92.68,2.16,3.27,95.85,5.11,4.96,88038.1,71.87,17.24,5.6,12 | |
| 71 | +597,51876,1695.3,48.0,26.59,92.29,1.51,4.27,96.53,5.46,6.09,57604.0,43.68,12.96,8.1,10 | |
| 72 | +601,63562,867.3,38.8,12.89,88.51,2.13,6.61,95.19,8.83,13.09,88363.9,72.9,13.17,6.3,21 | |
| 73 | +607,79090,1932.2,42.1,15.1,81.73,4.46,10.29,89.07,12.02,12.39,111062.0,79.51,32.42,5.9,17 | |
| 74 | +609,66645,1590.4,42.5,17.47,84.51,2.62,9.6,93.26,11.64,15.95,95655.6,72.33,23.06,5.9,15 | |
| 75 | +611,70409,1146.3,41.3,16.4,86.98,2.54,7.76,93.9,9.47,12.04,97504.2,77.85,22.23,5.2,16 | |
| 76 | +617,65430,2784.0,45.7,21.74,85.81,1.68,9.62,95.16,12.13,16.89,93212.2,74.66,22.57,6.1,14 | |
| 77 | +621,72908,1138.2,41.4,16.21,92.16,1.23,4.84,97.31,6.96,9.68,88060.2,75.41,16.47,5.7,23 | |
| 78 | +627,59251,40.6,51.6,26.38,93.22,1.17,4.58,97.47,2.0,1.44,62206.0,74.36,11.41,7.9,19 | |
| 79 | +629,71156,563.7,47.7,26.46,95.23,0.96,2.61,98.17,3.46,3.64,65642.4,55.8,16.83,6.8,15 | |
| 80 | +631,58649,125.6,39.9,15.69,94.97,1.6,2.0,97.36,2.94,2.7,74790.8,75.25,7.96,6.6,18 | |
| 81 | +637,61482,224.0,46.3,20.51,91.85,3.78,2.52,94.27,4.39,1.86,84292.0,79.02,20.53,7.1,14 | |
| 82 | +641,67015,64.4,53.8,27.66,90.32,5.11,2.59,92.69,4.98,1.93,63954.6,72.88,19.23,9.4,16 | |
| 83 | +649,61835,17.1,54.2,29.34,93.45,3.85,1.43,95.01,2.62,1.29,60029.1,72.33,14.85,10.5,14 | |
| 84 | +651,79806,2179.5,40.1,17.21,65.29,12.27,16.21,74.62,21.28,26.77,77872.2,46.22,45.9,10.7,14 | |
| 85 | +657,85981,1173.0,40.8,15.46,52.82,28.96,12.07,58.67,15.59,18.91,93247.2,71.78,37.32,8.5,13 | |
| 86 | +661,76752,143.2,45.4,18.09,75.16,17.36,3.89,77.79,5.37,4.4,101914.6,84.65,28.34,7.3,26 | |
| 87 | +667,78741,2277.8,43.8,19.04,79.2,6.75,9.59,86.42,12.13,14.56,79381.2,62.49,23.18,10.0,15 | |
| 88 | +669,69997,620.9,44.3,19.24,88.97,6.18,2.29,91.67,3.76,4.42,79539.8,72.91,14.92,8.2,20 | |
| 89 | +671,48909,13.8,44.2,20.92,90.97,6.11,1.45,92.7,1.91,2.08,78438.1,71.01,19.17,5.3,12 | |
| 90 | +681,45353,19.9,45.3,22.03,97.41,0.82,1.02,98.67,0.94,1.13,72304.9,72.6,12.88,6.1,14 | |
| 91 | +689,52770,219.5,43.1,19.87,93.03,2.47,2.76,96.06,2.01,3.17,75197.3,58.06,16.34,5.5,13 | |
| 92 | +691,64320,1923.1,47.6,28.39,91.64,1.22,5.36,96.82,6.0,9.07,57689.0,43.2,28.45,7.9,14 | |
| 93 | +693,63232,144.6,48.3,24.49,96.66,0.96,1.6,98.33,2.2,2.25,71852.9,74.83,19.74,6.0,17 | |
| 94 | +697,76888,378.6,52.9,29.55,92.42,1.12,4.98,97.94,1.72,1.6,57139.8,61.24,13.65,7.5,22 | |
| 95 | +701,68480,962.8,49.1,25.96,96.53,1.04,1.54,98.36,2.68,2.71,65751.2,65.96,17.73,6.1,20 | |
| 96 | +707,69150,2661.6,40.7,24.76,82.07,2.21,12.77,92.63,14.67,18.44,79992.6,43.21,60.05,8.1,15 | |
| 97 | +709,61599,1647.8,48.1,24.87,91.01,1.54,5.64,96.28,7.45,7.58,99679.2,75.51,50.22,5.6,14 | |
| 98 | +727,66308,77.8,45.1,21.87,97.05,1.14,0.9,98.26,1.63,1.38,78346.8,79.2,17.2,5.2,15 | |
| 99 | +729,74046,1158.9,40.6,17.12,93.42,1.89,3.11,96.91,4.86,5.15,91563.9,68.71,26.5,5.1,21 | |
| 100 | +749,68146,2233.8,47.2,25.49,90.95,1.02,6.3,96.77,8.1,9.71,74462.8,50.04,32.81,7.2,22 | |
| 101 | +751,63140,7250.7,45.6,25.85,88.29,2.36,7.14,94.51,10.14,9.77,56719.2,30.02,51.11,9.4,15 | |
| 102 | +757,58658,4322.3,43.3,23.45,89.54,1.24,7.06,96.32,10.84,12.56,50073.8,25.04,29.21,9.1,16 | |
| 103 | +761,75350,1668.9,44.0,19.16,95.14,0.79,2.86,98.3,5.08,5.23,86404.1,69.31,24.39,5.7,17 | |
| 104 | +767,67951,2478.3,46.5,25.53,92.25,1.03,5.09,97.48,8.05,8.37,73128.5,57.6,27.81,7.0,15 | |
| 105 | +769,85039,992.9,41.4,16.02,95.07,1.33,2.34,97.77,4.25,3.57,94543.0,76.62,26.76,5.8,22 | |
| 106 | +771,65739,67.0,50.4,26.6,97.46,0.88,0.98,98.74,2.24,1.48,78844.5,76.91,21.24,7.2,15 | |
| 107 | +777,64329,125.0,45.2,22.43,97.01,0.85,1.56,98.46,1.47,2.27,65958.5,70.34,13.47,5.5,17 | |
| 108 | +781,57258,53.5,42.4,20.66,97.35,0.72,1.33,98.68,1.5,1.87,77869.8,77.56,16.32,4.3,18 | |
| 109 | +787,73978,58.0,47.7,26.01,96.82,0.97,1.32,98.34,1.57,2.09,66333.0,74.66,12.98,5.4,21 | |
| 110 | +789,77526,770.7,42.0,17.24,96.3,1.03,1.71,98.28,3.05,3.12,99297.2,76.89,33.18,4.6,17 | |
| 111 | +791,56340,1585.1,48.8,27.05,94.93,1.0,2.82,97.96,4.36,5.01,71176.2,54.2,29.71,6.6,12 | |
| 112 | +797,63024,103.5,46.3,23.36,97.12,0.76,1.46,98.64,2.17,2.67,74650.3,74.56,16.94,5.3,19 | |
| 113 | +799,61225,30.6,52.7,29.43,98.33,0.53,0.74,99.09,1.16,1.23,61381.0,73.56,13.76,6.7,14 | |
| 114 | +801,63599,234.8,50.7,28.18,98.2,0.56,0.72,99.08,1.24,1.45,62381.9,69.96,15.8,6.9,12 | |
| 115 | +803,56729,981.6,47.9,25.68,97.39,0.69,1.22,98.93,2.15,2.22,69830.1,64.64,27.74,7.9,19 | |
| 116 | +807,56851,44.8,52.3,28.27,98.48,0.62,0.34,99.15,0.81,0.83,56902.8,71.72,12.79,8.5,18 | |
| 117 | +809,35672,10.9,54.0,29.62,91.52,6.51,0.75,92.63,1.87,1.37,62587.2,72.98,14.3,11.5,9 | |
| 118 | +811,12251,67.5,54.4,28.37,93.98,4.98,0.28,94.54,0.76,0.48,75500.0,75.56,16.99,12.2,1 | |
| 119 | +813,47903,404.4,43.6,18.97,78.37,7.95,11.23,88.32,1.7,1.8,87018.5,65.74,14.07,7.6,18 | |
| 120 | +817,40192,249.6,49.8,24.17,92.87,0.54,5.92,98.63,1.07,1.17,70259.7,71.78,13.92,7.9,10 | |
| 121 | +819,50216,77.9,45.7,22.21,98.11,0.88,0.42,98.75,0.77,0.74,79436.0,76.72,15.73,5.9,26 | |
| 122 | +821,53423,1169.6,47.4,26.73,96.63,0.76,1.84,98.59,1.93,3.14,66411.0,56.57,27.05,5.9,10 | |
| 123 | +823,53208,1063.5,46.4,24.4,98.01,0.74,0.66,99.01,1.11,1.51,70752.6,62.37,19.66,5.6,11 | |
| 124 | +827,50728,214.5,47.4,24.38,98.71,0.41,0.54,99.48,0.76,1.12,70821.8,72.74,14.77,5.4,18 | |
| 125 | +829,47203,29.0,50.3,26.69,98.49,0.35,0.72,99.54,0.89,1.04,64957.0,72.87,14.21,7.0,12 | |
| 126 | +831,45781,14.8,30.7,9.12,29.95,6.65,59.49,31.14,1.21,1.81,98836.5,35.62,9.62,7.6,33 | |
| 127 | +833,62053,5390.7,44.4,22.51,54.55,4.28,34.54,80.4,36.0,37.74,67848.6,41.03,34.94,9.4,19 | |
| 128 | +837,40585,14.8,53.8,30.31,87.45,10.24,0.85,88.58,1.14,0.63,61573.9,72.29,15.0,11.2,10 | |
modified
engine/qc26/api/public.py
+23 −2
@@ -19,10 +19,20 @@ from ..connectors.parties.documents import TOPIC_LABELS, TOPICS | ||
| 19 | 19 | from ..db import (AuditLog, CampaignEvent, Candidate, CompassQuestion, DocumentVersion, ElectionEvent, ForecastHouseEffect, |
| 20 | 20 | ForecastPartyResult, ForecastRidingResult, ForecastRidingSummary, ForecastRun, HistoricalResult, IngestionJob, |
| 21 | 21 | LiveRidingState, LiveSnapshot, NowcastRun, Party, PartyDocument, PartyPosition, Poll, PollWeight, Pollster, |
| 22 | − PolicyProposal, Region, Riding, RidingBaseline, SessionLocal, Source, get_setting, db_latency_ms) | |
| 22 | + PolicyProposal, Region, Riding, RidingBaseline, RidingDemographics, SessionLocal, Source, get_setting, db_latency_ms) | |
| 23 | + | |
| 24 | + | |
| 25 | +def _demo_payload(d: RidingDemographics | None) -> dict | None: | |
| 26 | + if not d: | |
| 27 | + return None | |
| 28 | + return {"population2021": d.population_2021, "density": d.density, "medianAge": d.median_age, "pct65Plus": d.pct_65_plus, | |
| 29 | + "pctFrenchMt": d.pct_french_mt, "pctEnglishMt": d.pct_english_mt, "pctOtherMt": d.pct_other_mt, "pctPlopFrench": d.pct_plop_french, | |
| 30 | + "pctImmigrants": d.pct_immigrants, "pctVisibleMinority": d.pct_visible_minority, "medianHouseholdIncome": d.median_household_income, | |
| 31 | + "pctOwner": d.pct_owner, "pctBachelorPlus": d.pct_bachelor_plus, "unemploymentRate": d.unemployment_rate, "adaCount": d.ada_count, | |
| 32 | + "source": d.source, "computedAt": d.computed_at.isoformat() if d.computed_at else None} | |
| 23 | 33 | from ..live.feed import broadcaster |
| 24 | 34 | from ..parties import OTHERS_COLOR, OTHERS_LABEL, PARTY_IDS, registry_payload |
| 25 | −from ..regions import POLL_REGION_LABELS | |
| 35 | +from ..regions import POLL_REGION_LABELS, poll_region_for | |
| 26 | 36 | from ..compass import positions as compass_positions |
| 27 | 37 | from ..compass.questions import AXES |
| 28 | 38 | |
@@ -208,12 +218,15 @@ def ridings(s: Session = Depends(db)): | ||
| 208 | 218 | pwin[fr.riding_code][fr.party_id] = round(fr.p_win, 3) |
| 209 | 219 | live = {st.riding_code: st for st in s.scalars(select(LiveRidingState)).all()} |
| 210 | 220 | regions = {r.id: r for r in s.scalars(select(Region)).all()} |
| 221 | + demo = {d.riding_code: d for d in s.scalars(select(RidingDemographics)).all()} | |
| 211 | 222 | out = [] |
| 212 | 223 | for r in s.scalars(select(Riding).order_by(Riding.sort_key)).all(): |
| 213 | 224 | x = summ.get(r.code) |
| 214 | 225 | st = live.get(r.code) |
| 226 | + dm = demo.get(r.code) | |
| 215 | 227 | out.append({"code": r.code, "name": r.name, "slug": r.slug, "region": r.region_id, "regionName": regions[r.region_id].name if r.region_id in regions else None, |
| 216 | 228 | "electors": r.electors_2026, "isNew": r.is_new_2026, "changed": r.boundary_changed_2026, |
| 229 | + "pctFrench": dm.pct_french_mt if dm else None, "pctImmigrants": dm.pct_immigrants if dm else None, "medianIncome": dm.median_household_income if dm else None, | |
| 217 | 230 | "incumbent": {"party": r.incumbent_party_id, "name": r.incumbent_name, "running": r.incumbent_running}, |
| 218 | 231 | "centroid": [r.centroid_lon, r.centroid_lat], |
| 219 | 232 | "forecast": {"favorite": x.favorite_party_id, "p": round(x.favorite_p, 3), "runnerUp": x.runner_up_party_id, "margin": round(x.margin_mean, 1), |
@@ -249,6 +262,8 @@ def riding_detail(slug: str, s: Session = Depends(db)): | ||
| 249 | 262 | "elected": h.elected, "turnout": h.turnout}) |
| 250 | 263 | d["history"] = dict(hist) |
| 251 | 264 | d["historyNote"] = "Résultats officiels sur les limites de l'époque (carte 2017) ; la ligne « 2022 transposé » applique les limites 2026." if hist else "Circonscription nouvelle ou renommée : pas d'historique sous ce nom." |
| 265 | + d["demographics"] = _demo_payload(s.get(RidingDemographics, r.code)) | |
| 266 | + d["pollRegion"] = poll_region_for(r.code) | |
| 252 | 267 | if run: |
| 253 | 268 | x = s.scalar(select(ForecastRidingSummary).where(ForecastRidingSummary.run_id == run.id, ForecastRidingSummary.riding_code == r.code)) |
| 254 | 269 | res = s.scalars(select(ForecastRidingResult).where(ForecastRidingResult.run_id == run.id, ForecastRidingResult.riding_code == r.code).order_by(desc(ForecastRidingResult.p_win))).all() |
@@ -673,6 +688,12 @@ def status(s: Session = Depends(db)): | ||
| 673 | 688 | |
| 674 | 689 | |
| 675 | 690 | CHANGELOG = [ |
| 691 | + {"version": "1.1.0", "date": "2026-09-06", "title": "Recensement 2021 et swing par segments", | |
| 692 | + "notes": ["Recensement de la population 2021 (Statistique Canada, profil des aires de diffusion agrégées + limites) agrégé par intersection surfacique aux 127 circonscriptions : langue maternelle, première langue officielle parlée, immigration, minorités visibles, âge, revenu, tenure, scolarité, chômage.", | |
| 693 | + "Excès de swing par langue : ventilations francophones / non-francophones des sondages (90 jours, pondérées récence × √n) comparées à une référence 2022 par régression écologique des résultats sur la part francophone ; appliqué à chaque circonscription selon sa composition (κ = 0,6).", | |
| 694 | + "Excès de swing par région : ventilations RMR de Montréal / RMR de Québec / reste du Québec comparées aux résultats 2022 agrégés par le même regroupement (κ = 0,5) ; regroupement affiné circonscription par circonscription.", | |
| 695 | + "Calibration : la moyenne des parts par circonscription pondérée par les électeurs inscrits est ramenée au vote national simulé à chaque tirage.", | |
| 696 | + "« Pourquoi ? » : facteurs linguistique et régional affichés avec les valeurs réellement utilisées."]}, | |
| 676 | 697 | {"version": "1.0.0", "date": "2026-09-06", "title": "Lancement de QC26", |
| 677 | 698 | "notes": ["Agrégation des sondages depuis novembre 2022 : pondération récence × taille × qualité, effets maison (leave-one-pollster-out, rétrécis, bornés à ±3 pt).", |
| 678 | 699 | "État latent du vote par filtre de Kalman (marche aléatoire, espace logit), lissage RTS, dérive de campagne jusqu'au scrutin.", |
modified
engine/qc26/config.py
+1 −1
@@ -52,7 +52,7 @@ class Settings(BaseSettings): | ||
| 52 | 52 | # ils dérivent de la liste officielle des circonscriptions (127 en 2026 → 64). |
| 53 | 53 | |
| 54 | 54 | # --- Modèle --- |
| 55 | − model_version: str = "1.0.0" | |
| 55 | + model_version: str = "1.1.0" | |
| 56 | 56 | n_simulations: int = 20000 |
| 57 | 57 | random_seed: int | None = None |
| 58 | 58 | recency_halflife_days: float = 14.0 # hors campagne |
modified
engine/qc26/connectors/polls/registry.py
+27 −0
@@ -219,6 +219,33 @@ def verify_bootstrap_documents(s: Session, limit: int = 20, only_ids: list[str] | ||
| 219 | 219 | return {"verified": done, "failed": failed, "skipped": skipped, "remaining": max(0, len(polls) - done - failed - skipped)} |
| 220 | 220 | |
| 221 | 221 | |
| 222 | +# Sondages de référence de fin de campagne 2022 (ventilations linguistiques) : servent de base | |
| 223 | +# 2022 par segment au modèle (model/segments.py). Documents originaux de la firme. | |
| 224 | +REFERENCE_POLLS_2022 = [ | |
| 225 | + ("leger", "https://leger360.com/wp-content/uploads/2024/02/20221002-Rapport-fin-de-campagne-Election-Quebec-2022.pdf"), | |
| 226 | + ("leger", "https://leger360.com/wp-content/uploads/2024/02/20220927-Rapport-Intentions-de-vote-post-debat-Radio-Canada-Election-Quebec-2022.pdf"), | |
| 227 | +] | |
| 228 | + | |
| 229 | + | |
| 230 | +def ensure_reference_polls(s: Session) -> int: | |
| 231 | + """Ingère (une seule fois) les sondages de référence 2022 s'ils sont absents de la base.""" | |
| 232 | + known = set(s.scalars(select(PollDocument.url)).all()) | |
| 233 | + n = 0 | |
| 234 | + for pid, url in REFERENCE_POLLS_2022: | |
| 235 | + if url in known: | |
| 236 | + continue | |
| 237 | + try: | |
| 238 | + poll = connector_for(pid).process(s, DocumentCandidate(url=url, pollster_id=pid)) | |
| 239 | + s.commit() | |
| 240 | + if poll: | |
| 241 | + n += 1 | |
| 242 | + log.info("sondage de référence 2022 ingéré : %s", poll.id) | |
| 243 | + except Exception as e: # noqa: BLE001 | |
| 244 | + log.warning("sondage de référence %s : %s", url, e) | |
| 245 | + s.rollback() | |
| 246 | + return n | |
| 247 | + | |
| 248 | + | |
| 222 | 249 | def watch_new_polls(s: Session) -> dict: |
| 223 | 250 | """Cycle du Poll Watcher : découverte chez chaque firme, ingestion des documents inconnus.""" |
| 224 | 251 | known_urls = {d.url for d in s.scalars(select(PollDocument.url)).all()} if False else set(s.scalars(select(PollDocument.url)).all()) |
modified
engine/qc26/db.py
+25 −0
@@ -93,6 +93,31 @@ class RidingBaseline(Base): | ||
| 93 | 93 | source: Mapped[str] = mapped_column(String(240)) |
| 94 | 94 | |
| 95 | 95 | |
| 96 | +class RidingDemographics(Base): | |
| 97 | + """Recensement 2021 (Statistique Canada) agrégé aux circonscriptions 2026 par intersection | |
| 98 | + surfacique des aires de diffusion agrégées (ADA). Parts en % de la population concernée.""" | |
| 99 | + __tablename__ = "riding_demographics" | |
| 100 | + riding_code: Mapped[int] = mapped_column(ForeignKey("ridings.code"), primary_key=True) | |
| 101 | + population_2021: Mapped[int] = mapped_column(Integer) | |
| 102 | + density: Mapped[float | None] = mapped_column(Float) | |
| 103 | + median_age: Mapped[float | None] = mapped_column(Float) | |
| 104 | + pct_65_plus: Mapped[float | None] = mapped_column(Float) | |
| 105 | + pct_french_mt: Mapped[float | None] = mapped_column(Float) # langue maternelle française (réponses uniques) | |
| 106 | + pct_english_mt: Mapped[float | None] = mapped_column(Float) | |
| 107 | + pct_other_mt: Mapped[float | None] = mapped_column(Float) | |
| 108 | + pct_plop_french: Mapped[float | None] = mapped_column(Float) # première langue officielle parlée : français | |
| 109 | + pct_immigrants: Mapped[float | None] = mapped_column(Float) | |
| 110 | + pct_visible_minority: Mapped[float | None] = mapped_column(Float) | |
| 111 | + median_household_income: Mapped[float | None] = mapped_column(Float) | |
| 112 | + pct_owner: Mapped[float | None] = mapped_column(Float) | |
| 113 | + pct_bachelor_plus: Mapped[float | None] = mapped_column(Float) # 25-64 ans | |
| 114 | + unemployment_rate: Mapped[float | None] = mapped_column(Float) | |
| 115 | + ada_count: Mapped[int] = mapped_column(Integer, default=0) | |
| 116 | + coverage: Mapped[float | None] = mapped_column(Float) # part de la population de la circ. couverte par des ADA | |
| 117 | + source: Mapped[str] = mapped_column(String(300)) | |
| 118 | + computed_at: Mapped[datetime] = mapped_column(DateTime(timezone=True), default=utcnow) | |
| 119 | + | |
| 120 | + | |
| 96 | 121 | class HistoricalResult(Base): |
| 97 | 122 | """Résultats officiels par circonscription (cartes antérieures).""" |
| 98 | 123 | __tablename__ = "historical_results" |
added
engine/qc26/ingest/census.py
+222 −0
@@ -0,0 +1,222 @@ | ||
| 1 | +"""Recensement de la population 2021 (Statistique Canada) → circonscriptions 2026. | |
| 2 | + | |
| 3 | +Produit : Profil du recensement 2021, aires de diffusion agrégées (98-401-X2021012, CSV) | |
| 4 | +Géométrie : fichier des limites des ADA 2021 (lada000b21a, Lambert Statistique Canada) | |
| 5 | + | |
| 6 | +Méthode : chaque ADA du Québec (1 175) est intersectée avec les 127 circonscriptions 2026 ; | |
| 7 | +les effectifs (population, langue, immigration, tenure…) sont répartis au prorata de la | |
| 8 | +surface intersectée, les indicateurs (âge médian, revenu médian, taux de chômage) sont | |
| 9 | +moyennés pondérés par la population allouée. Le résultat est écrit dans | |
| 10 | +`data/derived/riding_demographics_2021.csv` (versionné) puis chargé en base. | |
| 11 | +""" | |
| 12 | +from __future__ import annotations | |
| 13 | + | |
| 14 | +import csv | |
| 15 | +import io | |
| 16 | +import json | |
| 17 | +import logging | |
| 18 | +import zipfile | |
| 19 | +from collections import defaultdict | |
| 20 | +from pathlib import Path | |
| 21 | + | |
| 22 | +from sqlalchemy import select | |
| 23 | +from sqlalchemy.orm import Session | |
| 24 | + | |
| 25 | +from ..config import settings | |
| 26 | +from ..db import Riding, RidingDemographics, Source, utcnow | |
| 27 | + | |
| 28 | +log = logging.getLogger("qc26.census") | |
| 29 | +DERIVED = settings.data_dir / "derived" | |
| 30 | +OUT_CSV = DERIVED / "riding_demographics_2021.csv" | |
| 31 | +SOURCE = ("Statistique Canada, Recensement de la population 2021, Profil du recensement (98-401-X2021012, ADA) " | |
| 32 | + "+ limites des ADA 2021 ; agrégation surfacique QC26 vers la carte électorale 2026") | |
| 33 | + | |
| 34 | +# identifiants de caractéristiques (fichier français) | |
| 35 | +IDS = { | |
| 36 | + "pop": 1, "density": 6, "age65": 24, "median_age": 40, | |
| 37 | + "mt_total": 393, "mt_en": 396, "mt_fr": 397, "mt_other": 398, | |
| 38 | + "plop_total": 388, "plop_fr": 390, | |
| 39 | + "income_median": 243, | |
| 40 | + "tenure_total": 1414, "owner": 1415, | |
| 41 | + "immig_total": 1527, "immigrants": 1529, | |
| 42 | + "vismin_total": 1683, "vismin": 1684, | |
| 43 | + "educ_total": 2014, "bachelor": 2024, "unemployment": 2230, | |
| 44 | +} | |
| 45 | +BACHELOR_LABEL = "Baccalauréat ou grade supérieur" # id 2024 (25-64 ans) | |
| 46 | + | |
| 47 | + | |
| 48 | +def _num(x: str) -> float | None: | |
| 49 | + x = (x or "").strip() | |
| 50 | + if not x or x in ("...", "..", "x", "F"): | |
| 51 | + return None | |
| 52 | + try: | |
| 53 | + return float(x.replace(",", ".")) | |
| 54 | + except ValueError: | |
| 55 | + return None | |
| 56 | + | |
| 57 | + | |
| 58 | +def read_ada_profile(csv_path: Path) -> dict[str, dict[str, float | None]]: | |
| 59 | + """Lit (en flux) le CSV national et ne garde que les ADA du Québec (DGUID 2021S051624…).""" | |
| 60 | + out: dict[str, dict[str, float | None]] = {} | |
| 61 | + wanted = set(IDS.values()) | |
| 62 | + bachelor_id: int | None = None | |
| 63 | + with open(csv_path, encoding="latin-1", newline="") as fh: | |
| 64 | + r = csv.reader(fh) | |
| 65 | + next(r) | |
| 66 | + cur = None | |
| 67 | + for row in r: | |
| 68 | + dg = row[1] | |
| 69 | + if not dg.startswith("2021S051624"): | |
| 70 | + if out and cur and not dg.startswith("2021S0516"): | |
| 71 | + break | |
| 72 | + continue | |
| 73 | + cid = int(row[8]) | |
| 74 | + if cid not in wanted: | |
| 75 | + continue | |
| 76 | + if cid == IDS["bachelor"] and row[9].strip() != BACHELOR_LABEL: | |
| 77 | + log.warning("libellé inattendu pour l'id %d : %r", cid, row[9]) | |
| 78 | + bachelor_id = IDS["bachelor"] | |
| 79 | + if dg != cur: | |
| 80 | + cur = dg | |
| 81 | + out[dg] = {} | |
| 82 | + key = next(k for k, v in IDS.items() if v == cid) | |
| 83 | + out[dg][key] = _num(row[11]) # C1_CHIFFRE_TOTAL | |
| 84 | + log.info("ADA Québec lues : %d (id baccalauréat = %s)", len(out), bachelor_id) | |
| 85 | + return out | |
| 86 | + | |
| 87 | + | |
| 88 | +def read_ada_geometries(zip_path: Path) -> dict[str, object]: | |
| 89 | + import pyproj | |
| 90 | + import shapefile | |
| 91 | + from shapely.geometry import shape | |
| 92 | + from shapely.ops import transform | |
| 93 | + z = zipfile.ZipFile(zip_path) | |
| 94 | + base = next(n[:-4] for n in z.namelist() if n.endswith(".shp")) | |
| 95 | + rd = shapefile.Reader(shp=io.BytesIO(z.read(base + ".shp")), dbf=io.BytesIO(z.read(base + ".dbf")), shx=io.BytesIO(z.read(base + ".shx"))) | |
| 96 | + prj = z.read(base + ".prj").decode() | |
| 97 | + tr = pyproj.Transformer.from_crs(pyproj.CRS.from_wkt(prj), "EPSG:4326", always_xy=True).transform | |
| 98 | + fields = [f[0] for f in rd.fields[1:]] | |
| 99 | + i_dg = fields.index("IDUGD") | |
| 100 | + geoms = {} | |
| 101 | + for sr in rd.iterShapeRecords(): | |
| 102 | + dg = sr.record[i_dg] | |
| 103 | + if not str(dg).startswith("2021S051624"): | |
| 104 | + continue | |
| 105 | + g = shape(sr.shape.__geo_interface__) | |
| 106 | + if g.is_empty: | |
| 107 | + continue | |
| 108 | + geoms[dg] = transform(tr, g).buffer(0) | |
| 109 | + log.info("géométries ADA Québec : %d", len(geoms)) | |
| 110 | + return geoms | |
| 111 | + | |
| 112 | + | |
| 113 | +def aggregate(profile: dict, geoms: dict, ridings_geo: dict) -> dict[int, dict]: | |
| 114 | + """Répartition surfacique ADA → circonscriptions 2026.""" | |
| 115 | + from shapely.strtree import STRtree | |
| 116 | + codes = list(ridings_geo) | |
| 117 | + polys = [ridings_geo[c] for c in codes] | |
| 118 | + tree = STRtree(polys) | |
| 119 | + acc: dict[int, dict[str, float]] = defaultdict(lambda: defaultdict(float)) | |
| 120 | + counts: dict[int, set] = defaultdict(set) | |
| 121 | + count_keys = ["pop", "age65", "mt_total", "mt_en", "mt_fr", "mt_other", "plop_total", "plop_fr", "tenure_total", "owner", | |
| 122 | + "immig_total", "immigrants", "vismin_total", "vismin", "educ_total", "bachelor"] | |
| 123 | + rate_keys = ["median_age", "income_median", "unemployment", "density"] | |
| 124 | + for dg, g in geoms.items(): | |
| 125 | + p = profile.get(dg) | |
| 126 | + if not p or not p.get("pop"): | |
| 127 | + continue | |
| 128 | + area = g.area | |
| 129 | + if area <= 0: | |
| 130 | + continue | |
| 131 | + for idx in tree.query(g, predicate="intersects"): | |
| 132 | + inter = g.intersection(polys[idx]) | |
| 133 | + if inter.is_empty: | |
| 134 | + continue | |
| 135 | + w = inter.area / area | |
| 136 | + if w < 1e-4: | |
| 137 | + continue | |
| 138 | + code = codes[idx] | |
| 139 | + counts[code].add(dg) | |
| 140 | + pop_w = (p["pop"] or 0) * w | |
| 141 | + for k in count_keys: | |
| 142 | + if p.get(k) is not None: | |
| 143 | + acc[code][k] += p[k] * w | |
| 144 | + for k in rate_keys: | |
| 145 | + if p.get(k) is not None: | |
| 146 | + acc[code][f"{k}_wsum"] += p[k] * pop_w | |
| 147 | + acc[code][f"{k}_w"] += pop_w | |
| 148 | + out: dict[int, dict] = {} | |
| 149 | + for code, a in acc.items(): | |
| 150 | + pop = a["pop"] | |
| 151 | + def pct(n, d): | |
| 152 | + return round(100 * a[n] / a[d], 2) if a.get(d) else None | |
| 153 | + def wavg(k): | |
| 154 | + return round(a[f"{k}_wsum"] / a[f"{k}_w"], 1) if a.get(f"{k}_w") else None | |
| 155 | + out[code] = { | |
| 156 | + "population_2021": int(round(pop)), "density": wavg("density"), "median_age": wavg("median_age"), | |
| 157 | + "pct_65_plus": pct("age65", "pop"), "pct_french_mt": pct("mt_fr", "mt_total"), "pct_english_mt": pct("mt_en", "mt_total"), | |
| 158 | + "pct_other_mt": pct("mt_other", "mt_total"), "pct_plop_french": pct("plop_fr", "plop_total"), | |
| 159 | + "pct_immigrants": pct("immigrants", "immig_total"), "pct_visible_minority": pct("vismin", "vismin_total"), | |
| 160 | + "median_household_income": wavg("income_median"), "pct_owner": pct("owner", "tenure_total"), | |
| 161 | + "pct_bachelor_plus": pct("bachelor", "educ_total"), "unemployment_rate": wavg("unemployment"), | |
| 162 | + "ada_count": len(counts[code]), | |
| 163 | + } | |
| 164 | + return out | |
| 165 | + | |
| 166 | + | |
| 167 | +def build(csv_path: Path, geo_zip: Path) -> Path: | |
| 168 | + """Calcule le fichier dérivé (à exécuter là où les fichiers StatCan sont disponibles).""" | |
| 169 | + from shapely.geometry import shape | |
| 170 | + gj = json.loads((settings.official_dir / "circonscriptions_electorales_sans_eau_2026.json").read_text()) | |
| 171 | + ridings_geo = {int(f["properties"]["CO_CEP"]): shape(f["geometry"]).buffer(0) for f in gj["features"]} | |
| 172 | + profile = read_ada_profile(csv_path) | |
| 173 | + geoms = read_ada_geometries(geo_zip) | |
| 174 | + agg = aggregate(profile, geoms, ridings_geo) | |
| 175 | + DERIVED.mkdir(parents=True, exist_ok=True) | |
| 176 | + cols = ["riding_code", "population_2021", "density", "median_age", "pct_65_plus", "pct_french_mt", "pct_english_mt", "pct_other_mt", | |
| 177 | + "pct_plop_french", "pct_immigrants", "pct_visible_minority", "median_household_income", "pct_owner", "pct_bachelor_plus", | |
| 178 | + "unemployment_rate", "ada_count"] | |
| 179 | + with open(OUT_CSV, "w", newline="", encoding="utf-8") as fh: | |
| 180 | + w = csv.writer(fh) | |
| 181 | + w.writerow(cols) | |
| 182 | + for code in sorted(agg): | |
| 183 | + w.writerow([code] + [agg[code].get(c) for c in cols[1:]]) | |
| 184 | + (DERIVED / "riding_demographics_2021.SOURCE.txt").write_text(SOURCE + f"\ncomputed_at={utcnow().isoformat()}\nADA={len(geoms)}\n") | |
| 185 | + log.info("fichier dérivé écrit : %s (%d circonscriptions)", OUT_CSV, len(agg)) | |
| 186 | + return OUT_CSV | |
| 187 | + | |
| 188 | + | |
| 189 | +def load(s: Session, path: Path = OUT_CSV) -> int: | |
| 190 | + """Charge le fichier dérivé en base (idempotent).""" | |
| 191 | + if not path.exists(): | |
| 192 | + log.warning("fichier démographique absent : %s", path) | |
| 193 | + return 0 | |
| 194 | + ridings = {r.code for r in s.scalars(select(Riding)).all()} | |
| 195 | + total_pop = 0 | |
| 196 | + n = 0 | |
| 197 | + with open(path, encoding="utf-8") as fh: | |
| 198 | + for row in csv.DictReader(fh): | |
| 199 | + code = int(row["riding_code"]) | |
| 200 | + if code not in ridings: | |
| 201 | + continue | |
| 202 | + d = s.get(RidingDemographics, code) or RidingDemographics(riding_code=code) | |
| 203 | + for k, v in row.items(): | |
| 204 | + if k == "riding_code": | |
| 205 | + continue | |
| 206 | + val = None if v in ("", "None") else (int(float(v)) if k in ("population_2021", "ada_count") else float(v)) | |
| 207 | + setattr(d, k, val) | |
| 208 | + d.source = SOURCE | |
| 209 | + d.coverage = 1.0 | |
| 210 | + s.add(d) | |
| 211 | + total_pop += d.population_2021 | |
| 212 | + n += 1 | |
| 213 | + src = s.get(Source, "statcan-census-2021") | |
| 214 | + if not src: | |
| 215 | + src = Source(id="statcan-census-2021", name="Statistique Canada — Recensement 2021 (profil ADA + limites)", tier=1, kind="official", | |
| 216 | + url="https://www12.statcan.gc.ca/census-recensement/2021/dp-pd/prof/details/download-telecharger.cfm?Lang=F") | |
| 217 | + s.add(src) | |
| 218 | + src.last_checked_at = src.last_ok_at = utcnow() | |
| 219 | + src.status = "operational" | |
| 220 | + src.detail = f"{n} circonscriptions, population totale {total_pop:,}".replace(",", " ") | |
| 221 | + log.info("démographie chargée : %d circonscriptions, population %d", n, total_pop) | |
| 222 | + return n | |
modified
engine/qc26/main.py
+11 −0
@@ -33,6 +33,17 @@ def bootstrap() -> None: | ||
| 33 | 33 | log.info("base vide : ingestion des données officielles") |
| 34 | 34 | official.run_all(s, refresh=True) |
| 35 | 35 | polls_bootstrap.load(s) |
| 36 | + from .db import RidingDemographics | |
| 37 | + from .ingest import census | |
| 38 | + if not s.scalar(select(RidingDemographics).limit(1)): | |
| 39 | + census.load(s) # fichier dérivé versionné (data/derived) | |
| 40 | + try: | |
| 41 | + from .connectors.polls.registry import ensure_reference_polls | |
| 42 | + if ensure_reference_polls(s): | |
| 43 | + from .model.run import run_forecast | |
| 44 | + run_forecast(s, trigger="reference_polls_2022") | |
| 45 | + except Exception as e: # noqa: BLE001 | |
| 46 | + log.warning("sondages de référence 2022 : %s", e) | |
| 36 | 47 | if not s.scalar(select(ForecastRun).where(ForecastRun.status == "ok").limit(1)): |
| 37 | 48 | from .model.run import run_forecast |
| 38 | 49 | try: |
modified
engine/qc26/model/ridings.py
+37 −14
@@ -1,16 +1,17 @@ | ||
| 1 | −"""Modèle de circonscription : priors 2022 transposés, swing proportionnel régional, | |
| 2 | −effets candidats, simulation Monte Carlo à erreurs corrélées. | |
| 1 | +"""Modèle de circonscription (v1.1) : priors 2022 transposés, swing proportionnel national, | |
| 2 | +excès de swing par segments (langue × composition du recensement, région), effets candidats, | |
| 3 | +calibration nationale et simulation Monte Carlo à erreurs corrélées. | |
| 3 | 4 | |
| 4 | 5 | Pour chaque tirage national v (vote %), la part du parti p dans la circonscription r : |
| 5 | − share_r,p ∝ baseline_r,p × (v_p / national2022_p)^β (swing proportionnel, β=1) | |
| 6 | − × exp(ε_région,p) × exp(ε_circ,p) × exp(effets candidats) | |
| 7 | −puis normalisation. Les erreurs régionales et locales sont tirées par simulation | |
| 8 | −(indépendantes entre régions/circonscriptions, corrélées via le tirage national). | |
| 6 | + logit-ratio_r,p = log(baseline_r,p) − log(national2022_p) + adj_segments_r,p + effets_candidats_r,p | |
| 7 | + share_r,p ∝ exp(log v_p + logit-ratio_r,p + ε_région,p + ε_circ,p) | |
| 8 | +puis normalisation, puis calibration : la moyenne des parts pondérée par les électeurs | |
| 9 | +inscrits est ramenée exactement au vote national simulé (une itération multiplicative). | |
| 9 | 10 | Les partis absents d'une circonscription en 2022 reçoivent un plancher (0,3 %). |
| 10 | 11 | """ |
| 11 | 12 | from __future__ import annotations |
| 12 | 13 | |
| 13 | −from dataclasses import dataclass | |
| 14 | +from dataclasses import dataclass, field | |
| 14 | 15 | |
| 15 | 16 | import numpy as np |
| 16 | 17 | |
@@ -27,11 +28,14 @@ class RidingInput: | ||
| 27 | 28 | incumbent_party: str | None |
| 28 | 29 | incumbent_running: bool | None |
| 29 | 30 | candidates: set[str] # partis présentant un·e candidat·e (DGEQ) |
| 31 | + electors: int = 50000 | |
| 32 | + pct_french: float | None = None # langue maternelle française (recensement 2021) | |
| 33 | + poll_region: str = "reste" | |
| 30 | 34 | |
| 31 | 35 | |
| 32 | 36 | def simulate_ridings(ridings: list[RidingInput], national_draws: np.ndarray, national_2022: dict[str, float], |
| 33 | 37 | rng: np.random.Generator, regional_sd: float | None = None, riding_sd: float | None = None, |
| 34 | − candidate_effects: bool = True) -> tuple[np.ndarray, np.ndarray]: | |
| 38 | + candidate_effects: bool = True, seg_adj: np.ndarray | None = None, calibrate: bool = True) -> tuple[np.ndarray, np.ndarray]: | |
| 35 | 39 | """national_draws : (S, P) en % ; retourne (winners (S, R) indices de parti, shares (S, R, P) en %).""" |
| 36 | 40 | regional_sd = settings.regional_error_sd if regional_sd is None else regional_sd |
| 37 | 41 | riding_sd = settings.riding_error_sd if riding_sd is None else riding_sd |
@@ -40,9 +44,6 @@ def simulate_ridings(ridings: list[RidingInput], national_draws: np.ndarray, nat | ||
| 40 | 44 | regions = sorted({r.region_id for r in ridings}) |
| 41 | 45 | reg_idx = {g: i for i, g in enumerate(regions)} |
| 42 | 46 | nat22 = np.array([max(national_2022.get(p, 0.5), 0.5) for p in PARTIES_ALL]) / 100.0 |
| 43 | − # La règle « pas de candidature → part résiduelle » ne s'applique qu'aux partis dont la | |
| 44 | − # liste est (quasi) complète : avant la clôture des candidatures, une absence n'est pas | |
| 45 | − # une information. | |
| 46 | 47 | counts = {p: sum(1 for r in ridings if p in r.candidates) for p in PARTY_IDS} |
| 47 | 48 | complete = {p for p, n in counts.items() if n >= 0.9 * R} |
| 48 | 49 | base = np.zeros((R, P)) |
@@ -54,6 +55,8 @@ def simulate_ridings(ridings: list[RidingInput], national_draws: np.ndarray, nat | ||
| 54 | 55 | base[i, j] = 0.001 # pas de candidat·e officiel·le : part résiduelle |
| 55 | 56 | if candidate_effects and r.incumbent_party and r.incumbent_party in PARTIES_ALL and r.incumbent_running is False: |
| 56 | 57 | adj[i, PARTIES_ALL.index(r.incumbent_party)] -= settings.incumbent_retirement_penalty |
| 58 | + if seg_adj is not None: | |
| 59 | + adj = adj + seg_adj | |
| 57 | 60 | log_ratio = np.log(base) - np.log(nat22)[None, :] + adj # (R, P) |
| 58 | 61 | reg_eps = rng.normal(0, regional_sd, size=(S, len(regions), P)) |
| 59 | 62 | reg_eps[:, :, -1] *= 0.5 |
@@ -64,6 +67,14 @@ def simulate_ridings(ridings: list[RidingInput], national_draws: np.ndarray, nat | ||
| 64 | 67 | logits = nat[:, None, :] + log_ratio[None, :, :] + reg_eps[:, reg_of, :] + rid_eps |
| 65 | 68 | ex = np.exp(logits - logits.max(axis=2, keepdims=True)) |
| 66 | 69 | shares = ex / ex.sum(axis=2, keepdims=True) |
| 70 | + if calibrate: | |
| 71 | + w = np.array([max(r.electors or 1, 1) for r in ridings], dtype=float) | |
| 72 | + w = w / w.sum() | |
| 73 | + for _ in range(2): | |
| 74 | + implied = np.einsum("r,srp->sp", w, shares) # (S, P) | |
| 75 | + ratio = np.clip((national_draws / 100.0) / np.clip(implied, 1e-6, None), 0.5, 2.0) | |
| 76 | + shares = shares * ratio[:, None, :] | |
| 77 | + shares = shares / shares.sum(axis=2, keepdims=True) | |
| 67 | 78 | # « Autres » agrège des candidatures hétérogènes : jamais déclaré gagnant (v1) |
| 68 | 79 | winners = shares[:, :, : len(PARTY_IDS)].argmax(axis=2) |
| 69 | 80 | return winners, shares * 100.0 |
@@ -80,7 +91,8 @@ def rating(p_win: float) -> str: | ||
| 80 | 91 | |
| 81 | 92 | |
| 82 | 93 | def explain_factors(r: RidingInput, national_now: dict[str, float], national_2022: dict[str, float], |
| 83 | − favorite: str, share_now: dict[str, float]) -> list[dict]: | |
| 94 | + favorite: str, share_now: dict[str, float], seg_adj_fav: float | None = None, | |
| 95 | + seg_detail: dict | None = None) -> list[dict]: | |
| 84 | 96 | """Facteurs réellement utilisés par le modèle pour la circonscription (« Pourquoi ? »).""" |
| 85 | 97 | out = [] |
| 86 | 98 | b = r.baseline.get(favorite, 0.0) |
@@ -90,14 +102,25 @@ def explain_factors(r: RidingInput, national_now: dict[str, float], national_202 | ||
| 90 | 102 | out.append({"factor": "tendance_provinciale", "label": "Tendance provinciale (sondages agrégés)", |
| 91 | 103 | "detail": f"{favorite.upper()} : {national_now.get(favorite, 0):.1f} % aujourd'hui vs {national_2022.get(favorite, 0):.1f} % en 2022 ({swing:+.1f} pt)", |
| 92 | 104 | "direction": "+" if swing > 0 else "-"}) |
| 93 | − out.append({"factor": "swing_proportionnel", "label": "Swing proportionnel appliqué au territoire", | |
| 105 | + if r.pct_french is not None and seg_detail and seg_detail.get("language"): | |
| 106 | + lang = seg_detail["language"] | |
| 107 | + out.append({"factor": "composition_linguistique", "label": "Composition linguistique (Recensement 2021) × ventilations des sondages", | |
| 108 | + "detail": (f"{r.pct_french:.0f} % de langue maternelle française. {favorite.upper()} chez les francophones : {lang['fr_now']:.0f} % " | |
| 109 | + f"(vs {lang['fr_2022']:.0f} % en 2022) ; non-francophones : {lang['nonfr_now']:.0f} % (vs {lang['nonfr_2022']:.0f} %)."), | |
| 110 | + "direction": "+" if (seg_adj_fav or 0) > 0.02 else "-" if (seg_adj_fav or 0) < -0.02 else "±"}) | |
| 111 | + if seg_detail and seg_detail.get("region"): | |
| 112 | + reg = seg_detail["region"] | |
| 113 | + out.append({"factor": "region_sondages", "label": f"Ventilation régionale des sondages ({reg['label']})", | |
| 114 | + "detail": f"{favorite.upper()} : {reg['now']:.0f} % dans la région aujourd'hui vs {reg['2022']:.0f} % en 2022, contre {swing:+.1f} pt au national.", | |
| 115 | + "direction": "±"}) | |
| 116 | + out.append({"factor": "swing_proportionnel", "label": "Swing proportionnel appliqué au territoire, calibré au national", | |
| 94 | 117 | "detail": f"part projetée {share_now.get(favorite, 0):.1f} % (moyenne des simulations)", "direction": "±"}) |
| 95 | 118 | if r.incumbent_party: |
| 96 | 119 | if r.incumbent_running is False: |
| 97 | 120 | out.append({"factor": "sortant", "label": "Député·e sortant·e ne se représente pas", |
| 98 | 121 | "detail": f"pénalité de ~1,3 pt appliquée à {r.incumbent_party.upper()}", "direction": "-" if r.incumbent_party == favorite else "+"}) |
| 99 | 122 | elif r.incumbent_running: |
| 100 | − out.append({"factor": "sortant", "label": "Député·e sortant·e candidat·e", "detail": f"{r.incumbent_party.upper()} (aucun bonus appliqué ; effet neutre dans le modèle v1)", "direction": "±"}) | |
| 123 | + out.append({"factor": "sortant", "label": "Député·e sortant·e candidat·e", "detail": f"{r.incumbent_party.upper()} (aucun bonus appliqué ; effet neutre dans le modèle)", "direction": "±"}) | |
| 101 | 124 | missing = [p for p in PARTY_IDS if r.candidates and p not in r.candidates] |
| 102 | 125 | if missing: |
| 103 | 126 | out.append({"factor": "candidatures", "label": "Candidatures non encore enregistrées (DGEQ)", |
modified
engine/qc26/model/run.py
+22 −4
@@ -13,10 +13,12 @@ from sqlalchemy.orm import Session | ||
| 13 | 13 | |
| 14 | 14 | from ..config import settings |
| 15 | 15 | from ..db import (Candidate, ForecastHouseEffect, ForecastPartyResult, ForecastRidingResult, ForecastRidingSummary, |
| 16 | − ForecastRun, Party, Poll, PollWeight, Riding, RidingBaseline, get_setting, set_setting, utcnow) | |
| 16 | + ForecastRun, Party, Poll, PollWeight, Riding, RidingBaseline, RidingDemographics, get_setting, set_setting, utcnow) | |
| 17 | 17 | from ..parties import PARTY_IDS |
| 18 | +from ..regions import POLL_REGION_LABELS, poll_region_for | |
| 18 | 19 | from .aggregate import PARTIES_ALL, Obs, compute_weights, election_day_distribution, house_effects, kalman_latent |
| 19 | 20 | from .ridings import RidingInput, explain_factors, rating, simulate_ridings |
| 21 | +from .segments import build_segment_model, riding_adjustments | |
| 20 | 22 | |
| 21 | 23 | log = logging.getLogger("qc26.model") |
| 22 | 24 | |
@@ -51,8 +53,11 @@ def load_ridings(s: Session) -> tuple[list[RidingInput], dict[str, float]]: | ||
| 51 | 53 | for code, pid in s.execute(select(Candidate.riding_code, Candidate.party_id)).all(): |
| 52 | 54 | if pid: |
| 53 | 55 | cands[code].add(pid) |
| 56 | + demo = {d.riding_code: d for d in s.scalars(select(RidingDemographics)).all()} | |
| 54 | 57 | out = [RidingInput(code=r.code, region_id=r.region_id or "inconnue", baseline=baselines.get(r.code, {}), |
| 55 | − incumbent_party=r.incumbent_party_id, incumbent_running=r.incumbent_running, candidates=cands.get(r.code, set())) | |
| 58 | + incumbent_party=r.incumbent_party_id, incumbent_running=r.incumbent_running, candidates=cands.get(r.code, set()), | |
| 59 | + electors=r.electors_2026 or 50000, pct_french=(demo[r.code].pct_french_mt if r.code in demo else None), | |
| 60 | + poll_region=poll_region_for(r.code)) | |
| 56 | 61 | for r in rows if baselines.get(r.code)] |
| 57 | 62 | national_2022 = {p.id: (p.vote_2022 or 0.0) for p in s.scalars(select(Party)).all()} |
| 58 | 63 | national_2022["aut"] = max(0.5, 100.0 - sum(national_2022.values())) |
@@ -86,7 +91,10 @@ def run_forecast(s: Session, trigger: str = "manual", cutoff: date | None = None | ||
| 86 | 91 | national = _draw_national(mu, cov, n_sims, rng) # (S, P) % |
| 87 | 92 | |
| 88 | 93 | ridings, national_2022 = load_ridings(s) |
| 89 | − winners, shares = simulate_ridings(ridings, national, national_2022, rng) | |
| 94 | + national_now_mean = {p: float(national[:, j].mean()) for j, p in enumerate(PARTIES_ALL)} | |
| 95 | + seg_model, f_r = build_segment_model(s, cutoff, national_now_mean, national_2022) | |
| 96 | + seg_adj = riding_adjustments(seg_model, [r.code for r in ridings], f_r) | |
| 97 | + winners, shares = simulate_ridings(ridings, national, national_2022, rng, seg_adj=seg_adj) | |
| 90 | 98 | S, R, P = shares.shape |
| 91 | 99 | seats_total = int(get_setting(s, "seats_total", R)) |
| 92 | 100 | majority = int(get_setting(s, "majority_threshold", seats_total // 2 + 1)) |
@@ -139,10 +147,18 @@ def run_forecast(s: Session, trigger: str = "manual", cutoff: date | None = None | ||
| 139 | 147 | fav_id = PARTIES_ALL[fav] |
| 140 | 148 | margin = share_mean[fav_id] - share_mean[PARTIES_ALL[second]] |
| 141 | 149 | swing = share_mean[fav_id] - r.baseline.get(fav_id, 0.0) |
| 150 | + seg_detail = {} | |
| 151 | + if "fr" in seg_model.current and "nonfr" in seg_model.current and fav_id in seg_model.current["fr"]: | |
| 152 | + seg_detail["language"] = {"fr_now": seg_model.current["fr"][fav_id], "fr_2022": seg_model.baseline.get("fr", {}).get(fav_id, 0), | |
| 153 | + "nonfr_now": seg_model.current["nonfr"][fav_id], "nonfr_2022": seg_model.baseline.get("nonfr", {}).get(fav_id, 0)} | |
| 154 | + if r.poll_region in seg_model.current and fav_id in seg_model.current[r.poll_region]: | |
| 155 | + seg_detail["region"] = {"label": POLL_REGION_LABELS.get(r.poll_region, r.poll_region), "now": seg_model.current[r.poll_region][fav_id], | |
| 156 | + "2022": seg_model.baseline.get(r.poll_region, {}).get(fav_id, 0)} | |
| 142 | 157 | summ = ForecastRidingSummary(run_id=run.id, riding_code=r.code, favorite_party_id=fav_id, favorite_p=float(pw[fav]), |
| 143 | 158 | runner_up_party_id=PARTIES_ALL[second], margin_mean=float(margin), |
| 144 | 159 | volatility=float(1 - pw[fav]), rating=rating(float(pw[fav])), swing_from_2022=float(swing), |
| 145 | − factors=explain_factors(r, national_now, national_2022, fav_id, share_mean)) | |
| 160 | + factors=explain_factors(r, national_now, national_2022, fav_id, share_mean, | |
| 161 | + seg_adj_fav=float(seg_adj[i, fav]), seg_detail=seg_detail)) | |
| 146 | 162 | s.add(summ) |
| 147 | 163 | riding_rows.append((r.code, fav_id, float(pw[fav]), PARTIES_ALL[second], float(margin))) |
| 148 | 164 | battlegrounds.append((float(1 - pw[fav]), r.code)) |
@@ -185,6 +201,8 @@ def run_forecast(s: Session, trigger: str = "manual", cutoff: date | None = None | ||
| 185 | 201 | "leader": top[0][0], "poll_count": len(obs), "latest_poll": max(o.end for o in obs).isoformat(), |
| 186 | 202 | "national_now": national_now, "cutoff": cutoff.isoformat(), "days_to_election": (settings.election_date - cutoff).days, |
| 187 | 203 | "battlegrounds": [c for _, c in sorted(battlegrounds, reverse=True)[:15]], |
| 204 | + "segments": seg_model.payload(), | |
| 205 | + "demographics_coverage": sum(1 for r in ridings if r.pct_french is not None), | |
| 188 | 206 | "ridings": {str(c): {"fav": f, "p": round(p, 3), "second": sec, "margin": round(m, 1)} for c, f, p, sec, m in riding_rows}, |
| 189 | 207 | } |
| 190 | 208 | diff = None |
added
engine/qc26/model/segments.py
+235 −0
@@ -0,0 +1,235 @@ | ||
| 1 | +"""Swing par segments (langue, région) — modèle 1.1. | |
| 2 | + | |
| 3 | +Idée : le swing n'est pas uniforme. Les firmes publient des ventilations par langue | |
| 4 | +(francophones / non-francophones) et par région (RMR de Montréal / RMR de Québec / reste). | |
| 5 | +On en tire, pour chaque parti, l'écart entre le swing du segment et le swing national ; | |
| 6 | +cet « excès de swing » est appliqué à chaque circonscription selon sa composition : | |
| 7 | +- langue : part de langue maternelle française (Recensement 2021, agrégé aux circonscriptions) ; | |
| 8 | +- région : regroupement « sondeurs » de la circonscription. | |
| 9 | + | |
| 10 | +Références 2022 par segment : | |
| 11 | +- région : résultats 2022 transposés, agrégés par regroupement (exact) ; | |
| 12 | +- langue : ventilations francophones / non-francophones des sondages de fin de campagne 2022 | |
| 13 | + (terrain dans les 14 jours précédant le 3 octobre 2022, documents originaux archivés), | |
| 14 | + recalées sur le résultat officiel (chaque parti est multiplié par résultat/sondage national, | |
| 15 | + puis renormalisé). Sans sondage de référence, la dimension linguistique est désactivée. | |
| 16 | + La régression écologique (résultats 2022 ~ part francophone) est publiée à titre de | |
| 17 | + diagnostic mais n'est PAS utilisée comme référence : l'extrapolation à f = 0 est instable. | |
| 18 | +Les excès de swing sont rétrécis (κ) car les sous-échantillons sont bruités. | |
| 19 | +""" | |
| 20 | +from __future__ import annotations | |
| 21 | + | |
| 22 | +import math | |
| 23 | +from collections import defaultdict | |
| 24 | +from dataclasses import dataclass, field | |
| 25 | +from datetime import date | |
| 26 | + | |
| 27 | +import numpy as np | |
| 28 | +from sqlalchemy import select | |
| 29 | +from sqlalchemy.orm import Session | |
| 30 | + | |
| 31 | +from ..config import settings | |
| 32 | +from ..db import Poll, PollSubsample, RidingBaseline, RidingDemographics, Riding | |
| 33 | +from ..parties import PARTY_IDS | |
| 34 | +from ..regions import poll_region_for | |
| 35 | +from .aggregate import PARTIES_ALL | |
| 36 | + | |
| 37 | +KAPPA_LANGUAGE = 0.6 | |
| 38 | +KAPPA_REGION = 0.5 | |
| 39 | +WINDOW_DAYS = 90 | |
| 40 | +HALFLIFE_DAYS = 21.0 | |
| 41 | +MIN_POLLS = 2 | |
| 42 | + | |
| 43 | +LANGUAGE_MAP = {"francophones": "fr", "francophone": "fr", "non-francophones": "nonfr", "non francophones": "nonfr", "nonfrancophones": "nonfr", | |
| 44 | + "anglophones": "nonfr", "allophones": "nonfr"} | |
| 45 | +REGION_MAP = {"montréal rmr": "montreal_rmr", "montreal rmr": "montreal_rmr", "rmr de montréal": "montreal_rmr", "grand montréal": "montreal_rmr", | |
| 46 | + "québec rmr": "quebec_rmr", "quebec rmr": "quebec_rmr", "rmr de québec": "quebec_rmr", "région de québec": "quebec_rmr", | |
| 47 | + "reste du québec": "reste", "ailleurs au québec": "reste", "autres régions": "reste", "régions": "reste"} | |
| 48 | + | |
| 49 | + | |
| 50 | +def canonical(dimension: str, segment: str) -> str | None: | |
| 51 | + s = segment.strip().lower() | |
| 52 | + if dimension == "language": | |
| 53 | + return LANGUAGE_MAP.get(s) | |
| 54 | + if dimension == "region": | |
| 55 | + return REGION_MAP.get(s) | |
| 56 | + return None | |
| 57 | + | |
| 58 | + | |
| 59 | +def _logit(p: float) -> float: | |
| 60 | + p = min(max(p, 0.003), 0.997) | |
| 61 | + return math.log(p / (1 - p)) | |
| 62 | + | |
| 63 | + | |
| 64 | +@dataclass | |
| 65 | +class SegmentModel: | |
| 66 | + current: dict[str, dict[str, float]] = field(default_factory=dict) # seg → party → % | |
| 67 | + baseline: dict[str, dict[str, float]] = field(default_factory=dict) # seg → party → % (2022) | |
| 68 | + n_polls: dict[str, int] = field(default_factory=dict) | |
| 69 | + excess: dict[str, dict[str, float]] = field(default_factory=dict) # seg → party → excès de swing (logit) | |
| 70 | + national_now: dict[str, float] = field(default_factory=dict) | |
| 71 | + national_2022: dict[str, float] = field(default_factory=dict) | |
| 72 | + language_regression: dict[str, dict[str, float]] = field(default_factory=dict) # party → {a, b, r2} (diagnostic) | |
| 73 | + reference_polls_2022: list[str] = field(default_factory=list) | |
| 74 | + | |
| 75 | + def payload(self) -> dict: | |
| 76 | + return {"current": self.current, "baseline2022": self.baseline, "nPolls": self.n_polls, | |
| 77 | + "excessSwingLogit": {k: {p: round(v, 3) for p, v in d.items()} for k, d in self.excess.items()}, | |
| 78 | + "languageRegressionDiagnostic": self.language_regression, "referencePolls2022": self.reference_polls_2022, | |
| 79 | + "languageEnabled": "fr" in self.excess and "nonfr" in self.excess, | |
| 80 | + "kappa": {"language": KAPPA_LANGUAGE, "region": KAPPA_REGION}, "window_days": WINDOW_DAYS} | |
| 81 | + | |
| 82 | + | |
| 83 | +def current_segment_support(s: Session, cutoff: date) -> tuple[dict[str, dict[str, float]], dict[str, int]]: | |
| 84 | + since = date.fromordinal(cutoff.toordinal() - WINDOW_DAYS) | |
| 85 | + rows = s.execute(select(PollSubsample, Poll).join(Poll, Poll.id == PollSubsample.poll_id) | |
| 86 | + .where(Poll.excluded == False, Poll.validated == True, Poll.fieldwork_end >= since, Poll.fieldwork_end <= cutoff, # noqa: E712 | |
| 87 | + PollSubsample.dimension.in_(["language", "region"]))).all() | |
| 88 | + acc: dict[str, dict[str, float]] = defaultdict(lambda: defaultdict(float)) | |
| 89 | + wsum: dict[str, float] = defaultdict(float) | |
| 90 | + polls: dict[str, set] = defaultdict(set) | |
| 91 | + per_poll_seg: dict[tuple[str, str], dict[str, float]] = defaultdict(dict) | |
| 92 | + meta: dict[tuple[str, str], tuple[float, int | None]] = {} | |
| 93 | + for ss, poll in rows: | |
| 94 | + key = canonical(ss.dimension, ss.segment) | |
| 95 | + if not key or ss.party_id not in PARTY_IDS: | |
| 96 | + continue | |
| 97 | + per_poll_seg[(poll.id, key)][ss.party_id] = ss.value | |
| 98 | + age = (cutoff - poll.fieldwork_end).days | |
| 99 | + meta[(poll.id, key)] = (0.5 ** (age / HALFLIFE_DAYS) * math.sqrt(min(ss.n or 300, 1500) / 500.0), ss.n) | |
| 100 | + for (pid, key), vals in per_poll_seg.items(): | |
| 101 | + if len(vals) < 4: | |
| 102 | + continue | |
| 103 | + tot = sum(vals.values()) | |
| 104 | + if not (80 <= tot <= 105): | |
| 105 | + continue | |
| 106 | + w = meta[(pid, key)][0] | |
| 107 | + for party, v in vals.items(): | |
| 108 | + acc[key][party] += w * v | |
| 109 | + wsum[key] += w | |
| 110 | + polls[key].add(pid) | |
| 111 | + out: dict[str, dict[str, float]] = {} | |
| 112 | + for key, d in acc.items(): | |
| 113 | + if len(polls[key]) < MIN_POLLS: | |
| 114 | + continue | |
| 115 | + vals = {p: d[p] / wsum[key] for p in PARTY_IDS if p in d} | |
| 116 | + rest = max(0.0, 100 - sum(vals.values())) | |
| 117 | + vals["aut"] = rest | |
| 118 | + out[key] = {p: round(v, 2) for p, v in vals.items()} | |
| 119 | + return out, {k: len(v) for k, v in polls.items()} | |
| 120 | + | |
| 121 | + | |
| 122 | +ELECTION_2022 = date(2022, 10, 3) | |
| 123 | + | |
| 124 | + | |
| 125 | +def language_baseline_2022(s: Session, national_2022: dict[str, float]) -> tuple[dict[str, dict[str, float]], list[str]]: | |
| 126 | + """Soutien 2022 par langue : sondages de fin de campagne 2022 (≤ 14 j avant le vote), recalés sur le résultat.""" | |
| 127 | + since = date.fromordinal(ELECTION_2022.toordinal() - 14) | |
| 128 | + rows = s.execute(select(PollSubsample, Poll).join(Poll, Poll.id == PollSubsample.poll_id) | |
| 129 | + .where(Poll.excluded == False, Poll.validated == True, Poll.fieldwork_end >= since, Poll.fieldwork_end < ELECTION_2022, # noqa: E712 | |
| 130 | + PollSubsample.dimension == "language")).all() | |
| 131 | + per: dict[tuple[str, str], dict[str, float]] = defaultdict(dict) | |
| 132 | + nat: dict[str, dict[str, float]] = {} | |
| 133 | + used: set[str] = set() | |
| 134 | + for ss, poll in rows: | |
| 135 | + key = canonical(ss.dimension, ss.segment) | |
| 136 | + if not key or ss.party_id not in PARTY_IDS: | |
| 137 | + continue | |
| 138 | + per[(poll.id, key)][ss.party_id] = ss.value | |
| 139 | + nat[poll.id] = {r.party_id: r.value for r in poll.results} | |
| 140 | + used.add(poll.id) | |
| 141 | + if not per: | |
| 142 | + return {}, [] | |
| 143 | + acc: dict[str, dict[str, list[float]]] = defaultdict(lambda: defaultdict(list)) | |
| 144 | + for (pid, key), vals in per.items(): | |
| 145 | + if len(vals) < 4: | |
| 146 | + continue | |
| 147 | + for p, v in vals.items(): | |
| 148 | + # recalage : le sondage sous/sur-estimait le parti au national → même correction dans le segment | |
| 149 | + pn = nat[pid].get(p) | |
| 150 | + corr = (national_2022.get(p, pn) / pn) if pn and pn > 0 else 1.0 | |
| 151 | + acc[key][p].append(v * corr) | |
| 152 | + out: dict[str, dict[str, float]] = {} | |
| 153 | + for key, d in acc.items(): | |
| 154 | + vals = {p: sum(v) / len(v) for p, v in d.items()} | |
| 155 | + vals["aut"] = max(0.5, 100 - sum(vals.values())) | |
| 156 | + tot = sum(vals.values()) | |
| 157 | + out[key] = {p: round(100 * v / tot, 2) for p, v in vals.items()} | |
| 158 | + return out, sorted(used) | |
| 159 | + | |
| 160 | + | |
| 161 | +def baseline_segments(s: Session, national_2022: dict[str, float] | None = None) -> tuple[dict[str, dict[str, float]], dict[str, dict[str, float]], dict[int, float], list[str]]: | |
| 162 | + """Références 2022 : régions (agrégation exacte), langue (sondages de fin de campagne recalés) ; régression écologique = diagnostic.""" | |
| 163 | + ridings = {r.code: r for r in s.scalars(select(Riding)).all()} | |
| 164 | + demo = {d.riding_code: d for d in s.scalars(select(RidingDemographics)).all()} | |
| 165 | + votes: dict[int, dict[str, int]] = defaultdict(dict) | |
| 166 | + for b in s.scalars(select(RidingBaseline)).all(): | |
| 167 | + votes[b.riding_code][b.party_id] = b.votes | |
| 168 | + # régions | |
| 169 | + reg: dict[str, dict[str, int]] = defaultdict(lambda: defaultdict(int)) | |
| 170 | + for code, d in votes.items(): | |
| 171 | + pr = poll_region_for(code) | |
| 172 | + for p, v in d.items(): | |
| 173 | + reg[pr][p] += v | |
| 174 | + region_base = {} | |
| 175 | + for pr, d in reg.items(): | |
| 176 | + tot = sum(d.values()) or 1 | |
| 177 | + region_base[pr] = {p: round(100 * d.get(p, 0) / tot, 2) for p in PARTIES_ALL} | |
| 178 | + # langue : régression pondérée share = a + b·f | |
| 179 | + f_r: dict[int, float] = {} | |
| 180 | + for code in votes: | |
| 181 | + dm = demo.get(code) | |
| 182 | + if dm and dm.pct_french_mt is not None: | |
| 183 | + f_r[code] = dm.pct_french_mt / 100.0 | |
| 184 | + regression: dict[str, dict[str, float]] = {} | |
| 185 | + if len(f_r) >= 50: | |
| 186 | + codes = list(f_r) | |
| 187 | + X = np.array([f_r[c] for c in codes]) | |
| 188 | + W = np.array([sum(votes[c].values()) for c in codes], dtype=float) | |
| 189 | + for p in PARTIES_ALL: | |
| 190 | + y = np.array([100 * votes[c].get(p, 0) / max(1, sum(votes[c].values())) for c in codes]) | |
| 191 | + A = np.vstack([np.ones_like(X), X]).T * np.sqrt(W)[:, None] | |
| 192 | + coef, *_ = np.linalg.lstsq(A, y * np.sqrt(W), rcond=None) | |
| 193 | + a, b = float(coef[0]), float(coef[1]) | |
| 194 | + pred = a + b * X | |
| 195 | + ss_res = float((W * (y - pred) ** 2).sum()) | |
| 196 | + ss_tot = float((W * (y - np.average(y, weights=W)) ** 2).sum()) or 1.0 | |
| 197 | + regression[p] = {"a": round(a, 2), "b": round(b, 2), "r2": round(1 - ss_res / ss_tot, 3)} | |
| 198 | + lang_base, ref_polls = language_baseline_2022(s, national_2022 or {}) | |
| 199 | + base = {**region_base, **lang_base} | |
| 200 | + return base, regression, f_r, ref_polls | |
| 201 | + | |
| 202 | + | |
| 203 | +def build_segment_model(s: Session, cutoff: date, national_now: dict[str, float], national_2022: dict[str, float]) -> tuple[SegmentModel, dict[int, float]]: | |
| 204 | + cur, n_polls = current_segment_support(s, cutoff) | |
| 205 | + base, regression, f_r, ref_polls = baseline_segments(s, national_2022) | |
| 206 | + m = SegmentModel(current=cur, baseline=base, n_polls=n_polls, national_now=national_now, national_2022=national_2022, language_regression=regression) | |
| 207 | + m.reference_polls_2022 = ref_polls | |
| 208 | + nat_delta = {p: _logit(national_now.get(p, 1) / 100) - _logit(national_2022.get(p, 1) / 100) for p in PARTIES_ALL} | |
| 209 | + for seg, curvals in cur.items(): | |
| 210 | + if seg not in base: | |
| 211 | + continue | |
| 212 | + m.excess[seg] = {} | |
| 213 | + for p in PARTY_IDS: | |
| 214 | + if p in curvals and p in base[seg]: | |
| 215 | + d = _logit(curvals[p] / 100) - _logit(base[seg][p] / 100) | |
| 216 | + m.excess[seg][p] = d - nat_delta[p] | |
| 217 | + return m, f_r | |
| 218 | + | |
| 219 | + | |
| 220 | +def riding_adjustments(m: SegmentModel, ridings_codes: list[int], f_r: dict[int, float]) -> np.ndarray: | |
| 221 | + """Matrice (R, P) d'ajustements logit par circonscription.""" | |
| 222 | + R, P = len(ridings_codes), len(PARTIES_ALL) | |
| 223 | + adj = np.zeros((R, P)) | |
| 224 | + has_lang = "fr" in m.excess and "nonfr" in m.excess | |
| 225 | + for i, code in enumerate(ridings_codes): | |
| 226 | + for j, p in enumerate(PARTY_IDS): | |
| 227 | + a = 0.0 | |
| 228 | + if has_lang and code in f_r: | |
| 229 | + f = f_r[code] | |
| 230 | + a += KAPPA_LANGUAGE * (f * m.excess["fr"].get(p, 0.0) + (1 - f) * m.excess["nonfr"].get(p, 0.0)) | |
| 231 | + pr = poll_region_for(code) | |
| 232 | + if pr in m.excess: | |
| 233 | + a += KAPPA_REGION * m.excess[pr].get(p, 0.0) | |
| 234 | + adj[i, j] = a | |
| 235 | + return adj | |
modified
engine/qc26/regions.py
+20 −0
@@ -88,3 +88,23 @@ RIDING_REGION: dict[int, str] = { | ||
| 88 | 88 | } |
| 89 | 89 | |
| 90 | 90 | assert len(RIDING_REGION) == 127, len(RIDING_REGION) |
| 91 | + | |
| 92 | +# Regroupement « sondeurs » affiné par circonscription : les firmes ventilent Montréal RMR / | |
| 93 | +# Québec RMR / reste du Québec. Les régions administratives débordent des RMR ; on corrige ici. | |
| 94 | +_POLL_REGION_BY_REGION = {rid: pr for rid, _, pr in REGIONS} | |
| 95 | +_POLL_REGION_OVERRIDES: dict[int, str] = { | |
| 96 | + # Laurentides hors RMR de Montréal | |
| 97 | + 589: "reste", 641: "reste", 649: "reste", 637: "reste", | |
| 98 | + # Lanaudière hors RMR | |
| 99 | + 627: "reste", 629: "reste", 631: "reste", | |
| 100 | + # Montérégie hors RMR (Saint-Hyacinthe, Granby, Brome-Missisquoi, Iberville, Saint-Jean, Huntingdon, Richelieu, Beauharnois) | |
| 101 | + 239: "reste", 159: "reste", 157: "reste", 161: "reste", 167: "reste", 171: "reste", 241: "reste", 179: "reste", | |
| 102 | + # Capitale-Nationale hors RMR de Québec (Charlevoix, Portneuf) | |
| 103 | + 771: "reste", 727: "reste", | |
| 104 | + # Chaudière-Appalaches hors RMR (Beauces, Bellechasse, Côte-du-Sud, Lotbinière) | |
| 105 | + 781: "reste", 777: "reste", 797: "reste", 799: "reste", 787: "reste", | |
| 106 | +} | |
| 107 | + | |
| 108 | + | |
| 109 | +def poll_region_for(code: int) -> str: | |
| 110 | + return _POLL_REGION_OVERRIDES.get(code) or _POLL_REGION_BY_REGION.get(RIDING_REGION.get(code, ""), "reste") | |
modified
engine/qc26/scheduler.py
+6 −0
@@ -78,9 +78,15 @@ def job_daily_forecast(): | ||
| 78 | 78 | |
| 79 | 79 | @_guarded("documents") |
| 80 | 80 | def job_documents(): |
| 81 | + from sqlalchemy import select | |
| 81 | 82 | from .compass.positions import compute_positions |
| 82 | 83 | from .connectors.parties import documents |
| 83 | 84 | with session_scope() as s: |
| 85 | + # un travail documentaire lancé hors planificateur (admin, one-off) est peut-être en cours : ne pas doubler les appels LLM | |
| 86 | + recent = s.scalar(select(IngestionJob).where(IngestionJob.kind.like("party-docs-%"), IngestionJob.status == "running", | |
| 87 | + IngestionJob.started_at > utcnow() - timedelta(hours=3))) | |
| 88 | + if recent: | |
| 89 | + return {"skipped": f"travail {recent.kind} #{recent.id} en cours"} | |
| 84 | 90 | res = documents.sync_all(s, analyze=True) |
| 85 | 91 | pos = compute_positions(s, only_missing=True) |
| 86 | 92 | return {"docs": res, "positions": pos} |
modified
web/src/app/circonscription/[slug]/page.tsx
+28 −0
@@ -22,6 +22,8 @@ interface Detail { | ||
| 22 | 22 | forecast?: { runId: number; completedAt: string | null; modelVersion: string; favorite: string; p: number; runnerUp: string | null; margin: number; rating: string; volatility: number; swing: number; factors: { factor: string; label: string; detail: string; direction: string }[]; parties: { party: string; voteMean: number; vote80: [number, number]; pWin: number }[] }; |
| 23 | 23 | live: { status: string; updatedAt: string | null; bureaux: [number, number]; votesValid: number; turnout: number | null; leader: string | null; marginVotes: number | null; marginPct: number | null; candidates: { name: string; party_id: string; votes: number; share: number }[] | null; final: boolean } | null; |
| 24 | 24 | nearby?: { code: number; name: string; slug: string }[]; |
| 25 | + demographics?: { population2021: number; density: number | null; medianAge: number | null; pct65Plus: number | null; pctFrenchMt: number | null; pctEnglishMt: number | null; pctOtherMt: number | null; pctPlopFrench: number | null; pctImmigrants: number | null; pctVisibleMinority: number | null; medianHouseholdIncome: number | null; pctOwner: number | null; pctBachelorPlus: number | null; unemploymentRate: number | null; adaCount: number; source: string } | null; | |
| 26 | + pollRegion?: string; | |
| 25 | 27 | } |
| 26 | 28 | |
| 27 | 29 | export async function generateMetadata({ params }: { params: Promise<{ slug: string }> }): Promise<Metadata> { |
@@ -103,6 +105,32 @@ export default async function Page({ params }: { params: Promise<{ slug: string | ||
| 103 | 105 | </div> |
| 104 | 106 | </section> |
| 105 | 107 | |
| 108 | + {d.demographics && ( | |
| 109 | + <section> | |
| 110 | + <SectionHeader eyebrow="Démographie" title="Recensement 2021" description="Statistique Canada, agrégé aux limites 2026 par intersection des aires de diffusion agrégées. La part francophone alimente le swing par segment du modèle." action={<SourceBadge source={{ name: "Statistique Canada — Recensement 2021 (profil ADA + limites)", url: "https://www12.statcan.gc.ca/census-recensement/2021/dp-pd/prof/details/download-telecharger.cfm?Lang=F", detail: `${d.demographics.adaCount} aires de diffusion agrégées intersectées`, tier: "Officiel (StatCan) · agrégation QC26" }} />} /> | |
| 111 | + <div className="grid grid-cols-2 md:grid-cols-4 lg:grid-cols-7 gap-3"> | |
| 112 | + {([ | |
| 113 | + ["Population 2021", int(d.demographics.population2021), ""], | |
| 114 | + ["Langue maternelle française", d.demographics.pctFrenchMt, " %"], | |
| 115 | + ["Langue maternelle anglaise", d.demographics.pctEnglishMt, " %"], | |
| 116 | + ["Autres langues maternelles", d.demographics.pctOtherMt, " %"], | |
| 117 | + ["Immigrant·es", d.demographics.pctImmigrants, " %"], | |
| 118 | + ["Minorités visibles", d.demographics.pctVisibleMinority, " %"], | |
| 119 | + ["65 ans et plus", d.demographics.pct65Plus, " %"], | |
| 120 | + ["Âge médian", d.demographics.medianAge, " ans"], | |
| 121 | + ["Revenu médian des ménages", d.demographics.medianHouseholdIncome ? int(d.demographics.medianHouseholdIncome) : null, " $"], | |
| 122 | + ["Propriétaires", d.demographics.pctOwner, " %"], | |
| 123 | + ["Baccalauréat ou plus (25-64)", d.demographics.pctBachelorPlus, " %"], | |
| 124 | + ["Taux de chômage", d.demographics.unemploymentRate, " %"], | |
| 125 | + ["Densité", d.demographics.density ? int(d.demographics.density) : null, " hab./km²"], | |
| 126 | + ["Regroupement sondeurs", d.pollRegion === "montreal_rmr" ? "Grand Montréal" : d.pollRegion === "quebec_rmr" ? "Région de Québec" : "Reste du Québec", ""], | |
| 127 | + ] as [string, string | number | null, string][]).map(([label, value, suffix]) => ( | |
| 128 | + <div key={label} className="card p-3"><div className="text-[11px] text-ink-3 leading-tight">{label}</div><div className="num font-bold text-[18px] mt-1">{value === null || value === undefined ? "—" : typeof value === "number" ? value.toLocaleString("fr-CA", { maximumFractionDigits: 1 }) : value}<span className="text-[12px] font-medium text-ink-2">{value === null ? "" : suffix}</span></div></div> | |
| 129 | + ))} | |
| 130 | + </div> | |
| 131 | + </section> | |
| 132 | + )} | |
| 133 | + | |
| 106 | 134 | {years.length > 0 && ( |
| 107 | 135 | <section> |
| 108 | 136 | <SectionHeader eyebrow="Historique" title="Résultats officiels" description={d.historyNote} /> |
modified
web/src/app/methodologie/page.tsx
+6 −2
@@ -30,7 +30,11 @@ export default async function Page() { | ||
| 30 | 30 | |
| 31 | 31 | <Section id="temporel" title="Modèle temporel"><p>L'état latent de l'opinion est estimé par un <b>filtre de Kalman</b> (marche aléatoire quotidienne, écart-type 0,012 dans l'espace logit) avec lissage de Rauch–Tung–Striebel. La variance d'observation de chaque sondage vient de sa taille effective (n divisé par un design effect selon le mode : web 1,30, IVR 1,40, téléphone 1,20) et d'une erreur non-échantillonnale de 1,5 point (« total survey error »). La projection au jour du vote ajoute une dérive proportionnelle aux jours restants (inflation 1,5) et une <b>erreur d'industrie</b> corrélée entre partis (σ ≈ 2,2 pt), calibrée sur l'écart entre sondages finaux et résultats en 2018 et 2022 : les firmes se trompent ensemble, et le modèle le sait.</p></Section> |
| 32 | 32 | |
| 33 | − <Section id="regional" title="Projection régionale"><p>Version 1 : le swing est <b>proportionnel</b> et uniforme à l'échelle du Québec, puis perturbé par une erreur régionale (σ 0,12 logit) indépendante entre 17 régions (classification QC26 alignée sur les régions administratives). Les ventilations régionales publiées par les firmes sont archivées (visibles dans le Poll Inspector) mais n'entrent pas encore dans le modèle ; c'est la prochaine étape du changelog.</p></Section> | |
| 33 | + <Section id="regional" title="Projection régionale et swing par segments"><p>Le swing de base est <b>proportionnel</b> à l'échelle du Québec. Depuis la version 1.1, il est corrigé par un <b>excès de swing par segment</b> tiré des ventilations publiées par les firmes dans les 90 derniers jours (pondérées récence × √n) :</p> | |
| 34 | + <ul className="list-disc pl-5 space-y-1"><li><b>Langue</b> : soutien chez les francophones et les non-francophones aujourd'hui, comparé à la référence 2022 (sondages Léger de fin de campagne 2022, documents archivés, recalés sur le résultat officiel). L'écart entre le swing du segment et le swing national est appliqué à chaque circonscription selon sa part de langue maternelle française au <b>Recensement 2021</b> (Statistique Canada, aires de diffusion agrégées intersectées avec les limites 2026), avec un rétrécissement κ = 0,6.</li> | |
| 35 | + <li><b>Région</b> : RMR de Montréal, RMR de Québec, reste du Québec, comparés aux résultats 2022 agrégés selon le même regroupement (affiné circonscription par circonscription), κ = 0,5.</li> | |
| 36 | + <li><b>Calibration</b> : à chaque tirage, la moyenne des parts pondérée par les électeurs inscrits est ramenée au vote national simulé.</li></ul> | |
| 37 | + <p>Une erreur régionale (σ 0,12 logit) indépendante entre 17 régions s'ajoute. La régression écologique des résultats 2022 sur la part francophone est publiée dans l'API à titre de diagnostic mais n'est pas utilisée : son extrapolation aux circonscriptions sans francophones est instable.</p></Section> | |
| 34 | 38 | |
| 35 | 39 | <Section id="circonscriptions" title="Modèle de circonscription"><p>Le point de départ est le <b>résultat 2022 transposé sur la carte 2026</b> : les résultats officiels par section de vote (Élections Québec) sont affectés à la circonscription 2026 contenant le centroïde de la section, puis les totaux officiels de chaque circonscription 2022 sont répartis selon cette distribution (hypothèse : les votes par anticipation suivent les votes du jour). Cette transposition reproduit exactement les chiffres publiés lors de la refonte de la carte (CAQ 93, PLQ 20, QS 11, PQ 3). Pour chaque simulation, la part du parti p dans la circonscription r vaut baseline<sub>r,p</sub> × (vote national simulé<sub>p</sub> / vote national 2022<sub>p</sub>), multipliée par les erreurs régionale et locale (σ 0,20 logit) et par un effet candidat : −0,06 logit (≈ 1,3 pt) quand le·la député·e sortant·e ne se représente pas. Un parti sans candidature n'est pénalisé que lorsque sa liste est quasi complète (≥ 90 % des circonscriptions), pour ne pas interpréter un retard d'enregistrement.</p></Section> |
| 36 | 40 | |
@@ -42,7 +46,7 @@ export default async function Page() { | ||
| 42 | 46 | |
| 43 | 47 | <Section id="boussole" title="Boussole"><p>48 énoncés sur 6 axes. Les positions des partis sont dérivées exclusivement des propositions extraites de leurs documents officiels (page + citation vérifiée dans le texte) : un modèle de langage situe le parti sur l'échelle −2…+2 <i>uniquement</i> si les propositions le permettent, sinon la position est « introuvable » et la question ne compte pas pour ce parti. La concordance utilisateur–parti est 1 − |u − p|/4, pondérée par l'importance. Les réponses restent dans le navigateur.</p></Section> |
| 44 | 48 | |
| 45 | − <Section id="sources" title="Sources"><ul className="list-disc pl-5 space-y-1"><li><b>Élections Québec</b> (données ouvertes officielles) : liste et géométrie des 127 circonscriptions 2026, électeurs inscrits, candidatures, partis autorisés et chef·fes, résultats 2018 et 2022 par circonscription et par section de vote, résultats du soir du vote.</li><li><b>Firmes de sondage</b> : documents originaux (Léger, Pallas Data, Synopsis, Angus Reid, Mainstreet, SEGMA, Abacus, Research Co., etc.).</li><li><b>Sites officiels des partis</b> : plateformes, programmes, engagements, cadres financiers, communiqués — versionnés.</li><li>Un index tiers (Wikipédia) sert uniquement à amorcer la liste des sondages et à repérer les documents originaux ; il est remplacé dès que le document est archivé.</li></ul></Section> | |
| 49 | + <Section id="sources" title="Sources"><ul className="list-disc pl-5 space-y-1"><li><b>Élections Québec</b> (données ouvertes officielles) : liste et géométrie des 127 circonscriptions 2026, électeurs inscrits, candidatures, partis autorisés et chef·fes, résultats 2018 et 2022 par circonscription et par section de vote, résultats du soir du vote.</li><li><b>Firmes de sondage</b> : documents originaux (Léger, Pallas Data, Synopsis, Angus Reid, Mainstreet, SEGMA, Abacus, Research Co., etc.).</li><li><b>Sites officiels des partis</b> : plateformes, programmes, engagements, cadres financiers, communiqués — versionnés.</li><li><b>Statistique Canada</b> : Recensement de la population 2021, profil des aires de diffusion agrégées et fichier des limites, agrégés aux 127 circonscriptions (langue, immigration, âge, revenu, tenure, scolarité, chômage).</li><li>Un index tiers (Wikipédia) sert uniquement à amorcer la liste des sondages et à repérer les documents originaux ; il est remplacé dès que le document est archivé.</li></ul></Section> | |
| 46 | 50 | |
| 47 | 51 | <Section id="limites" title="Limites"><ul className="list-disc pl-5 space-y-1"><li>Le swing est uniforme ; les dynamiques régionales spécifiques (Québec, Montréal, régions) ne sont captées que par un bruit régional, pas encore par les ventilations des sondages.</li><li>Les effets candidats se limitent au retrait du sortant ; notoriété locale et vedettes ne sont pas modélisées.</li><li>L'erreur d'industrie est calibrée sur deux élections seulement.</li><li>Le nombre de sondages par firme est faible pour plusieurs maisons : leurs effets maison ne sont pas estimés (aucune correction appliquée).</li><li>L'extraction documentaire dépend d'un modèle de langage ; les propositions à faible confiance sont mises en file de révision et non publiées.</li></ul><div className="mt-4"><Disclaimer /></div></Section> |
| 48 | 52 | </article> |
| 49 | 53 | |