Vague 3 : avis/notes sans clé API + menus livraison + permis d alcool
- yelp-scrape : notes/avis Yelp SANS clé (1 page de recherche yelp.ca/resto via Scrapfly ASP, cache Apollo embarqué) -> details.yelp, croisement conservateur nom+civique+ville, budget 80 pages/cycle, re-visite 30 j, indépendants d abord. Yelp Fusion (yelp) reprend la main si clé fournie. - ubereats : menus livraison + notes via sitemaps publics (index local 37 087 slugs CA, 30 j) + __REACT_QUERY_STATE__ des pages resto ; croisement conservateur (téléphone / postal+civique / GPS<120 m + nom) avec restos EXISTANTS seulement, menus en price_context delivery, budget 60 pages/cycle. - permits (racj) : permis d alcool en vigueur (Données Québec, CC-BY 4.0), regroupés par établissement et croisés comme MAPAQ -> details.permis_alcool (1 924 restos croisés au premier run). - /api/stats : with_yelp_rating, with_ubereats, with_alcohol_permit ; ingest.enrich appelle permits.sync (guard 6 j) ; registre sources.json (9) + docs générées ; 73 tests verts (9 nouveaux, tout hors ligne). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
22 changed files +1,480 −10
modified
.gitignore
+1 −0
@@ -14,3 +14,4 @@ frontend/node_modules/ | ||
| 14 | 14 | frontend/dist/ |
| 15 | 15 | frontend/tsconfig.tsbuildinfo |
| 16 | 16 | .DS_Store |
| 17 | +data/ubereats-stores.json | |
modified
data/sources.json
+39 −0
@@ -282,6 +282,45 @@ | ||
| 282 | 282 | "acces_legal": "API officielle avec clé, Display Requirements Yelp : attribution affichée avec la note.", |
| 283 | 283 | "cadence": "hebdomadaire, re-vérification 30 jours", |
| 284 | 284 | "status": "clé requise (YELP_API_KEY absente de .env — connecteur livré, SkipSource)" |
| 285 | + }, | |
| 286 | + { | |
| 287 | + "id": "yelp-scrape", | |
| 288 | + "name": "Yelp (avis et notes, sans clé)", | |
| 289 | + "url": "https://www.yelp.ca", | |
| 290 | + "platform": "yelp", | |
| 291 | + "connector": "yelp-scrape", | |
| 292 | + "tier": 5, | |
| 293 | + "price_context": "aucun", | |
| 294 | + "extraction": "Une page de RECHERCHE yelp.ca par resto (find_desc=nom&find_loc=ville) via Scrapfly ASP : le cache Apollo embarqué (scripts data-apollo-state) expose alias, note, nb d'avis, fourchette de prix, catégories, ville et civique par commerce -> details.yelp (mêmes clés que Yelp Fusion). Croisement conservateur : nom similaire + même numéro civique + ville compatible ; ambigu = ignoré. Budget YELP_SCRAPE_BUDGET (80 pages/cycle), re-visite 30 j (hit comme miss), priorité aux restos avec menu. N'émet aucune fiche. Le connecteur `yelp` (API Fusion) reprend la main dès qu'une YELP_API_KEY existe.", | |
| 295 | + "acces_legal": "Pages publiques yelp.ca, extraction minimale (note agrégée + lien fiche), attribution Yelp affichée avec la note (mêmes conditions d'affichage que Fusion). Anti-bot réel : Scrapfly ASP, ~30 crédits/page — budget serré. À remplacer par l'API Fusion (clé gratuite) dès que possible (CLAUDE.md §15).", | |
| 296 | + "cadence": "quotidien par petits lots (80 pages), re-vérification 30 jours", | |
| 297 | + "status": "actif" | |
| 298 | + }, | |
| 299 | + { | |
| 300 | + "id": "ubereats", | |
| 301 | + "name": "Uber Eats (menus livraison + notes)", | |
| 302 | + "url": "https://www.ubereats.com", | |
| 303 | + "platform": "ubereats", | |
| 304 | + "connector": "ubereats", | |
| 305 | + "tier": 2, | |
| 306 | + "price_context": "delivery", | |
| 307 | + "extraction": "Découverte par sitemaps publics (robots.txt -> sitemap-store-*.xml.gz, URLs /ca/store/, index local data/ubereats-stores.json rafraîchi 30 j). Chaque page resto embarque __REACT_QUERY_STATE__ (getStoreV1) : menu complet (sections/items/prix en cents), note, adresse, téléphone, GPS. Croisement conservateur avec un resto EXISTANT (téléphone identique OU postal+civique OU GPS<120 m + nom similaire) — n'émet aucune fiche. Menu -> table menus en price_context delivery (prix majorés ~25-30 %), note -> details.ubereats. Budget UBEREATS_BUDGET (60 pages/cycle), verdicts en cache, re-visite 30 j.", | |
| 308 | + "acces_legal": "Palier 2 (CLAUDE.md §9, §15) : CGU restrictives, Cloudflare — Scrapfly ASP (~1 crédit/page constaté). Usage d'appoint minimal : menus pour restos sans source dine-in/takeout, prix TOUJOURS étiquetés delivery, lien vers la fiche source conservé. Préférer UEAT/site du resto dès que disponible.", | |
| 309 | + "cadence": "quotidien par petits lots (60 pages), re-vérification 30 jours", | |
| 310 | + "status": "actif" | |
| 311 | + }, | |
| 312 | + { | |
| 313 | + "id": "racj", | |
| 314 | + "name": "RACJ — Permis d'alcool en vigueur", | |
| 315 | + "url": "https://www.donneesquebec.ca/recherche/dataset/racj-alcool-detaillant", | |
| 316 | + "platform": "donnees-quebec", | |
| 317 | + "connector": "permits (module d'enrichissement, pas de fiches)", | |
| 318 | + "tier": 5, | |
| 319 | + "price_context": "aucun", | |
| 320 | + "extraction": "CSV ouvert racj-alcool-detaillant.csv (Données Québec, registre des permis de détaillant d'alcool en vigueur) : regroupement par établissement (NoEtablissement) puis croisement CONSERVATEUR avec les restaurants (nom normalisé + code postal/ville+civique, ou postal + civique + similarité de nom — mêmes règles que MAPAQ). Catégories de permis (Bar, Restaurant pour vendre/servir…), capacité et titulaire -> details.permis_alcool.", | |
| 321 | + "acces_legal": "Données ouvertes du gouvernement du Québec (RACJ), licence CC-BY 4.0 — attribution RACJ/Données Québec affichée avec l'information de permis.", | |
| 322 | + "cadence": "hebdomadaire (guard 6 jours dans ingest.enrich)", | |
| 323 | + "status": "actif" | |
| 285 | 324 | } |
| 286 | 325 | ] |
| 287 | 326 | } |
modified
data/ueat-discovered.json
+72 −0
@@ -87,6 +87,78 @@ | ||
| 87 | 87 | { |
| 88 | 88 | "key": "e4de92c1-12d0-49db-bb59-7baf0202faf8", |
| 89 | 89 | "label": "Nora Gray (découvert via noragray.com)" |
| 90 | + }, | |
| 91 | + { | |
| 92 | + "key": "5f4f560b-830d-4af0-892d-b15be0ed653c", | |
| 93 | + "label": "Lucille’s Oyster Dive (découvert via lucillesoyster.com)" | |
| 94 | + }, | |
| 95 | + { | |
| 96 | + "key": "12fe0fc4-f986-40e3-abbc-670c01f37247", | |
| 97 | + "label": "Jun I (découvert via juni.ca)" | |
| 98 | + }, | |
| 99 | + { | |
| 100 | + "key": "484faa4f-19f5-435b-9428-d5c8fd439f75", | |
| 101 | + "label": "La Boîte À Huîtres (découvert via laboiteauxhuitres.com)" | |
| 102 | + }, | |
| 103 | + { | |
| 104 | + "key": "791c777b-4e38-4fa6-850d-4e9f322a90d8", | |
| 105 | + "label": "Sushi Sakura (découvert via sushisakura.ca)" | |
| 106 | + }, | |
| 107 | + { | |
| 108 | + "key": "23c2c440-9ad3-4d3f-b198-3e2d565f7999", | |
| 109 | + "label": "Boîte Geisha (découvert via boitegeisha.com)" | |
| 110 | + }, | |
| 111 | + { | |
| 112 | + "key": "d94f44b9-67ac-42f1-99e5-a56040210f87", | |
| 113 | + "label": "Monza (découvert via restaurantmonza.com)" | |
| 114 | + }, | |
| 115 | + { | |
| 116 | + "key": "c147a584-441d-4db2-9ab5-384c250ab925", | |
| 117 | + "label": "Barranco (découvert via barrancorestaurant.com)" | |
| 118 | + }, | |
| 119 | + { | |
| 120 | + "key": "403f3002-0bae-4f8c-9ebb-391a2e69caeb", | |
| 121 | + "label": "Seau de Crabe (découvert via seaudecrabe.ca)" | |
| 122 | + }, | |
| 123 | + { | |
| 124 | + "key": "2881e6e6-58fd-4b26-b8b9-8cfc4569fa14", | |
| 125 | + "label": "Shūshū — Mile-End (découvert via shushumtl.com)" | |
| 126 | + }, | |
| 127 | + { | |
| 128 | + "key": "073dc1c3-04eb-49f7-8c06-7c71c572ddf7", | |
| 129 | + "label": "L'Œil du dragon (découvert via oeildudragon.com)" | |
| 130 | + }, | |
| 131 | + { | |
| 132 | + "key": "3ad538c1-56c1-43e4-a308-fb0ef963cf07", | |
| 133 | + "label": "Ristorante Giorgio (découvert via giorgio.ca)" | |
| 134 | + }, | |
| 135 | + { | |
| 136 | + "key": "89422724-84db-4302-ac12-ebf532434604", | |
| 137 | + "label": "Pizzeria Zac (découvert via chezzacpizza.com)" | |
| 138 | + }, | |
| 139 | + { | |
| 140 | + "key": "ab96cf7a-cf9c-4178-8d52-9b3729709aec", | |
| 141 | + "label": "La galette libanaise (découvert via lagalettelibanaise.com)" | |
| 142 | + }, | |
| 143 | + { | |
| 144 | + "key": "4d9578ba-f905-472e-90fa-e39e184b0597", | |
| 145 | + "label": "Stratos Pizzeria (découvert via stratos-pizzeria.com)" | |
| 146 | + }, | |
| 147 | + { | |
| 148 | + "key": "51236cf5-13f5-4fb4-a08c-d30b31519709", | |
| 149 | + "label": "Umai Soo She (découvert via umaisooshe.ca)" | |
| 150 | + }, | |
| 151 | + { | |
| 152 | + "key": "388d78df-d9c9-4aaa-a974-7b8d696312eb", | |
| 153 | + "label": "Torii Izakaya (découvert via toriiizakaya.ca)" | |
| 154 | + }, | |
| 155 | + { | |
| 156 | + "key": "7ec72df5-e520-4335-9935-e2be36db5761", | |
| 157 | + "label": "Küto (découvert via kuto.ca)" | |
| 158 | + }, | |
| 159 | + { | |
| 160 | + "key": "d42e5dd2-815c-4803-930d-ae0ee95b1e4b", | |
| 161 | + "label": "Bol de Vie (découvert via boldevie.ca)" | |
| 90 | 162 | } |
| 91 | 163 | ] |
| 92 | 164 | } |
| \ No newline at end of file | ||
modified
docs/connecteurs/INDEX.md
+5 −2
@@ -1,8 +1,8 @@ | ||
| 1 | 1 | # Resto-Ka — Index des connecteurs |
| 2 | 2 | |
| 3 | −_Généré automatiquement par `scripts/gen_connector_docs.py` le 2026-08-18 03:30 — ne pas éditer à la main, régénérer._ | |
| 3 | +_Généré automatiquement par `scripts/gen_connector_docs.py` le 2026-08-18 06:06 — ne pas éditer à la main, régénérer._ | |
| 4 | 4 | |
| 5 | −**6 sources au registre** · **14989 restos actifs** en BD. | |
| 5 | +**9 sources au registre** · **14989 restos actifs** en BD. | |
| 6 | 6 | |
| 7 | 7 | | Source | Nom | Tier | Type d'accès | Backend | Actifs/Total | GPS | Horaires | Cuisines | État | Dernier sync OK | |
| 8 | 8 | |---|---|---|---|---|---|---|---|---|---|---| |
@@ -12,3 +12,6 @@ _Généré automatiquement par `scripts/gen_connector_docs.py` le 2026-08-18 03: | ||
| 12 | 12 | | [`sthubert`](sthubert.md) | St-Hubert (commande en ligne) | T6 | — | — | 0/0 | — | — | — | à faire | — | |
| 13 | 13 | | [`mapaq`](mapaq.md) | MAPAQ — Inspections alimentaires (conda… | T5 | jeu de données ouvert (CSV, Données Qué… | direct | 0/0 | — | — | — | actif | 2026-08-18 02:53 | |
| 14 | 14 | | [`yelp`](yelp.md) | Yelp Fusion (avis et notes) | T5 | API JSON | direct | 0/0 | — | — | — | clé requise | — | |
| 15 | +| [`yelp-scrape`](yelp-scrape.md) | Yelp (avis et notes, sans clé) | T5 | pages HTML (rendu serveur) | Scrapfly | 0/0 | — | — | — | actif | — | | |
| 16 | +| [`ubereats`](ubereats.md) | Uber Eats (menus livraison + notes) | T2 | sitemap XML + pages HTML | Scrapfly | 0/0 | — | — | — | actif | — | | |
| 17 | +| [`racj`](racj.md) | RACJ — Permis d'alcool en vigueur | T5 | jeu de données ouvert (CSV, Données Qué… | direct | 0/0 | — | — | — | actif | 2026-08-18 05:47 | | |
modified
docs/connecteurs/mapaq.md
+1 −1
@@ -1,6 +1,6 @@ | ||
| 1 | 1 | # MAPAQ — Inspections alimentaires (condamnations) — connecteur `mapaq` |
| 2 | 2 | |
| 3 | −_Fiche générée automatiquement par `scripts/gen_connector_docs.py` le 2026-08-18 03:30 — ne pas éditer à la main, régénérer._ | |
| 3 | +_Fiche générée automatiquement par `scripts/gen_connector_docs.py` le 2026-08-18 06:06 — ne pas éditer à la main, régénérer._ | |
| 4 | 4 | |
| 5 | 5 | **État : actif** · Palier (tier) : 5 · Backend : direct |
| 6 | 6 | |
modified
docs/connecteurs/osm.md
+3 −1
@@ -1,6 +1,6 @@ | ||
| 1 | 1 | # OpenStreetMap (découverte) — connecteur `osm` |
| 2 | 2 | |
| 3 | −_Fiche générée automatiquement par `scripts/gen_connector_docs.py` le 2026-08-18 03:30 — ne pas éditer à la main, régénérer._ | |
| 3 | +_Fiche générée automatiquement par `scripts/gen_connector_docs.py` le 2026-08-18 06:06 — ne pas éditer à la main, régénérer._ | |
| 4 | 4 | |
| 5 | 5 | **État : actif** · Palier (tier) : 4 · Backend : direct · Restos actifs : 13969/13969 |
| 6 | 6 | |
@@ -47,6 +47,8 @@ Le connecteur alimente la table `restaurants` (et `menus` le cas échéant). Com | ||
| 47 | 47 | | `images` | Photos (JSON) | 0 % | — | |
| 48 | 48 | | `url` | URL de la fiche source | 100 % | https://www.openstreetmap.org/node/112637139 | |
| 49 | 49 | |
| 50 | +**Menus** : 13 menus rattachés (728 items, 1 contexte(s) de prix, dernière capture 2026-08-18T10:05:33Z) — table `menus` (sections → items → options, prix CAD). | |
| 51 | + | |
| 50 | 52 | ## Fréquence & budget |
| 51 | 53 | |
| 52 | 54 | - **Cadence déclarée (registre)** : mensuel |
added
docs/connecteurs/racj.md
+54 −0
@@ -0,0 +1,54 @@ | ||
| 1 | +# RACJ — Permis d'alcool en vigueur — connecteur `racj` | |
| 2 | + | |
| 3 | +_Fiche générée automatiquement par `scripts/gen_connector_docs.py` le 2026-08-18 06:06 — ne pas éditer à la main, régénérer._ | |
| 4 | + | |
| 5 | +**État : actif** · Palier (tier) : 5 · Backend : direct | |
| 6 | + | |
| 7 | +## Description de la source | |
| 8 | + | |
| 9 | +Enrichissement RACJ — permis d'alcool en vigueur (Régie des alcools, des courses et des jeux, Données Québec, licence CC-BY 4.0). Télécharge le registre CSV des permis de détaillant d'alcool (bar, restaurant pour vendre/servir…), regroupe par établissement et croise de façon CONSERVATRICE avec les restaurants (mêmes règles que les inspections MAPAQ : nom exact + lieu, ou postal + civique + nom similaire — jamais fusionner deux établissements). Résultat dans details.permis_alcool (catégories de permis, capacité, titulaire). Gratuit, sans clé API. Cadence hebdomadaire (guard 6 jours). | |
| 10 | + | |
| 11 | +- **Plateforme** : donnees-quebec · **Site** : https://www.donneesquebec.ca/recherche/dataset/racj-alcool-detaillant | |
| 12 | +- **Extraction (registre)** : CSV ouvert racj-alcool-detaillant.csv (Données Québec, registre des permis de détaillant d'alcool en vigueur) : regroupement par établissement (NoEtablissement) puis croisement CONSERVATEUR avec les restaurants (nom normalisé + code postal/ville+civique, ou postal + civique + similarité de nom — mêmes règles que MAPAQ). Catégories de permis (Bar, Restaurant pour vendre/servir…), capacité et titulaire -> details.permis_alcool. | |
| 13 | +- **Contexte de prix** : aucun | |
| 14 | +- **Module** : `restoka/permits.py` — `(module fonctionnel, pas de classe)` | |
| 15 | + | |
| 16 | +## Accès | |
| 17 | + | |
| 18 | +- **Type d'accès** : jeu de données ouvert (CSV, Données Québec) | |
| 19 | +- **Endpoint de base** : https://www.donneesquebec.ca/recherche/dataset/d817c9f7-76c7-44af-882d-0d673056ef86/resource/6b69360c-af8d-4c57-b5bf-1c38d5461de3/download/racj-alcool-detaillant.csv | |
| 20 | +- **URLs du module** : https://www.donneesquebec.ca/recherche/dataset/ | |
| 21 | +- **Pagination** : réponse unique (pas de pagination) | |
| 22 | +- **Backend anti-bot / rendu** : requests direct (session UA RestoKaBot, throttling poli) | |
| 23 | +- **Authentification** : aucune — accès public/anonyme | |
| 24 | + | |
| 25 | +## Champs récupérés → schéma cible | |
| 26 | + | |
| 27 | +Aucune fiche en BD pour cette source (connecteur en attente ou clé manquante) — schéma cible : table `restaurants` + `menus`. | |
| 28 | + | |
| 29 | +## Fréquence & budget | |
| 30 | + | |
| 31 | +- **Cadence déclarée (registre)** : hebdomadaire (guard 6 jours dans ingest.enrich) | |
| 32 | +- **Cadence observée** (médiane sync_log) : — | |
| 33 | +- **Dernier passage OK** : 2026-08-18 05:47 — 18686 trouvées, +1924 / ~0 / -0 | |
| 34 | + | |
| 35 | +## Volumétrie & complétude | |
| 36 | + | |
| 37 | +- Aucune donnée en BD pour cette source. | |
| 38 | +- **Runs journalisés (60 derniers)** : 1, dont 0 en erreur | |
| 39 | + | |
| 40 | +## Erreurs connues & dépannage | |
| 41 | + | |
| 42 | +Aucune erreur dans les 60 derniers runs journalisés. | |
| 43 | + | |
| 44 | +Rejouer la source seule : `python3 run.py sync racj` · vérifier `sync_log` (`SELECT * FROM sync_log WHERE source='racj' ORDER BY ts DESC LIMIT 5;`). | |
| 45 | + | |
| 46 | +## Licence, attribution & conditions | |
| 47 | + | |
| 48 | +- **Cadre d'accès (registre, `acces_legal`)** : Données ouvertes du gouvernement du Québec (RACJ), licence CC-BY 4.0 — attribution RACJ/Données Québec affichée avec l'information de permis. | |
| 49 | +- **Licence** : donnée ouverte CC-BY 4.0 (Données Québec) — mention de la source « MAPAQ / Données Québec » affichée. | |
| 50 | +- Retrait sur demande : contact@spboucher.ai. | |
| 51 | + | |
| 52 | +## Historique | |
| 53 | + | |
| 54 | +- 2026-08-18 — vague d'enrichissement : standardisation de la documentation des connecteurs (fiche générée par `scripts/gen_connector_docs.py`). | |
modified
docs/connecteurs/site-resto.md
+1 −1
@@ -1,6 +1,6 @@ | ||
| 1 | 1 | # Sites web des restos (menus maison) — connecteur `site-resto` |
| 2 | 2 | |
| 3 | −_Fiche générée automatiquement par `scripts/gen_connector_docs.py` le 2026-08-18 03:30 — ne pas éditer à la main, régénérer._ | |
| 3 | +_Fiche générée automatiquement par `scripts/gen_connector_docs.py` le 2026-08-18 06:06 — ne pas éditer à la main, régénérer._ | |
| 4 | 4 | |
| 5 | 5 | **État : actif** · Palier (tier) : 3 · Backend : Firecrawl · Restos actifs : 86/86 |
| 6 | 6 | |
modified
docs/connecteurs/sthubert.md
+1 −1
@@ -1,6 +1,6 @@ | ||
| 1 | 1 | # St-Hubert (commande en ligne) — connecteur `sthubert` |
| 2 | 2 | |
| 3 | −_Fiche générée automatiquement par `scripts/gen_connector_docs.py` le 2026-08-18 03:30 — ne pas éditer à la main, régénérer._ | |
| 3 | +_Fiche générée automatiquement par `scripts/gen_connector_docs.py` le 2026-08-18 06:06 — ne pas éditer à la main, régénérer._ | |
| 4 | 4 | |
| 5 | 5 | **État : à faire** · Palier (tier) : 6 · Backend : — |
| 6 | 6 | |
added
docs/connecteurs/ubereats.md
+55 −0
@@ -0,0 +1,55 @@ | ||
| 1 | +# Uber Eats (menus livraison + notes) — connecteur `ubereats` | |
| 2 | + | |
| 3 | +_Fiche générée automatiquement par `scripts/gen_connector_docs.py` le 2026-08-18 06:06 — ne pas éditer à la main, régénérer._ | |
| 4 | + | |
| 5 | +**État : actif** · Palier (tier) : 2 · Backend : Scrapfly | |
| 6 | + | |
| 7 | +## Description de la source | |
| 8 | + | |
| 9 | +Connecteur d'ENRICHISSEMENT Uber Eats (menus livraison + notes) — SANS clé API. Découverte par les sitemaps publics (robots.txt -> sitemap-store-*.xml.gz, URLs /ca/store/<slug>/<uuid>, index local data/ubereats-stores.json rafraîchi tous les 30 j). Chaque page de resto embarque un état React Query (__REACT_QUERY_STATE__) avec le menu complet (sections + items + prix en cents), la note Uber Eats, l'adresse, le téléphone et les GPS. Croisement CONSERVATEUR avec un resto EXISTANT seulement (téléphone identique, OU code postal + numéro civique identiques, OU GPS <120 m + nom similaire) : n'émet AUCUNE fiche. Menu -> table menus en price_context `delivery` (prix majorés ~25-30 %, CLAUDE.md §6.2), note -> details.ubereats. Économe : Scrapfly ASP (1 page ≈ 1 crédit constaté), budget UBEREATS_BUDGET (défaut 60) pages resto/cycle, verdict par magasin mis en cache (detail_cache), re-visite 30 j. Conformité §15 : extraction minimale, usage d'appoint (restos sans menu ailleurs), prix toujours étiquetés `delivery`, lien source conservé. | |
| 10 | + | |
| 11 | +- **Plateforme** : ubereats · **Site** : https://www.ubereats.com | |
| 12 | +- **Extraction (registre)** : Découverte par sitemaps publics (robots.txt -> sitemap-store-*.xml.gz, URLs /ca/store/, index local data/ubereats-stores.json rafraîchi 30 j). Chaque page resto embarque __REACT_QUERY_STATE__ (getStoreV1) : menu complet (sections/items/prix en cents), note, adresse, téléphone, GPS. Croisement conservateur avec un resto EXISTANT (téléphone identique OU postal+civique OU GPS<120 m + nom similaire) — n'émet aucune fiche. Menu -> table menus en price_context delivery (prix majorés ~25-30 %), note -> details.ubereats. Budget UBEREATS_BUDGET (60 pages/cycle), verdicts en cache, re-visite 30 j. | |
| 13 | +- **Contexte de prix** : delivery | |
| 14 | +- **Module** : `restoka/connectors/ubereats.py` — `UberEatsConnector` | |
| 15 | + | |
| 16 | +## Accès | |
| 17 | + | |
| 18 | +- **Type d'accès** : sitemap XML + pages HTML | |
| 19 | +- **Endpoint de base** : https://www.ubereats.com/robots.txt | |
| 20 | +- **URLs du module** : https://www.ubereats.com/robots.txt | |
| 21 | +- **Pagination** : réponse unique (pas de pagination) | |
| 22 | +- **Backend anti-bot / rendu** : Scrapfly (asp + render_js — contournement anti-bot) | |
| 23 | +- **Politesse** : 1.0 s entre requêtes, timeout 60 s, UA `RestoKaBot/1.0 (+https://www.resto-ka.com/bot)` | |
| 24 | +- **Authentification** : aucune — accès public/anonyme | |
| 25 | + | |
| 26 | +## Champs récupérés → schéma cible | |
| 27 | + | |
| 28 | +Aucune fiche en BD pour cette source (connecteur en attente ou clé manquante) — schéma cible : table `restaurants` + `menus`. | |
| 29 | + | |
| 30 | +## Fréquence & budget | |
| 31 | + | |
| 32 | +- **Cadence déclarée (registre)** : quotidien par petits lots (60 pages), re-vérification 30 jours | |
| 33 | +- **Cadence observée** (médiane sync_log) : — | |
| 34 | +- **Caps / budgets du module** : `MAX_CANDIDATES` = 3, `MAX_CONSECUTIVE_FAILURES` = 3, `MAX_GPS_M` = 120.0 | |
| 35 | + | |
| 36 | +## Volumétrie & complétude | |
| 37 | + | |
| 38 | +- Aucune donnée en BD pour cette source. | |
| 39 | +- **Runs journalisés (60 derniers)** : 0, dont 0 en erreur | |
| 40 | + | |
| 41 | +## Erreurs connues & dépannage | |
| 42 | + | |
| 43 | +Aucune erreur dans les 60 derniers runs journalisés. | |
| 44 | + | |
| 45 | +Rejouer la source seule : `python3 run.py sync ubereats` · vérifier `sync_log` (`SELECT * FROM sync_log WHERE source='ubereats' ORDER BY ts DESC LIMIT 5;`). | |
| 46 | + | |
| 47 | +## Licence, attribution & conditions | |
| 48 | + | |
| 49 | +- **Cadre d'accès (registre, `acces_legal`)** : Palier 2 (CLAUDE.md §9, §15) : CGU restrictives, Cloudflare — Scrapfly ASP (~1 crédit/page constaté). Usage d'appoint minimal : menus pour restos sans source dine-in/takeout, prix TOUJOURS étiquetés delivery, lien vers la fiche source conservé. Préférer UEAT/site du resto dès que disponible. | |
| 50 | +- **Scraping / API** : User-Agent identifiable `RestoKaBot/1.0 (+https://www.resto-ka.com/bot; contact@spboucher.ai)`, throttling poli, aucun contournement d'accès ; les fiches pointent vers la source d'origine. | |
| 51 | +- Retrait sur demande : contact@spboucher.ai. | |
| 52 | + | |
| 53 | +## Historique | |
| 54 | + | |
| 55 | +- 2026-08-18 — vague d'enrichissement : standardisation de la documentation des connecteurs (fiche générée par `scripts/gen_connector_docs.py`). | |
modified
docs/connecteurs/ueat.md
+1 −1
@@ -1,6 +1,6 @@ | ||
| 1 | 1 | # UEAT (commande en ligne) — connecteur `ueat` |
| 2 | 2 | |
| 3 | −_Fiche générée automatiquement par `scripts/gen_connector_docs.py` le 2026-08-18 03:30 — ne pas éditer à la main, régénérer._ | |
| 3 | +_Fiche générée automatiquement par `scripts/gen_connector_docs.py` le 2026-08-18 06:06 — ne pas éditer à la main, régénérer._ | |
| 4 | 4 | |
| 5 | 5 | **État : actif** · Palier (tier) : 1 · Backend : direct · Restos actifs : 934/936 |
| 6 | 6 | |
added
docs/connecteurs/yelp-scrape.md
+55 −0
@@ -0,0 +1,55 @@ | ||
| 1 | +# Yelp (avis et notes, sans clé) — connecteur `yelp-scrape` | |
| 2 | + | |
| 3 | +_Fiche générée automatiquement par `scripts/gen_connector_docs.py` le 2026-08-18 06:06 — ne pas éditer à la main, régénérer._ | |
| 4 | + | |
| 5 | +**État : actif** · Palier (tier) : 5 · Backend : Scrapfly | |
| 6 | + | |
| 7 | +## Description de la source | |
| 8 | + | |
| 9 | +Connecteur d'ENRICHISSEMENT Yelp SANS clé API (avis/notes par scraping léger). Une seule page de RECHERCHE yelp.ca par resto (find_desc=<nom>&find_loc=<ville>) via Scrapfly ASP : la page embarque un cache Apollo (scripts data-apollo-state) avec, PAR commerce, alias, nom, note, nombre d'avis, fourchette de prix, catégories et adresse (ville + civique). Croisement CONSERVATEUR : nom similaire + même numéro civique + ville compatible (jamais fusionner deux établissements). Résultat -> details.yelp {rating, review_count, price, categories, url} (mêmes clés que Yelp Fusion — le connecteur officiel `yelp` reprend la main dès qu'une YELP_API_KEY existe). Économe : budget YELP_SCRAPE_BUDGET (défaut 80) pages/cycle, re-visite 30 jours (hit comme miss), priorité aux restos avec menu. Base légale : pages publiques, extraction minimale (note agrégée + lien vers la fiche Yelp, attribution affichée) — voir CLAUDE.md §15. | |
| 10 | + | |
| 11 | +- **Plateforme** : yelp · **Site** : https://www.yelp.ca | |
| 12 | +- **Extraction (registre)** : Une page de RECHERCHE yelp.ca par resto (find_desc=nom&find_loc=ville) via Scrapfly ASP : le cache Apollo embarqué (scripts data-apollo-state) expose alias, note, nb d'avis, fourchette de prix, catégories, ville et civique par commerce -> details.yelp (mêmes clés que Yelp Fusion). Croisement conservateur : nom similaire + même numéro civique + ville compatible ; ambigu = ignoré. Budget YELP_SCRAPE_BUDGET (80 pages/cycle), re-visite 30 j (hit comme miss), priorité aux restos avec menu. N'émet aucune fiche. Le connecteur `yelp` (API Fusion) reprend la main dès qu'une YELP_API_KEY existe. | |
| 13 | +- **Contexte de prix** : aucun | |
| 14 | +- **Module** : `restoka/connectors/yelp_scrape.py` — `YelpScrapeConnector` | |
| 15 | + | |
| 16 | +## Accès | |
| 17 | + | |
| 18 | +- **Type d'accès** : pages HTML (rendu serveur) | |
| 19 | +- **Endpoint de base** : https://www.yelp.ca/search?find_desc={desc}&find_loc={loc} | |
| 20 | +- **URLs du module** : https://www.yelp.ca/search?find_desc={desc}&find_loc={loc} · https://www.yelp.ca/biz/ | |
| 21 | +- **Pagination** : réponse unique (pas de pagination) | |
| 22 | +- **Backend anti-bot / rendu** : Scrapfly (asp + render_js — contournement anti-bot) | |
| 23 | +- **Politesse** : 1.0 s entre requêtes, timeout 60 s, UA `RestoKaBot/1.0 (+https://www.resto-ka.com/bot)` | |
| 24 | +- **Authentification** : aucune — accès public/anonyme | |
| 25 | + | |
| 26 | +## Champs récupérés → schéma cible | |
| 27 | + | |
| 28 | +Aucune fiche en BD pour cette source (connecteur en attente ou clé manquante) — schéma cible : table `restaurants` + `menus`. | |
| 29 | + | |
| 30 | +## Fréquence & budget | |
| 31 | + | |
| 32 | +- **Cadence déclarée (registre)** : quotidien par petits lots (80 pages), re-vérification 30 jours | |
| 33 | +- **Cadence observée** (médiane sync_log) : — | |
| 34 | +- **Caps / budgets du module** : `MAX_CONSECUTIVE_FAILURES` = 3 | |
| 35 | + | |
| 36 | +## Volumétrie & complétude | |
| 37 | + | |
| 38 | +- Aucune donnée en BD pour cette source. | |
| 39 | +- **Runs journalisés (60 derniers)** : 0, dont 0 en erreur | |
| 40 | + | |
| 41 | +## Erreurs connues & dépannage | |
| 42 | + | |
| 43 | +Aucune erreur dans les 60 derniers runs journalisés. | |
| 44 | + | |
| 45 | +Rejouer la source seule : `python3 run.py sync yelp-scrape` · vérifier `sync_log` (`SELECT * FROM sync_log WHERE source='yelp-scrape' ORDER BY ts DESC LIMIT 5;`). | |
| 46 | + | |
| 47 | +## Licence, attribution & conditions | |
| 48 | + | |
| 49 | +- **Cadre d'accès (registre, `acces_legal`)** : Pages publiques yelp.ca, extraction minimale (note agrégée + lien fiche), attribution Yelp affichée avec la note (mêmes conditions d'affichage que Fusion). Anti-bot réel : Scrapfly ASP, ~30 crédits/page — budget serré. À remplacer par l'API Fusion (clé gratuite) dès que possible (CLAUDE.md §15). | |
| 50 | +- **Scraping / API** : User-Agent identifiable `RestoKaBot/1.0 (+https://www.resto-ka.com/bot; contact@spboucher.ai)`, throttling poli, aucun contournement d'accès ; les fiches pointent vers la source d'origine. | |
| 51 | +- Retrait sur demande : contact@spboucher.ai. | |
| 52 | + | |
| 53 | +## Historique | |
| 54 | + | |
| 55 | +- 2026-08-18 — vague d'enrichissement : standardisation de la documentation des connecteurs (fiche générée par `scripts/gen_connector_docs.py`). | |
modified
docs/connecteurs/yelp.md
+1 −1
@@ -1,6 +1,6 @@ | ||
| 1 | 1 | # Yelp Fusion (avis et notes) — connecteur `yelp` |
| 2 | 2 | |
| 3 | −_Fiche générée automatiquement par `scripts/gen_connector_docs.py` le 2026-08-18 03:30 — ne pas éditer à la main, régénérer._ | |
| 3 | +_Fiche générée automatiquement par `scripts/gen_connector_docs.py` le 2026-08-18 06:06 — ne pas éditer à la main, régénérer._ | |
| 4 | 4 | |
| 5 | 5 | **État : clé requise** · Palier (tier) : 5 · Backend : direct |
| 6 | 6 | |
added
restoka/connectors/ubereats.py
+393 −0
@@ -0,0 +1,393 @@ | ||
| 1 | +# ============================================================================== | |
| 2 | +# Author: Simon-Pierre Boucher <contact@spboucher.ai> | |
| 3 | +# File: restoka/connectors/ubereats.py | |
| 4 | +# Desc: Connecteur d'ENRICHISSEMENT Uber Eats (menus livraison + notes) — | |
| 5 | +# SANS clé API. Découverte par les sitemaps publics (robots.txt -> | |
| 6 | +# sitemap-store-*.xml.gz, URLs /ca/store/<slug>/<uuid>, index local | |
| 7 | +# data/ubereats-stores.json rafraîchi tous les 30 j). Chaque page de | |
| 8 | +# resto embarque un état React Query (__REACT_QUERY_STATE__) avec le | |
| 9 | +# menu complet (sections + items + prix en cents), la note Uber Eats, | |
| 10 | +# l'adresse, le téléphone et les GPS. Croisement CONSERVATEUR avec un | |
| 11 | +# resto EXISTANT seulement (téléphone identique, OU code postal + numéro | |
| 12 | +# civique identiques, OU GPS <120 m + nom similaire) : n'émet AUCUNE | |
| 13 | +# fiche. Menu -> table menus en price_context `delivery` (prix majorés | |
| 14 | +# ~25-30 %, CLAUDE.md §6.2), note -> details.ubereats. | |
| 15 | +# | |
| 16 | +# Économe : Scrapfly ASP (1 page ≈ 1 crédit constaté), budget | |
| 17 | +# UBEREATS_BUDGET (défaut 60) pages resto/cycle, verdict par magasin | |
| 18 | +# mis en cache (detail_cache), re-visite 30 j. Conformité §15 : | |
| 19 | +# extraction minimale, usage d'appoint (restos sans menu ailleurs), | |
| 20 | +# prix toujours étiquetés `delivery`, lien source conservé. | |
| 21 | +# ============================================================================== | |
| 22 | +from __future__ import annotations | |
| 23 | + | |
| 24 | +import base64 | |
| 25 | +import datetime | |
| 26 | +import gzip | |
| 27 | +import json | |
| 28 | +import os | |
| 29 | +import re | |
| 30 | +import sys | |
| 31 | +import time | |
| 32 | +import urllib.parse | |
| 33 | +from pathlib import Path | |
| 34 | + | |
| 35 | +from ..inspections import _CIVIC_RE, _name_similar, norm_name | |
| 36 | +from ..normalize import normalize_phone | |
| 37 | +from ..regions import strip_accents | |
| 38 | +from ..schema import Restaurant | |
| 39 | +from .base import BaseConnector, SkipSource | |
| 40 | + | |
| 41 | +ROBOTS_URL = "https://www.ubereats.com/robots.txt" | |
| 42 | +INDEX_PATH = Path(__file__).resolve().parents[2] / "data" / "ubereats-stores.json" | |
| 43 | +INDEX_REFRESH_DAYS = 30 # les sitemaps bougent peu | |
| 44 | +MAX_BUDGET = int(os.environ.get("UBEREATS_BUDGET", "60")) # pages resto/cycle | |
| 45 | +REFRESH_DAYS = 30 # re-visite d'un resto (hit comme miss) | |
| 46 | +MAX_CANDIDATES = 3 # magasins testés par resto, max | |
| 47 | +MAX_CONSECUTIVE_FAILURES = 3 | |
| 48 | +MAX_GPS_M = 120.0 | |
| 49 | + | |
| 50 | +_STATE_RE = re.compile(r"__REACT_QUERY_STATE__\">(.*?)</script>", re.S) | |
| 51 | +_STORE_URL_RE = re.compile(r"/ca/store/([^/]+)/([A-Za-z0-9_-]{20,})") | |
| 52 | + | |
| 53 | + | |
| 54 | +def decode_state(page_html: str) -> dict | None: | |
| 55 | + """Décode le bloc __REACT_QUERY_STATE__ (JSON avec « " » encodé \\u0022 | |
| 56 | + et « \\ » encodé %5C) d'une page Uber Eats.""" | |
| 57 | + m = _STATE_RE.search(page_html) | |
| 58 | + if not m: | |
| 59 | + return None | |
| 60 | + t = m.group(1).strip().replace("\\u0022", '"').replace("%5C", "\\") | |
| 61 | + try: | |
| 62 | + return json.loads(t) | |
| 63 | + except ValueError: | |
| 64 | + return None | |
| 65 | + | |
| 66 | + | |
| 67 | +def store_payload(state: dict) -> dict | None: | |
| 68 | + """Extrait le payload getStoreV1 (fiche resto complète) de l'état.""" | |
| 69 | + for q in (state or {}).get("queries") or []: | |
| 70 | + qk = q.get("queryKey") | |
| 71 | + if isinstance(qk, list) and qk and qk[0] == "getStoreV1": | |
| 72 | + data = (q.get("state") or {}).get("data") | |
| 73 | + if isinstance(data, dict): | |
| 74 | + return data | |
| 75 | + return None | |
| 76 | + | |
| 77 | + | |
| 78 | +def build_menu(store: dict, captured_at: str) -> dict | None: | |
| 79 | + """Menu standard Resto·Ka (CLAUDE.md §5.2) depuis catalogSectionsMap. | |
| 80 | + Prix Uber Eats en cents -> dollars, contexte `delivery` obligatoire.""" | |
| 81 | + csm = store.get("catalogSectionsMap") or {} | |
| 82 | + sections_out: list[dict] = [] | |
| 83 | + seen: set[str] = set() | |
| 84 | + for sections in csm.values(): | |
| 85 | + for sec in sections or []: | |
| 86 | + payload = ((sec.get("payload") or {}) | |
| 87 | + .get("standardItemsPayload") or {}) | |
| 88 | + name = ((payload.get("title") or {}).get("text") or "").strip() | |
| 89 | + items_in = payload.get("catalogItems") or [] | |
| 90 | + if not name or not items_in or name in seen: | |
| 91 | + continue | |
| 92 | + items = [] | |
| 93 | + for it in items_in: | |
| 94 | + title = (it.get("title") or "").strip() | |
| 95 | + if not title: | |
| 96 | + continue | |
| 97 | + price = it.get("price") | |
| 98 | + items.append({ | |
| 99 | + "name": title, | |
| 100 | + "description": (it.get("itemDescription") or "").strip(), | |
| 101 | + "price": round(price / 100.0, 2) | |
| 102 | + if isinstance(price, (int, float)) and price > 0 else None, | |
| 103 | + }) | |
| 104 | + if items: | |
| 105 | + seen.add(name) | |
| 106 | + sections_out.append({"name": name, "items": items}) | |
| 107 | + break # une seule vue du catalogue suffit | |
| 108 | + if not sections_out: | |
| 109 | + return None | |
| 110 | + return { | |
| 111 | + "price_context": "delivery", | |
| 112 | + "price_source": "ubereats", | |
| 113 | + "currency": store.get("currencyCode") or "CAD", | |
| 114 | + "captured_at": captured_at, | |
| 115 | + "sections": sections_out, | |
| 116 | + } | |
| 117 | + | |
| 118 | + | |
| 119 | +def _review_count(store: dict) -> int: | |
| 120 | + raw = str((store.get("rating") or {}).get("reviewCount") or "") | |
| 121 | + m = re.search(r"\d+", raw.replace(",", "").replace(" ", "")) | |
| 122 | + return int(m.group(0)) if m else 0 | |
| 123 | + | |
| 124 | + | |
| 125 | +def _haversine_m(lat1, lng1, lat2, lng2) -> float: | |
| 126 | + import math | |
| 127 | + r = 6371000.0 | |
| 128 | + p1, p2 = math.radians(lat1), math.radians(lat2) | |
| 129 | + dp, dl = math.radians(lat2 - lat1), math.radians(lng2 - lng1) | |
| 130 | + a = (math.sin(dp / 2) ** 2 | |
| 131 | + + math.cos(p1) * math.cos(p2) * math.sin(dl / 2) ** 2) | |
| 132 | + return 2 * r * math.asin(math.sqrt(a)) | |
| 133 | + | |
| 134 | + | |
| 135 | +def verify_store(row, store: dict) -> str: | |
| 136 | + """Vérifie qu'une page magasin Uber Eats correspond bien au resto de la | |
| 137 | + base — CONSERVATEUR : téléphone identique, OU code postal + numéro civique | |
| 138 | + identiques, OU GPS <120 m + nom similaire. Retourne le mode de croisement | |
| 139 | + ('' = pas le même établissement).""" | |
| 140 | + loc = store.get("location") or {} | |
| 141 | + phone = normalize_phone(store.get("phoneNumber")) | |
| 142 | + if phone and normalize_phone(row["phone"]) == phone: | |
| 143 | + return "telephone" | |
| 144 | + postal = (loc.get("postalCode") or "").replace(" ", "").upper() | |
| 145 | + rpostal = (row["postal_code"] or "").replace(" ", "").upper() | |
| 146 | + civic_m = _CIVIC_RE.match(loc.get("streetAddress") or "") | |
| 147 | + rcivic_m = _CIVIC_RE.match(row["address"] or "") | |
| 148 | + if postal and rpostal == postal and civic_m and rcivic_m \ | |
| 149 | + and civic_m.group(1) == rcivic_m.group(1): | |
| 150 | + return "postal+civique" | |
| 151 | + lat, lng = loc.get("latitude"), loc.get("longitude") | |
| 152 | + if (lat is not None and lng is not None | |
| 153 | + and row["lat"] is not None and row["lng"] is not None | |
| 154 | + and _haversine_m(row["lat"], row["lng"], lat, lng) <= MAX_GPS_M): | |
| 155 | + addr_tokens = set(re.sub(r"[^a-z0-9]+", " ", strip_accents( | |
| 156 | + (loc.get("address") or "").lower())).split()) | |
| 157 | + if _name_similar(norm_name(row["name"]), | |
| 158 | + norm_name(store.get("title") or ""), addr_tokens): | |
| 159 | + return "gps+nom" | |
| 160 | + return "" | |
| 161 | + | |
| 162 | + | |
| 163 | +def slugify(name: str) -> str: | |
| 164 | + """Slug façon Uber Eats : « McDonald's » -> « mcdonalds », | |
| 165 | + « Café Dépôt » -> « cafe-depot » (apostrophes supprimées, pas remplacées).""" | |
| 166 | + s = strip_accents((name or "").lower()) | |
| 167 | + s = re.sub(r"['’´`.]", "", s) | |
| 168 | + s = re.sub(r"&", " and ", s) | |
| 169 | + s = re.sub(r"[^a-z0-9]+", "-", s).strip("-") | |
| 170 | + return s | |
| 171 | + | |
| 172 | + | |
| 173 | +class UberEatsConnector(BaseConnector): | |
| 174 | + source_id = "ubereats" | |
| 175 | + request_delay = 1.0 | |
| 176 | + timeout = 60 | |
| 177 | + use_detail_cache = False # verdicts gérés à la main (detail_cache) | |
| 178 | + enrichment_only = True # n'émet aucune fiche (ingest.run) | |
| 179 | + | |
| 180 | + # -- index sitemap --------------------------------------------------------- | |
| 181 | + def _fetch_sitemap(self, url: str) -> str: | |
| 182 | + result = self.scrapfly(url, render_js=False) | |
| 183 | + content = result.get("content") or "" | |
| 184 | + if url.endswith(".gz") or content[:20].startswith("H4sI"): | |
| 185 | + try: | |
| 186 | + return gzip.decompress(base64.b64decode(content)) \ | |
| 187 | + .decode("utf-8", "replace") | |
| 188 | + except (ValueError, OSError): | |
| 189 | + return content | |
| 190 | + return content | |
| 191 | + | |
| 192 | + def load_index(self) -> dict[str, list[str]]: | |
| 193 | + """Index slug -> [URLs /ca/store/…] depuis les sitemaps publics, | |
| 194 | + en cache local 30 jours (data/ubereats-stores.json).""" | |
| 195 | + if INDEX_PATH.exists(): | |
| 196 | + try: | |
| 197 | + data = json.loads(INDEX_PATH.read_text(encoding="utf-8")) | |
| 198 | + if time.time() - float(data.get("fetched_at") or 0) \ | |
| 199 | + < INDEX_REFRESH_DAYS * 86400: | |
| 200 | + return data.get("stores") or {} | |
| 201 | + except (ValueError, OSError): | |
| 202 | + pass | |
| 203 | + robots = self._fetch_sitemap(ROBOTS_URL) | |
| 204 | + sitemaps = re.findall(r"Sitemap:\s*(\S*sitemap-store\S*)", robots) | |
| 205 | + if not sitemaps: | |
| 206 | + raise RuntimeError("robots.txt Uber Eats sans sitemap-store " | |
| 207 | + "(blocage ?)") | |
| 208 | + stores: dict[str, list[str]] = {} | |
| 209 | + for sm_url in sitemaps: | |
| 210 | + xml = self._fetch_sitemap(sm_url) | |
| 211 | + for loc in re.findall(r"<loc>([^<]+)</loc>", xml): | |
| 212 | + mm = _STORE_URL_RE.search(loc) | |
| 213 | + if mm: | |
| 214 | + slug = urllib.parse.unquote(mm.group(1)).lower() | |
| 215 | + stores.setdefault(slug, []) | |
| 216 | + if loc not in stores[slug]: | |
| 217 | + stores[slug].append(loc) | |
| 218 | + INDEX_PATH.parent.mkdir(parents=True, exist_ok=True) | |
| 219 | + INDEX_PATH.write_text(json.dumps( | |
| 220 | + {"fetched_at": time.time(), "stores": stores}, | |
| 221 | + ensure_ascii=False), encoding="utf-8") | |
| 222 | + print(f"[resto-ka] ubereats: index sitemap rafraîchi — " | |
| 223 | + f"{len(stores)} slug(s) canadiens") | |
| 224 | + return stores | |
| 225 | + | |
| 226 | + @staticmethod | |
| 227 | + def candidates_for(name: str, slugs: list[str]) -> list[str]: | |
| 228 | + """Slugs Uber Eats candidats pour un nom de resto (préfixe strict).""" | |
| 229 | + base = slugify(name) | |
| 230 | + if len(base) < 5: | |
| 231 | + return [] | |
| 232 | + import bisect | |
| 233 | + i = bisect.bisect_left(slugs, base) | |
| 234 | + out = [] | |
| 235 | + while i < len(slugs) and len(out) < MAX_CANDIDATES: | |
| 236 | + s = slugs[i] | |
| 237 | + if s == base or s.startswith(base + "-"): | |
| 238 | + out.append(s) | |
| 239 | + i += 1 | |
| 240 | + else: | |
| 241 | + break | |
| 242 | + return out | |
| 243 | + | |
| 244 | + # -- cycle ----------------------------------------------------------------- | |
| 245 | + def _now(self) -> str: | |
| 246 | + return datetime.datetime.now(datetime.timezone.utc) \ | |
| 247 | + .strftime("%Y-%m-%dT%H:%M:%SZ") | |
| 248 | + | |
| 249 | + def _fresh(self, stamp: str, now: float, stale_s: float) -> bool: | |
| 250 | + try: | |
| 251 | + ts = datetime.datetime.strptime(stamp, "%Y-%m-%dT%H:%M:%SZ") \ | |
| 252 | + .replace(tzinfo=datetime.timezone.utc).timestamp() | |
| 253 | + return ts > now - stale_s | |
| 254 | + except (ValueError, TypeError): | |
| 255 | + return False | |
| 256 | + | |
| 257 | + def _details_payload(self, store: dict, url: str, how: str) -> dict: | |
| 258 | + rating = store.get("rating") or {} | |
| 259 | + return { | |
| 260 | + "rating": rating.get("ratingValue"), | |
| 261 | + "review_count": _review_count(store), | |
| 262 | + "price_bucket": store.get("priceBucket"), | |
| 263 | + "url": url.split("?")[0], | |
| 264 | + "uuid": store.get("uuid"), | |
| 265 | + "matched_by": how, | |
| 266 | + "fetched_at": self._now(), | |
| 267 | + } | |
| 268 | + | |
| 269 | + def _probe_store(self, con, row, url: str) -> tuple[bool, bool]: | |
| 270 | + """Visite une page magasin et l'attache au resto si c'est le même | |
| 271 | + établissement. Retourne (matched, menu_added).""" | |
| 272 | + from .. import db | |
| 273 | + mm = _STORE_URL_RE.search(url) | |
| 274 | + store_uuid = mm.group(2) if mm else url | |
| 275 | + cached = db.get_cached_detail(con, self.source_id, store_uuid, | |
| 276 | + "verdict-v1") | |
| 277 | + if cached is not None and cached.get("matched_uid") != row["uid"]: | |
| 278 | + return False, False # déjà identifié comme un autre resto | |
| 279 | + result = self.scrapfly(url, render_js=False) | |
| 280 | + status = result.get("status_code") or 0 | |
| 281 | + if status in (400, 404, 410, 451): # magasin retiré d'Uber Eats | |
| 282 | + db.put_cached_detail(con, self.source_id, store_uuid, | |
| 283 | + "verdict-v1", {"matched_uid": None, | |
| 284 | + "gone": status}) | |
| 285 | + return False, False | |
| 286 | + if status != 200: | |
| 287 | + raise RuntimeError(f"HTTP {status}") | |
| 288 | + state = decode_state(result.get("content") or "") | |
| 289 | + store = store_payload(state or {}) | |
| 290 | + if not store: | |
| 291 | + raise RuntimeError("payload getStoreV1 absent") | |
| 292 | + how = verify_store(row, store) | |
| 293 | + if not how: | |
| 294 | + db.put_cached_detail(con, self.source_id, store_uuid, | |
| 295 | + "verdict-v1", {"matched_uid": None}) | |
| 296 | + return False, False | |
| 297 | + db.put_cached_detail(con, self.source_id, store_uuid, "verdict-v1", | |
| 298 | + {"matched_uid": row["uid"]}) | |
| 299 | + menu = build_menu(store, self._now()) | |
| 300 | + menu_added = False | |
| 301 | + if menu: | |
| 302 | + # validation stricte du schéma menu (prix implausibles, etc.) | |
| 303 | + Restaurant(source=self.source_id, external_id=store_uuid, | |
| 304 | + name=row["name"], menu=menu)._validate_menu() | |
| 305 | + db.upsert_menu(con, row["uid"], menu, time.time()) | |
| 306 | + menu_added = True | |
| 307 | + db.merge_details(con, row["uid"], | |
| 308 | + {"ubereats": self._details_payload(store, url, how)}) | |
| 309 | + con.commit() | |
| 310 | + return True, menu_added | |
| 311 | + | |
| 312 | + def fetch(self) -> list[Restaurant]: | |
| 313 | + if not os.environ.get("SCRAPFLY_KEY"): | |
| 314 | + raise SkipSource("SCRAPFLY_KEY manquant (.env) — Uber Eats est " | |
| 315 | + "derrière Cloudflare, scraping direct impossible") | |
| 316 | + from .. import db | |
| 317 | + index = self.load_index() | |
| 318 | + slugs = sorted(index.keys()) | |
| 319 | + con = db.connect() | |
| 320 | + now = time.time() | |
| 321 | + stale_s = REFRESH_DAYS * 86400.0 | |
| 322 | + budget = MAX_BUDGET | |
| 323 | + matched = menus = misses = failures_row = 0 | |
| 324 | + rows = con.execute( | |
| 325 | + "SELECT uid, name, address, city, postal_code, phone, lat, lng," | |
| 326 | + " details," | |
| 327 | + " EXISTS (SELECT 1 FROM menus m WHERE m.uid=restaurants.uid)" | |
| 328 | + " AS has_menu" | |
| 329 | + " FROM restaurants WHERE active=1 AND dup_of IS NULL AND name<>''" | |
| 330 | + " AND (phone<>'' OR (postal_code<>'' AND address<>'')" | |
| 331 | + " OR (lat IS NOT NULL AND lng IS NOT NULL))" | |
| 332 | + " ORDER BY has_menu ASC, phone<>'' DESC, updated_at DESC" | |
| 333 | + ).fetchall() | |
| 334 | + for row in rows: | |
| 335 | + if budget <= 0: | |
| 336 | + break | |
| 337 | + if failures_row >= MAX_CONSECUTIVE_FAILURES: | |
| 338 | + print("[resto-ka] ubereats: Scrapfly bloqué " | |
| 339 | + f"{failures_row} fois de suite — arrêt du cycle", | |
| 340 | + file=sys.stderr) | |
| 341 | + break | |
| 342 | + try: | |
| 343 | + details = json.loads(row["details"] or "{}") | |
| 344 | + except ValueError: | |
| 345 | + details = {} | |
| 346 | + ue = details.get("ubereats") or {} | |
| 347 | + if self._fresh(ue.get("fetched_at", ""), now, stale_s): | |
| 348 | + continue # déjà frais (<30 j) | |
| 349 | + probe = details.get("ubereats_probe") or {} | |
| 350 | + if self._fresh(probe.get("fetched_at", ""), now, stale_s): | |
| 351 | + continue # échec récent : re-visite dans 30 j | |
| 352 | + if ue.get("url"): # déjà croisé : rafraîchir directement | |
| 353 | + urls = [ue["url"]] | |
| 354 | + else: | |
| 355 | + cand = self.candidates_for(row["name"], slugs) | |
| 356 | + urls = [u for s in cand for u in index.get(s, [])] | |
| 357 | + if not urls: | |
| 358 | + continue # aucun candidat : pas de marqueur, | |
| 359 | + # l'index du mois prochain peut changer | |
| 360 | + found = False | |
| 361 | + for url in urls[:MAX_CANDIDATES]: | |
| 362 | + if budget <= 0: | |
| 363 | + break | |
| 364 | + budget -= 1 | |
| 365 | + try: | |
| 366 | + ok, menu_added = self._probe_store(con, row, url) | |
| 367 | + except Exception as exc: | |
| 368 | + failures_row += 1 | |
| 369 | + print(f"[resto-ka] ubereats: {row['uid']} erreur: {exc}", | |
| 370 | + file=sys.stderr) | |
| 371 | + continue | |
| 372 | + failures_row = 0 | |
| 373 | + if ok: | |
| 374 | + matched += 1 | |
| 375 | + menus += 1 if menu_added else 0 | |
| 376 | + found = True | |
| 377 | + break | |
| 378 | + if not found and not ue.get("url"): | |
| 379 | + db.merge_details(con, row["uid"], | |
| 380 | + {"ubereats_probe": | |
| 381 | + {"miss": "aucun magasin correspondant", | |
| 382 | + "fetched_at": self._now()}}) | |
| 383 | + con.commit() | |
| 384 | + misses += 1 | |
| 385 | + con.commit() | |
| 386 | + con.close() | |
| 387 | + self.enriched_count = matched | |
| 388 | + self.enrich_message = (f"{matched} resto(s) croisés Uber Eats " | |
| 389 | + f"({menus} menu(s) livraison), {misses} sans " | |
| 390 | + f"correspondance, budget restant " | |
| 391 | + f"{max(budget, 0)} page(s)") | |
| 392 | + print(f"[resto-ka] ubereats: {self.enrich_message}") | |
| 393 | + return [] | |
added
restoka/connectors/yelp_scrape.py
+238 −0
@@ -0,0 +1,238 @@ | ||
| 1 | +# ============================================================================== | |
| 2 | +# Author: Simon-Pierre Boucher <contact@spboucher.ai> | |
| 3 | +# File: restoka/connectors/yelp_scrape.py | |
| 4 | +# Desc: Connecteur d'ENRICHISSEMENT Yelp SANS clé API (avis/notes par | |
| 5 | +# scraping léger). Une seule page de RECHERCHE yelp.ca par resto | |
| 6 | +# (find_desc=<nom>&find_loc=<ville>) via Scrapfly ASP : la page embarque | |
| 7 | +# un cache Apollo (scripts data-apollo-state) avec, PAR commerce, | |
| 8 | +# alias, nom, note, nombre d'avis, fourchette de prix, catégories et | |
| 9 | +# adresse (ville + civique). Croisement CONSERVATEUR : nom similaire | |
| 10 | +# + même numéro civique + ville compatible (jamais fusionner deux | |
| 11 | +# établissements). Résultat -> details.yelp {rating, review_count, | |
| 12 | +# price, categories, url} (mêmes clés que Yelp Fusion — le connecteur | |
| 13 | +# officiel `yelp` reprend la main dès qu'une YELP_API_KEY existe). | |
| 14 | +# | |
| 15 | +# Économe : budget YELP_SCRAPE_BUDGET (défaut 80) pages/cycle, | |
| 16 | +# re-visite 30 jours (hit comme miss), priorité aux restos avec menu. | |
| 17 | +# Base légale : pages publiques, extraction minimale (note agrégée + | |
| 18 | +# lien vers la fiche Yelp, attribution affichée) — voir CLAUDE.md §15. | |
| 19 | +# ============================================================================== | |
| 20 | +from __future__ import annotations | |
| 21 | + | |
| 22 | +import datetime | |
| 23 | +import html as html_mod | |
| 24 | +import json | |
| 25 | +import os | |
| 26 | +import re | |
| 27 | +import sys | |
| 28 | +import urllib.parse | |
| 29 | + | |
| 30 | +from ..inspections import _CIVIC_RE, _name_similar, norm_name | |
| 31 | +from ..regions import strip_accents | |
| 32 | +from ..schema import Restaurant | |
| 33 | +from .base import BaseConnector, SkipSource | |
| 34 | + | |
| 35 | +SEARCH_URL = "https://www.yelp.ca/search?find_desc={desc}&find_loc={loc}" | |
| 36 | +MAX_BUDGET = int(os.environ.get("YELP_SCRAPE_BUDGET", "80")) # pages/cycle | |
| 37 | +REFRESH_DAYS = 30 # re-visite des notes ET des échecs | |
| 38 | +MAX_CONSECUTIVE_FAILURES = 3 # Yelp bloque même Scrapfly -> on arrête | |
| 39 | + | |
| 40 | +_APOLLO_RE = re.compile( | |
| 41 | + r"<script[^>]*data-apollo-state[^>]*>\s*<!--(.*?)-->\s*</script>", re.S) | |
| 42 | + | |
| 43 | + | |
| 44 | +def parse_apollo_businesses(page_html: str) -> list[dict]: | |
| 45 | + """Extrait les commerces du cache Apollo embarqué dans une page de | |
| 46 | + recherche yelp.ca : [{alias, name, rating, review_count, price, | |
| 47 | + categories, city, address, closed}].""" | |
| 48 | + cache: dict = {} | |
| 49 | + for m in _APOLLO_RE.finditer(page_html): | |
| 50 | + try: | |
| 51 | + cache.update(json.loads(html_mod.unescape(m.group(1)))) | |
| 52 | + except ValueError: | |
| 53 | + continue | |
| 54 | + if not cache: | |
| 55 | + return [] | |
| 56 | + categories = {k.split(":", 1)[1]: (v.get("title") or "") | |
| 57 | + for k, v in cache.items() | |
| 58 | + if k.startswith("BusinessCategory:") and isinstance(v, dict)} | |
| 59 | + locations = {k.split(":", 1)[1]: (v.get("address") or {}) | |
| 60 | + for k, v in cache.items() | |
| 61 | + if k.startswith("BusinessLocation:") and isinstance(v, dict)} | |
| 62 | + out = [] | |
| 63 | + for k, v in cache.items(): | |
| 64 | + if not k.startswith("Business:") or not isinstance(v, dict): | |
| 65 | + continue | |
| 66 | + if not v.get("alias") or not v.get("name"): | |
| 67 | + continue | |
| 68 | + loc_ref = ((v.get("location") or {}).get("__ref") or "") | |
| 69 | + addr = locations.get(loc_ref.split(":", 1)[-1], {}) | |
| 70 | + cats = [] | |
| 71 | + for c in v.get("categories") or []: | |
| 72 | + ref = (c.get("__ref") or "").split(":", 1)[-1] | |
| 73 | + if categories.get(ref): | |
| 74 | + cats.append(categories[ref]) | |
| 75 | + closed = any((a or {}).get("type") == "permclosed" | |
| 76 | + for a in [v.get(kk) for kk in v | |
| 77 | + if kk.startswith("activeAlert")]) | |
| 78 | + price = v.get("priceRange") | |
| 79 | + out.append({ | |
| 80 | + "alias": v["alias"], | |
| 81 | + "name": v["name"], | |
| 82 | + "rating": v.get("rating"), | |
| 83 | + "review_count": v.get("reviewCount") or 0, | |
| 84 | + "price": (price or {}).get("display") if isinstance(price, dict) | |
| 85 | + else price, | |
| 86 | + "categories": cats, | |
| 87 | + "city": (addr or {}).get("city") or "", | |
| 88 | + "address": (addr or {}).get("addressLine1") or "", | |
| 89 | + "closed": closed, | |
| 90 | + }) | |
| 91 | + return out | |
| 92 | + | |
| 93 | + | |
| 94 | +def _city_key(city: str) -> str: | |
| 95 | + """« Quebec City » / « Québec » -> « quebec » ; « Montréal » -> « montreal ».""" | |
| 96 | + s = strip_accents((city or "").lower()) | |
| 97 | + s = re.sub(r"[^a-z0-9]+", " ", s).strip() | |
| 98 | + return re.sub(r"\bcity\b", "", s).strip() | |
| 99 | + | |
| 100 | + | |
| 101 | +def match_business(row, candidates: list[dict]) -> dict | None: | |
| 102 | + """Croisement CONSERVATEUR d'un resto avec les résultats Yelp : | |
| 103 | + nom similaire ET même numéro civique ET ville compatible. Sans civique | |
| 104 | + des deux côtés : nom normalisé IDENTIQUE + même ville. Ambigu -> None.""" | |
| 105 | + rname = norm_name(row["name"]) | |
| 106 | + rcity = _city_key(row["city"]) | |
| 107 | + rcivic_m = _CIVIC_RE.match(row["address"] or "") | |
| 108 | + rcivic = rcivic_m.group(1) if rcivic_m else "" | |
| 109 | + addr_tokens = set(re.sub(r"[^a-z0-9]+", " ", strip_accents( | |
| 110 | + (row["address"] or "").lower())).split()) | set(rcity.split()) | |
| 111 | + hits = [] | |
| 112 | + for c in candidates: | |
| 113 | + if c["closed"]: | |
| 114 | + continue | |
| 115 | + ccity = _city_key(c["city"]) | |
| 116 | + if rcity and ccity and rcity != ccity \ | |
| 117 | + and rcity not in ccity and ccity not in rcity: | |
| 118 | + continue | |
| 119 | + cname = norm_name(c["name"]) | |
| 120 | + ccivic_m = _CIVIC_RE.match(c["address"] or "") | |
| 121 | + ccivic = ccivic_m.group(1) if ccivic_m else "" | |
| 122 | + if rcivic and ccivic: | |
| 123 | + if rcivic == ccivic and _name_similar(rname, cname, addr_tokens): | |
| 124 | + hits.append(c) | |
| 125 | + elif rname and rname == cname and rcity and ccity: | |
| 126 | + hits.append(c) # nom exact + ville, sans civique | |
| 127 | + aliases = {h["alias"] for h in hits} | |
| 128 | + if len(aliases) == 1: | |
| 129 | + return hits[0] | |
| 130 | + return None # rien ou ambigu : on ne fusionne pas | |
| 131 | + | |
| 132 | + | |
| 133 | +class YelpScrapeConnector(BaseConnector): | |
| 134 | + source_id = "yelp-scrape" | |
| 135 | + request_delay = 1.0 | |
| 136 | + timeout = 60 | |
| 137 | + use_detail_cache = False | |
| 138 | + enrichment_only = True # n'émet aucune fiche (ingest.run) | |
| 139 | + | |
| 140 | + def _now(self) -> str: | |
| 141 | + return datetime.datetime.now(datetime.timezone.utc) \ | |
| 142 | + .strftime("%Y-%m-%dT%H:%M:%SZ") | |
| 143 | + | |
| 144 | + def _fresh(self, stamp: str, now: float, stale_s: float) -> bool: | |
| 145 | + try: | |
| 146 | + ts = datetime.datetime.strptime(stamp, "%Y-%m-%dT%H:%M:%SZ") \ | |
| 147 | + .replace(tzinfo=datetime.timezone.utc).timestamp() | |
| 148 | + return ts > now - stale_s | |
| 149 | + except (ValueError, TypeError): | |
| 150 | + return False | |
| 151 | + | |
| 152 | + def fetch(self) -> list[Restaurant]: | |
| 153 | + if not os.environ.get("SCRAPFLY_KEY"): | |
| 154 | + raise SkipSource("SCRAPFLY_KEY manquant (.env) — scraping Yelp " | |
| 155 | + "impossible sans contournement anti-bot") | |
| 156 | + import time as _time | |
| 157 | + from .. import db | |
| 158 | + con = db.connect() | |
| 159 | + now = _time.time() | |
| 160 | + stale_s = REFRESH_DAYS * 86400.0 | |
| 161 | + budget = MAX_BUDGET | |
| 162 | + enriched = misses = failures_row = 0 | |
| 163 | + rows = con.execute( | |
| 164 | + "SELECT uid, name, address, city, details," | |
| 165 | + " EXISTS (SELECT 1 FROM menus m WHERE m.uid=restaurants.uid)" | |
| 166 | + " AS has_menu" | |
| 167 | + " FROM restaurants WHERE active=1 AND dup_of IS NULL" | |
| 168 | + " AND name<>'' AND city<>''" | |
| 169 | + # indépendants d'abord : les succursales de chaînes n'ont presque | |
| 170 | + # jamais d'avis Yelp (budget mieux investi ailleurs) | |
| 171 | + " ORDER BY has_menu DESC, chain IS NULL DESC, phone<>'' DESC," | |
| 172 | + " updated_at DESC" | |
| 173 | + ).fetchall() | |
| 174 | + for row in rows: | |
| 175 | + if budget <= 0: | |
| 176 | + break | |
| 177 | + if failures_row >= MAX_CONSECUTIVE_FAILURES: | |
| 178 | + print("[resto-ka] yelp-scrape: Scrapfly bloqué " | |
| 179 | + f"{failures_row} fois de suite — arrêt du cycle", | |
| 180 | + file=sys.stderr) | |
| 181 | + break | |
| 182 | + try: | |
| 183 | + details = json.loads(row["details"] or "{}") | |
| 184 | + except ValueError: | |
| 185 | + details = {} | |
| 186 | + yelp = details.get("yelp") or {} | |
| 187 | + if self._fresh(yelp.get("fetched_at", ""), now, stale_s): | |
| 188 | + continue # note déjà fraîche (<30 j) | |
| 189 | + probe = details.get("yelp_scrape") or {} | |
| 190 | + if self._fresh(probe.get("fetched_at", ""), now, stale_s): | |
| 191 | + continue # échec récent : re-visite dans 30 j | |
| 192 | + url = SEARCH_URL.format( | |
| 193 | + desc=urllib.parse.quote(row["name"][:64]), | |
| 194 | + loc=urllib.parse.quote(f"{row['city']}, QC")) | |
| 195 | + budget -= 1 | |
| 196 | + try: | |
| 197 | + result = self.scrapfly(url, render_js=False) | |
| 198 | + except Exception as exc: | |
| 199 | + failures_row += 1 | |
| 200 | + print(f"[resto-ka] yelp-scrape: {row['uid']} erreur: {exc}", | |
| 201 | + file=sys.stderr) | |
| 202 | + continue | |
| 203 | + if (result.get("status_code") or 0) != 200: | |
| 204 | + failures_row += 1 | |
| 205 | + continue | |
| 206 | + failures_row = 0 | |
| 207 | + candidates = parse_apollo_businesses(result.get("content") or "") | |
| 208 | + hit = match_business(row, candidates) | |
| 209 | + if hit and hit.get("rating") is not None: | |
| 210 | + db.merge_details(con, row["uid"], {"yelp": { | |
| 211 | + "url": f"https://www.yelp.ca/biz/" | |
| 212 | + f"{urllib.parse.quote(hit['alias'])}", | |
| 213 | + "name": hit["name"], | |
| 214 | + "rating": hit["rating"], | |
| 215 | + "review_count": hit["review_count"], | |
| 216 | + "price": hit.get("price"), | |
| 217 | + "categories": hit.get("categories") or [], | |
| 218 | + "matched_by": "nom+adresse", | |
| 219 | + "via": "scrape", | |
| 220 | + "fetched_at": self._now(), | |
| 221 | + }}) | |
| 222 | + enriched += 1 | |
| 223 | + else: | |
| 224 | + reason = ("fiche sans note" if hit else | |
| 225 | + "introuvable ou ambigu") | |
| 226 | + db.merge_details(con, row["uid"], | |
| 227 | + {"yelp_scrape": {"miss": reason, | |
| 228 | + "fetched_at": self._now()}}) | |
| 229 | + misses += 1 | |
| 230 | + con.commit() | |
| 231 | + con.commit() | |
| 232 | + con.close() | |
| 233 | + self.enriched_count = enriched | |
| 234 | + self.enrich_message = (f"{enriched} resto(s) notés (scraping), " | |
| 235 | + f"{misses} sans correspondance, " | |
| 236 | + f"budget restant {max(budget, 0)} page(s)") | |
| 237 | + print(f"[resto-ka] yelp-scrape: {self.enrich_message}") | |
| 238 | + return [] | |
modified
restoka/ingest.py
+5 −0
@@ -104,6 +104,11 @@ def enrich() -> None: | ||
| 104 | 104 | inspections.sync() |
| 105 | 105 | except Exception as exc: |
| 106 | 106 | print(f"[resto-ka] mapaq: erreur non bloquante: {exc}", file=sys.stderr) |
| 107 | + try: # permis d'alcool RACJ (Données Québec) — cadence hebdo, croisement | |
| 108 | + from . import permits | |
| 109 | + permits.sync() | |
| 110 | + except Exception as exc: | |
| 111 | + print(f"[resto-ka] racj: erreur non bloquante: {exc}", file=sys.stderr) | |
| 107 | 112 | |
| 108 | 113 | |
| 109 | 114 | def watch(interval_seconds: int = 7 * 86400 // 7) -> None: |
added
restoka/permits.py
+218 −0
@@ -0,0 +1,218 @@ | ||
| 1 | +# ============================================================================== | |
| 2 | +# Author: Simon-Pierre Boucher <contact@spboucher.ai> | |
| 3 | +# File: restoka/permits.py | |
| 4 | +# Desc: Enrichissement RACJ — permis d'alcool en vigueur (Régie des alcools, | |
| 5 | +# des courses et des jeux, Données Québec, licence CC-BY 4.0). | |
| 6 | +# Télécharge le registre CSV des permis de détaillant d'alcool | |
| 7 | +# (bar, restaurant pour vendre/servir…), regroupe par établissement et | |
| 8 | +# croise de façon CONSERVATRICE avec les restaurants (mêmes règles que | |
| 9 | +# les inspections MAPAQ : nom exact + lieu, ou postal + civique + nom | |
| 10 | +# similaire — jamais fusionner deux établissements). Résultat dans | |
| 11 | +# details.permis_alcool (catégories de permis, capacité, titulaire). | |
| 12 | +# Gratuit, sans clé API. Cadence hebdomadaire (guard 6 jours). | |
| 13 | +# ============================================================================== | |
| 14 | +from __future__ import annotations | |
| 15 | + | |
| 16 | +import csv | |
| 17 | +import io | |
| 18 | +import json | |
| 19 | +import re | |
| 20 | +import sys | |
| 21 | +import time | |
| 22 | +from datetime import date | |
| 23 | + | |
| 24 | +import requests | |
| 25 | + | |
| 26 | +from .connectors.base import USER_AGENT | |
| 27 | +from .inspections import (_CIVIC_RE, _NUMBERED_CO_RE, _name_similar, | |
| 28 | + core_name, norm_name) | |
| 29 | +from .regions import strip_accents | |
| 30 | + | |
| 31 | +SOURCE_ID = "racj" | |
| 32 | +REFRESH_DAYS = 6 # au plus une fois par cycle hebdo | |
| 33 | +CSV_URL = ("https://www.donneesquebec.ca/recherche/dataset/" | |
| 34 | + "d817c9f7-76c7-44af-882d-0d673056ef86/resource/" | |
| 35 | + "6b69360c-af8d-4c57-b5bf-1c38d5461de3/download/" | |
| 36 | + "racj-alcool-detaillant.csv") | |
| 37 | + | |
| 38 | + | |
| 39 | +def _download() -> list[dict]: | |
| 40 | + resp = requests.get(CSV_URL, headers={"User-Agent": USER_AGENT}, timeout=120) | |
| 41 | + resp.raise_for_status() | |
| 42 | + text = None | |
| 43 | + for enc in ("utf-8-sig", "latin-1"): | |
| 44 | + try: | |
| 45 | + text = resp.content.decode(enc) | |
| 46 | + break | |
| 47 | + except UnicodeDecodeError: | |
| 48 | + continue | |
| 49 | + if text is None: | |
| 50 | + raise RuntimeError("encodage CSV RACJ inconnu") | |
| 51 | + return list(csv.DictReader(io.StringIO(text))) | |
| 52 | + | |
| 53 | + | |
| 54 | +def group_establishments(records: list[dict]) -> list[dict]: | |
| 55 | + """Regroupe les lignes CSV (une par local/terrasse/permis) par | |
| 56 | + établissement (NoEtablissement) : catégories de permis, capacité totale, | |
| 57 | + nom, adresse, ville, code postal, titulaire.""" | |
| 58 | + by_no: dict[str, dict] = {} | |
| 59 | + for rec in records: | |
| 60 | + no = (rec.get("NoEtablissement") or "").strip() | |
| 61 | + if not no: | |
| 62 | + continue | |
| 63 | + e = by_no.setdefault(no, { | |
| 64 | + "no": no, | |
| 65 | + "nom": (rec.get("RaisonSociale") or "").strip(), | |
| 66 | + "titulaire": (rec.get("Titulaire") or "").strip(), | |
| 67 | + "adresse": (rec.get("Adresse") or "").strip(), | |
| 68 | + "ville": (rec.get("Ville") or "").strip(), | |
| 69 | + "postal": (rec.get("CodePostal") or "").replace(" ", "").upper(), | |
| 70 | + "categories": set(), | |
| 71 | + "capacite": 0, | |
| 72 | + "permis": set(), | |
| 73 | + }) | |
| 74 | + cat = (rec.get("Categorie") or "").strip() | |
| 75 | + if cat: | |
| 76 | + e["categories"].add(cat) | |
| 77 | + e["permis"].add((rec.get("NoPermis") or "").strip()) | |
| 78 | + try: | |
| 79 | + e["capacite"] += int(rec.get("Capacite") or 0) | |
| 80 | + except ValueError: | |
| 81 | + pass | |
| 82 | + return list(by_no.values()) | |
| 83 | + | |
| 84 | + | |
| 85 | +def match(con, establishments: list[dict]) -> dict: | |
| 86 | + """Croisement CONSERVATEUR permis <-> restaurants (mêmes règles que | |
| 87 | + inspections.match) : | |
| 88 | + | |
| 89 | + Règle A « nom+lieu » : nom commercial normalisé IDENTIQUE ET (code postal | |
| 90 | + identique OU même ville + même numéro civique). | |
| 91 | + Règle B « adresse+nom » : code postal identique ET numéro civique | |
| 92 | + identique ET similarité de nom (hors tokens d'adresse/ville). | |
| 93 | + """ | |
| 94 | + restos = con.execute( | |
| 95 | + "SELECT uid, name, chain, city, postal_code, address FROM restaurants" | |
| 96 | + " WHERE active=1 AND dup_of IS NULL").fetchall() | |
| 97 | + by_name: dict[str, list] = {} | |
| 98 | + by_postal: dict[str, list] = {} | |
| 99 | + for r in restos: | |
| 100 | + info = { | |
| 101 | + "uid": r["uid"], | |
| 102 | + "nname": norm_name(r["name"]), | |
| 103 | + "nchain": norm_name(r["chain"] or ""), | |
| 104 | + "ncity": strip_accents((r["city"] or "").lower()).strip(), | |
| 105 | + "postal": (r["postal_code"] or "").replace(" ", "").upper(), | |
| 106 | + } | |
| 107 | + mm = _CIVIC_RE.match(r["address"] or "") | |
| 108 | + info["civic"] = mm.group(1) if mm else "" | |
| 109 | + if info["nname"]: | |
| 110 | + by_name.setdefault(info["nname"], []).append(info) | |
| 111 | + core = core_name(info["nname"]) | |
| 112 | + if core and core != info["nname"]: | |
| 113 | + by_name.setdefault(core, []).append(info) | |
| 114 | + if info["postal"]: | |
| 115 | + by_postal.setdefault(info["postal"], []).append(info) | |
| 116 | + | |
| 117 | + today = date.today().isoformat() | |
| 118 | + matched = 0 | |
| 119 | + seen_uids: set[str] = set() | |
| 120 | + for est in establishments: | |
| 121 | + names = [] | |
| 122 | + if est["nom"]: | |
| 123 | + names.append(norm_name(est["nom"])) | |
| 124 | + tit = est["titulaire"] | |
| 125 | + if tit and not _NUMBERED_CO_RE.match(strip_accents(tit.lower())): | |
| 126 | + names.append(norm_name(tit)) | |
| 127 | + names = [n for n in names if n] | |
| 128 | + names += [c for c in (core_name(n) for n in names) | |
| 129 | + if c and c not in names] | |
| 130 | + if not names: | |
| 131 | + continue | |
| 132 | + ncity = strip_accents(est["ville"].lower()).strip() | |
| 133 | + civic_m = _CIVIC_RE.match(est["adresse"]) | |
| 134 | + civic = civic_m.group(1) if civic_m else "" | |
| 135 | + postal = est["postal"] | |
| 136 | + | |
| 137 | + hit, how = None, "" | |
| 138 | + # Règle A : nom exact + code postal ou ville+civique | |
| 139 | + for n in names: | |
| 140 | + for info in by_name.get(n, []): | |
| 141 | + same_place = ((postal and info["postal"] == postal) | |
| 142 | + or (ncity and info["ncity"] | |
| 143 | + and (ncity == info["ncity"] | |
| 144 | + or ncity in info["ncity"] | |
| 145 | + or info["ncity"] in ncity) | |
| 146 | + and civic and info["civic"] == civic)) | |
| 147 | + if same_place: | |
| 148 | + hit, how = info["uid"], "nom+lieu" | |
| 149 | + break | |
| 150 | + if hit: | |
| 151 | + break | |
| 152 | + # Règle B : code postal + civique + similarité de nom | |
| 153 | + if hit is None and postal and civic: | |
| 154 | + addr_tokens = set(re.sub(r"[^a-z0-9]+", " ", strip_accents( | |
| 155 | + est["adresse"].lower())).split()) | {ncity} | |
| 156 | + for info in by_postal.get(postal, []): | |
| 157 | + if info["civic"] != civic: | |
| 158 | + continue | |
| 159 | + if any(_name_similar(n, info["nname"], addr_tokens) | |
| 160 | + or (info["nchain"] | |
| 161 | + and _name_similar(n, info["nchain"], addr_tokens)) | |
| 162 | + for n in names): | |
| 163 | + hit, how = info["uid"], "adresse+nom" | |
| 164 | + break | |
| 165 | + if hit is None or hit in seen_uids: | |
| 166 | + continue # jamais deux établissements sur 1 fiche | |
| 167 | + seen_uids.add(hit) | |
| 168 | + from . import db as _db | |
| 169 | + _db.merge_details(con, hit, {"permis_alcool": { | |
| 170 | + "categories": sorted(est["categories"]), | |
| 171 | + "nb_permis": len(est["permis"]), | |
| 172 | + "capacite": est["capacite"] or None, | |
| 173 | + "titulaire": est["titulaire"], | |
| 174 | + "matched_by": how, | |
| 175 | + "source": "RACJ / Données Québec (CC-BY 4.0)", | |
| 176 | + "maj": today, | |
| 177 | + }}) | |
| 178 | + matched += 1 | |
| 179 | + con.commit() | |
| 180 | + total = con.execute( | |
| 181 | + "SELECT COUNT(*) c FROM restaurants WHERE active=1 AND dup_of IS NULL" | |
| 182 | + " AND details LIKE '%\"permis_alcool\"%'").fetchone()["c"] | |
| 183 | + return {"matched": matched, "with_permit": total} | |
| 184 | + | |
| 185 | + | |
| 186 | +def sync(con=None, force: bool = False) -> dict | None: | |
| 187 | + """Télécharge le registre RACJ, regroupe et croise. Cadence hebdo.""" | |
| 188 | + from . import db | |
| 189 | + own = con is None | |
| 190 | + if own: | |
| 191 | + con = db.connect() | |
| 192 | + try: | |
| 193 | + last = con.execute( | |
| 194 | + "SELECT MAX(ts) ts FROM sync_log WHERE source=? AND ok=1", | |
| 195 | + (SOURCE_ID,)).fetchone()["ts"] | |
| 196 | + if not force and last and time.time() - last < REFRESH_DAYS * 86400: | |
| 197 | + return None # déjà à jour cette semaine | |
| 198 | + records = _download() | |
| 199 | + establishments = group_establishments(records) | |
| 200 | + m = match(con, establishments) | |
| 201 | + stats = {"rows": len(records), "etablissements": len(establishments), | |
| 202 | + **m} | |
| 203 | + con.execute( | |
| 204 | + "INSERT INTO sync_log (source, ts, found, added, updated, removed," | |
| 205 | + " ok, message, stats) VALUES (?,?,?,?,0,0,1,?,?)", | |
| 206 | + (SOURCE_ID, time.time(), len(establishments), m["matched"], | |
| 207 | + f"permis d'alcool croisés : {m['with_permit']} resto(s)", | |
| 208 | + json.dumps(stats, ensure_ascii=False))) | |
| 209 | + con.commit() | |
| 210 | + print(f"[resto-ka] racj: {stats}") | |
| 211 | + return stats | |
| 212 | + except Exception as exc: | |
| 213 | + db.log_failure(con, SOURCE_ID, str(exc)) | |
| 214 | + print(f"[resto-ka] racj: erreur non bloquante: {exc}", file=sys.stderr) | |
| 215 | + return {"error": str(exc)} | |
| 216 | + finally: | |
| 217 | + if own: | |
| 218 | + con.close() | |
modified
restoka/web.py
+6 −1
@@ -452,13 +452,18 @@ def stats(): | ||
| 452 | 452 | SUM(uid IS NOT NULL) inspections_matched, |
| 453 | 453 | COUNT(DISTINCT uid) restaurants_with_inspections |
| 454 | 454 | FROM inspections""").fetchone()) |
| 455 | + enr = dict(con.execute( | |
| 456 | + """SELECT SUM(details LIKE '%"yelp":%') with_yelp_rating, | |
| 457 | + SUM(details LIKE '%"ubereats":%') with_ubereats, | |
| 458 | + SUM(details LIKE '%"permis_alcool"%') with_alcohol_permit | |
| 459 | + FROM restaurants WHERE active=1 AND dup_of IS NULL""").fetchone()) | |
| 455 | 460 | con.close() |
| 456 | 461 | try: # nb de sources au registre (data/sources.json), actives ou en attente |
| 457 | 462 | registry = json.loads(SOURCES_PATH.read_text(encoding="utf-8"))["sources"] |
| 458 | 463 | registered = len(registry) |
| 459 | 464 | except (ValueError, OSError, KeyError): |
| 460 | 465 | registered = None |
| 461 | − return {**head, **m, **insp, "sources_registry": registered, | |
| 466 | + return {**head, **m, **insp, **enr, "sources_registry": registered, | |
| 462 | 467 | "by_region": by_region, "by_context": by_context, |
| 463 | 468 | "recent_syncs": log} |
| 464 | 469 | |
modified
scripts/gen_connector_docs.py
+2 −1
@@ -26,7 +26,8 @@ from pathlib import Path | ||
| 26 | 26 | |
| 27 | 27 | ROOT = Path(__file__).resolve().parents[1] |
| 28 | 28 | CONN_DIR = ROOT / "restoka" / "connectors" |
| 29 | −EXTRA_MODULES = [ROOT / "restoka" / "inspections.py"] # mapaq (enrichissement) | |
| 29 | +EXTRA_MODULES = [ROOT / "restoka" / "inspections.py", # mapaq (enrichissement) | |
| 30 | + ROOT / "restoka" / "permits.py"] # racj (enrichissement) | |
| 30 | 31 | DOCS_DIR = ROOT / "docs" / "connecteurs" |
| 31 | 32 | DB_PATH = ROOT / "data" / "restoka.db" |
| 32 | 33 | SOURCES_JSON = ROOT / "data" / "sources.json" |
added
tests/test_permits.py
+102 −0
@@ -0,0 +1,102 @@ | ||
| 1 | +# ============================================================================== | |
| 2 | +# Author: Simon-Pierre Boucher <contact@spboucher.ai> | |
| 3 | +# File: tests/test_permits.py | |
| 4 | +# Desc: Enrichissement RACJ (permis d'alcool) — regroupement des lignes CSV | |
| 5 | +# par établissement, croisement CONSERVATEUR avec les restaurants | |
| 6 | +# (nom+lieu, adresse+nom, jamais de fusion hasardeuse) et écriture | |
| 7 | +# de details.permis_alcool. Tout HORS LIGNE (aucun téléchargement). | |
| 8 | +# ============================================================================== | |
| 9 | +import json | |
| 10 | + | |
| 11 | +from restoka import permits | |
| 12 | +from restoka import db as rdb | |
| 13 | +from restoka.schema import Restaurant | |
| 14 | + | |
| 15 | +RECORDS = [ | |
| 16 | + { # établissement 1 : deux lignes (salle + terrasse), deux permis | |
| 17 | + "NoEtablissement": "22723", "RaisonSociale": "Chez Mamy", | |
| 18 | + "Titulaire": "9312-5581 Québec Inc.", "Neq": "1170486295", | |
| 19 | + "Adresse": "1999 Rue Sainte-Famille ", "CodeVille": "94068", | |
| 20 | + "Ville": "Saguenay", "CodePostal": "G7X4X5", "RegAdmin": "02", | |
| 21 | + "NoPermis": "100000067-1", "Categorie": "Restaurant pour vendre", | |
| 22 | + "TypeLocal": "Autre", "Capacite": "80", | |
| 23 | + }, | |
| 24 | + { | |
| 25 | + "NoEtablissement": "22723", "RaisonSociale": "Chez Mamy", | |
| 26 | + "Titulaire": "9312-5581 Québec Inc.", "Neq": "1170486295", | |
| 27 | + "Adresse": "1999 Rue Sainte-Famille ", "CodeVille": "94068", | |
| 28 | + "Ville": "Saguenay", "CodePostal": "G7X4X5", "RegAdmin": "02", | |
| 29 | + "NoPermis": "100000068-1", "Categorie": "Bar", | |
| 30 | + "TypeLocal": "Terrasse", "Capacite": "40", | |
| 31 | + }, | |
| 32 | + { # établissement 2 : croisement par postal + civique + nom similaire | |
| 33 | + "NoEtablissement": "31000", "RaisonSociale": "Resto-Bar Pizza Bella", | |
| 34 | + "Titulaire": "Gestion Pizza Bella Inc.", "Neq": "", | |
| 35 | + "Adresse": "300 Rue Principale", "CodeVille": "23027", | |
| 36 | + "Ville": "Québec", "CodePostal": "G1K3Y2", "RegAdmin": "03", | |
| 37 | + "NoPermis": "200000001-1", "Categorie": "Restaurant pour servir", | |
| 38 | + "TypeLocal": "Autre", "Capacite": "120", | |
| 39 | + }, | |
| 40 | + { # piège : homonyme dans une AUTRE ville — ne doit RIEN croiser | |
| 41 | + "NoEtablissement": "40000", "RaisonSociale": "Chez Mamy", | |
| 42 | + "Titulaire": "8888-0000 Québec Inc.", "Neq": "", | |
| 43 | + "Adresse": "714 Rang Ouest", "CodeVille": "11111", | |
| 44 | + "Ville": "Saint-Jean-de-Dieu", "CodePostal": "G0L3M0", | |
| 45 | + "RegAdmin": "01", "NoPermis": "300000001-1", "Categorie": "Bar", | |
| 46 | + "TypeLocal": "Autre", "Capacite": "60", | |
| 47 | + }, | |
| 48 | +] | |
| 49 | + | |
| 50 | + | |
| 51 | +def _seed(con): | |
| 52 | + restos = [ | |
| 53 | + Restaurant(source="osm", external_id="node/1", name="Chez Mamy", | |
| 54 | + address="1999 Rue Sainte-Famille", city="Saguenay", | |
| 55 | + postal_code="G7X 4X5", lat=48.42, lng=-71.06), | |
| 56 | + Restaurant(source="osm", external_id="node/2", name="Bella Pizza Resto", | |
| 57 | + address="300 Rue Principale", city="Québec", | |
| 58 | + postal_code="G1K 3Y2", lat=46.81, lng=-71.21), | |
| 59 | + # piège : même nom mais autre ville (le permis 40000 ne doit pas venir ici) | |
| 60 | + Restaurant(source="osm", external_id="node/3", name="Chez Mamy", | |
| 61 | + address="12 Rue Untel", city="Gatineau", | |
| 62 | + postal_code="J8X 1A1", lat=45.47, lng=-75.70), | |
| 63 | + ] | |
| 64 | + rdb.sync_source(con, "osm", [r.finalize() for r in restos]) | |
| 65 | + | |
| 66 | + | |
| 67 | +def test_group_establishments(): | |
| 68 | + ests = permits.group_establishments(RECORDS) | |
| 69 | + assert len(ests) == 3 | |
| 70 | + e = next(x for x in ests if x["no"] == "22723") | |
| 71 | + assert e["categories"] == {"Restaurant pour vendre", "Bar"} | |
| 72 | + assert e["capacite"] == 120 # 80 + 40 (salle + terrasse) | |
| 73 | + assert len(e["permis"]) == 2 | |
| 74 | + assert e["postal"] == "G7X4X5" | |
| 75 | + | |
| 76 | + | |
| 77 | +def test_match_conservateur(con): | |
| 78 | + _seed(con) | |
| 79 | + stats = permits.match(con, permits.group_establishments(RECORDS)) | |
| 80 | + assert stats["matched"] == 2 and stats["with_permit"] == 2 | |
| 81 | + | |
| 82 | + # nom+lieu -> le Chez Mamy de Saguenay, jamais celui de Gatineau | |
| 83 | + d1 = json.loads(con.execute( | |
| 84 | + "SELECT details FROM restaurants WHERE uid='osm:node/1'") | |
| 85 | + .fetchone()["details"]) | |
| 86 | + pa = d1["permis_alcool"] | |
| 87 | + assert pa["categories"] == ["Bar", "Restaurant pour vendre"] | |
| 88 | + assert pa["capacite"] == 120 and pa["nb_permis"] == 2 | |
| 89 | + assert pa["matched_by"] == "nom+lieu" | |
| 90 | + | |
| 91 | + # adresse+nom (CP + civique + token « pizza/bella » partagé) | |
| 92 | + d2 = json.loads(con.execute( | |
| 93 | + "SELECT details FROM restaurants WHERE uid='osm:node/2'") | |
| 94 | + .fetchone()["details"]) | |
| 95 | + assert d2["permis_alcool"]["matched_by"] == "adresse+nom" | |
| 96 | + assert d2["permis_alcool"]["categories"] == ["Restaurant pour servir"] | |
| 97 | + | |
| 98 | + # l'homonyme d'une autre ville n'est PAS épinglé sur le resto de Gatineau | |
| 99 | + d3 = json.loads(con.execute( | |
| 100 | + "SELECT details FROM restaurants WHERE uid='osm:node/3'") | |
| 101 | + .fetchone()["details"] or "{}") | |
| 102 | + assert "permis_alcool" not in d3 | |
added
tests/test_ubereats.py
+127 −0
@@ -0,0 +1,127 @@ | ||
| 1 | +# ============================================================================== | |
| 2 | +# Author: Simon-Pierre Boucher <contact@spboucher.ai> | |
| 3 | +# File: tests/test_ubereats.py | |
| 4 | +# Desc: Connecteur Uber Eats (menus livraison + notes) — décodage HORS LIGNE | |
| 5 | +# de l'état __REACT_QUERY_STATE__ (« " » en ", « \ » en %5C), | |
| 6 | +# construction du menu standard (prix en cents -> dollars, contexte | |
| 7 | +# delivery), vérification CONSERVATRICE d'un magasin (téléphone, | |
| 8 | +# postal+civique, GPS+nom) et slugs/candidats de découverte. | |
| 9 | +# ============================================================================== | |
| 10 | +import json | |
| 11 | + | |
| 12 | +from restoka.connectors.ubereats import (UberEatsConnector, _review_count, | |
| 13 | + build_menu, decode_state, slugify, | |
| 14 | + store_payload, verify_store) | |
| 15 | + | |
| 16 | +STORE = { | |
| 17 | + "title": "Chez Testo (Rosemont)", | |
| 18 | + "uuid": "a52d9225-dfe5-4f43-ac49-19833a61628f", | |
| 19 | + "slug": "chez-testo-rosemont", | |
| 20 | + "citySlug": "montreal", | |
| 21 | + "currencyCode": "CAD", | |
| 22 | + "rating": {"ratingValue": 4.3, "reviewCount": "15000+"}, | |
| 23 | + "priceBucket": "$$", | |
| 24 | + "phoneNumber": "+15142774385", | |
| 25 | + "location": { | |
| 26 | + "address": "1275, Boul. Rosemont, Montréal, QC H2S 3L9", | |
| 27 | + "streetAddress": "1275, Boul. Rosemont", | |
| 28 | + "city": "Montréal", "country": "CA", "postalCode": "H2S 3L9", | |
| 29 | + "region": "QC", "latitude": 45.5366491, "longitude": -73.5936433, | |
| 30 | + }, | |
| 31 | + "catalogSectionsMap": { | |
| 32 | + "cat-1": [ | |
| 33 | + {"payload": {"standardItemsPayload": { | |
| 34 | + "title": {"text": "Poutines"}, | |
| 35 | + "catalogItems": [ | |
| 36 | + {"title": "Poutine classique", "price": 1139, | |
| 37 | + "itemDescription": "Frites, sauce, fromage"}, | |
| 38 | + {"title": "Poutine galvaude", "price": 1499, | |
| 39 | + "itemDescription": ""}, | |
| 40 | + {"title": "Item sans prix", "price": 0, | |
| 41 | + "itemDescription": ""}, | |
| 42 | + ]}}}, | |
| 43 | + {"payload": {"standardItemsPayload": { | |
| 44 | + "title": {"text": "Breuvages"}, | |
| 45 | + "catalogItems": [ | |
| 46 | + {"title": "Liqueur", "price": 349, | |
| 47 | + "itemDescription": ""}, | |
| 48 | + ]}}}, | |
| 49 | + {"payload": {"standardItemsPayload": { | |
| 50 | + "title": {"text": "Section vide"}, "catalogItems": []}}}, | |
| 51 | + ], | |
| 52 | + }, | |
| 53 | +} | |
| 54 | + | |
| 55 | + | |
| 56 | +def _encode_state(state: dict) -> str: | |
| 57 | + """Encode comme sur ubereats.com : « " » -> \\u0022, « \\ » -> %5C.""" | |
| 58 | + blob = json.dumps(state, ensure_ascii=False) | |
| 59 | + blob = blob.replace("\\", "%5C").replace('"', "\\u0022") | |
| 60 | + return ('<html><script type="application/json" ' | |
| 61 | + f'id="__REACT_QUERY_STATE__">{blob}</script></html>') | |
| 62 | + | |
| 63 | + | |
| 64 | +def _state() -> dict: | |
| 65 | + return {"mutations": [], "queries": [ | |
| 66 | + {"state": {"data": {"isLoggedIn": False}}, | |
| 67 | + "queryKey": ["getUserV1", {}, None], | |
| 68 | + "queryHash": '[\\"getUserV1\\",{},null]'}, | |
| 69 | + {"state": {"data": STORE}, | |
| 70 | + "queryKey": ["getStoreV1", {"storeUuid": STORE["uuid"]}]}, | |
| 71 | + ]} | |
| 72 | + | |
| 73 | + | |
| 74 | +def test_decode_state_et_store_payload(): | |
| 75 | + page = _encode_state(_state()) | |
| 76 | + state = decode_state(page) | |
| 77 | + assert state and len(state["queries"]) == 2 | |
| 78 | + store = store_payload(state) | |
| 79 | + assert store and store["title"] == "Chez Testo (Rosemont)" | |
| 80 | + assert store["location"]["postalCode"] == "H2S 3L9" | |
| 81 | + assert _review_count(store) == 15000 | |
| 82 | + | |
| 83 | + | |
| 84 | +def test_build_menu_delivery(): | |
| 85 | + menu = build_menu(STORE, "2026-08-18T00:00:00Z") | |
| 86 | + assert menu["price_context"] == "delivery" | |
| 87 | + assert menu["price_source"] == "ubereats" | |
| 88 | + assert menu["currency"] == "CAD" | |
| 89 | + assert [s["name"] for s in menu["sections"]] == ["Poutines", "Breuvages"] | |
| 90 | + items = menu["sections"][0]["items"] | |
| 91 | + assert items[0]["price"] == 11.39 # cents -> dollars | |
| 92 | + assert items[2]["price"] is None # 0 = pas de prix | |
| 93 | + assert menu["sections"][1]["items"][0]["price"] == 3.49 | |
| 94 | + | |
| 95 | + | |
| 96 | +def _row(**kw): | |
| 97 | + base = {"uid": "osm:node/9", "name": "Chez Testo", | |
| 98 | + "address": "1275 Boulevard Rosemont", "city": "Montréal", | |
| 99 | + "postal_code": "H2S 3L9", "phone": "+15142774385", | |
| 100 | + "lat": 45.5366, "lng": -73.5936} | |
| 101 | + base.update(kw) | |
| 102 | + return base | |
| 103 | + | |
| 104 | + | |
| 105 | +def test_verify_store_conservateur(): | |
| 106 | + assert verify_store(_row(), STORE) == "telephone" | |
| 107 | + # sans téléphone : postal + civique | |
| 108 | + assert verify_store(_row(phone=""), STORE) == "postal+civique" | |
| 109 | + # sans téléphone ni postal : GPS <120 m + nom similaire | |
| 110 | + assert verify_store(_row(phone="", postal_code=""), STORE) == "gps+nom" | |
| 111 | + # un AUTRE resto au bout de la rue : rien ne matche -> jamais fusionné | |
| 112 | + autre = _row(phone="+15140000000", postal_code="H2S 9Z9", | |
| 113 | + address="2001 Boulevard Rosemont", name="Autre Bistro", | |
| 114 | + lat=45.55, lng=-73.60) | |
| 115 | + assert verify_store(autre, STORE) == "" | |
| 116 | + | |
| 117 | + | |
| 118 | +def test_slugify_et_candidats(): | |
| 119 | + assert slugify("McDonald's") == "mcdonalds" | |
| 120 | + assert slugify("Café Dépôt") == "cafe-depot" | |
| 121 | + assert slugify("Chez Testo") == "chez-testo" | |
| 122 | + slugs = sorted(["chez-testo", "chez-testo-rosemont", "chez-testons", | |
| 123 | + "mcdonalds-rosemont", "pizza-x"]) | |
| 124 | + cand = UberEatsConnector.candidates_for("Chez Testo", slugs) | |
| 125 | + assert cand == ["chez-testo", "chez-testo-rosemont"] | |
| 126 | + # nom trop court/générique : aucun candidat | |
| 127 | + assert UberEatsConnector.candidates_for("Ora", slugs) == [] | |
added
tests/test_yelp_scrape.py
+100 −0
@@ -0,0 +1,100 @@ | ||
| 1 | +# ============================================================================== | |
| 2 | +# Author: Simon-Pierre Boucher <contact@spboucher.ai> | |
| 3 | +# File: tests/test_yelp_scrape.py | |
| 4 | +# Desc: Connecteur yelp-scrape (avis/notes sans clé API) — parsing HORS | |
| 5 | +# LIGNE du cache Apollo embarqué dans une page de recherche yelp.ca | |
| 6 | +# (scripts data-apollo-state HTML-échappés) et croisement CONSERVATEUR | |
| 7 | +# nom + civique + ville (ambigu ou fermé = ignoré). | |
| 8 | +# ============================================================================== | |
| 9 | +import html | |
| 10 | +import json | |
| 11 | + | |
| 12 | +from restoka.connectors.yelp_scrape import (_city_key, match_business, | |
| 13 | + parse_apollo_businesses) | |
| 14 | + | |
| 15 | +# Cache Apollo minimal, encodé comme sur yelp.ca (" dans un commentaire) | |
| 16 | +_CACHE = { | |
| 17 | + "BusinessCategory:cat1": {"__typename": "BusinessCategory", | |
| 18 | + "encid": "cat1", "title": "Poutineries"}, | |
| 19 | + "BusinessLocation:loc1": { | |
| 20 | + "__typename": "BusinessLocation", "encid": "loc1", | |
| 21 | + "address": {"__typename": "BusinessAddress", | |
| 22 | + "city": "Quebec City", | |
| 23 | + "addressLine1": "640 Grande Allée E"}, | |
| 24 | + "neighborhoods": [], | |
| 25 | + }, | |
| 26 | + "Business:biz1": { | |
| 27 | + "__typename": "Business", "encid": "biz1", | |
| 28 | + "alias": "chez-ashton-québec-2", "name": "Chez Ashton", | |
| 29 | + "rating": 3.5, "reviewCount": 42, | |
| 30 | + "categories": [{"__ref": "BusinessCategory:cat1"}], | |
| 31 | + "priceRange": {"__typename": "PriceRange", "display": "$"}, | |
| 32 | + "location": {"__ref": "BusinessLocation:loc1"}, | |
| 33 | + 'activeAlert({"deviceType":"WWW"})': None, | |
| 34 | + }, | |
| 35 | + "BusinessLocation:loc2": { | |
| 36 | + "__typename": "BusinessLocation", "encid": "loc2", | |
| 37 | + "address": {"__typename": "BusinessAddress", | |
| 38 | + "city": "Levis", "addressLine1": "5430 Rue Wilfrid-Hallé"}, | |
| 39 | + "neighborhoods": [], | |
| 40 | + }, | |
| 41 | + "Business:biz2": { # fermé définitivement -> jamais croisé | |
| 42 | + "__typename": "Business", "encid": "biz2", | |
| 43 | + "alias": "chez-ashton-levis", "name": "Chez Ashton", | |
| 44 | + "rating": 4.0, "reviewCount": 3, | |
| 45 | + "categories": [], "priceRange": None, | |
| 46 | + "location": {"__ref": "BusinessLocation:loc2"}, | |
| 47 | + 'activeAlert({"deviceType":"WWW"})': {"__typename": "BusinessAlert", | |
| 48 | + "type": "permclosed"}, | |
| 49 | + }, | |
| 50 | +} | |
| 51 | + | |
| 52 | + | |
| 53 | +def _page(cache: dict) -> str: | |
| 54 | + blob = html.escape(json.dumps(cache, ensure_ascii=False), quote=True) | |
| 55 | + return ('<html><body><script data-apollo-state="x" ' | |
| 56 | + f'type="application/json"><!--{blob}--></script></body></html>') | |
| 57 | + | |
| 58 | + | |
| 59 | +def test_parse_apollo_businesses(): | |
| 60 | + biz = parse_apollo_businesses(_page(_CACHE)) | |
| 61 | + assert len(biz) == 2 | |
| 62 | + b1 = next(b for b in biz if b["alias"] == "chez-ashton-québec-2") | |
| 63 | + assert b1["rating"] == 3.5 and b1["review_count"] == 42 | |
| 64 | + assert b1["price"] == "$" and b1["categories"] == ["Poutineries"] | |
| 65 | + assert b1["city"] == "Quebec City" | |
| 66 | + assert b1["address"] == "640 Grande Allée E" | |
| 67 | + assert b1["closed"] is False | |
| 68 | + b2 = next(b for b in biz if b["alias"] == "chez-ashton-levis") | |
| 69 | + assert b2["closed"] is True | |
| 70 | + | |
| 71 | + | |
| 72 | +def test_city_key(): | |
| 73 | + assert _city_key("Quebec City") == "quebec" | |
| 74 | + assert _city_key("Québec") == "quebec" | |
| 75 | + assert _city_key("Montréal") == "montreal" | |
| 76 | + assert _city_key("Saint-Nicolas") == "saint nicolas" | |
| 77 | + | |
| 78 | + | |
| 79 | +def test_match_business_conservateur(): | |
| 80 | + candidates = parse_apollo_businesses(_page(_CACHE)) | |
| 81 | + resto = {"name": "Chez Ashton", "city": "Québec", | |
| 82 | + "address": "640 Grande Allée Est"} | |
| 83 | + hit = match_business(resto, candidates) | |
| 84 | + assert hit and hit["alias"] == "chez-ashton-québec-2" | |
| 85 | + | |
| 86 | + # mauvais civique -> aucun croisement (voisin de la même rue) | |
| 87 | + assert match_business({"name": "Chez Ashton", "city": "Québec", | |
| 88 | + "address": "700 Grande Allée Est"}, | |
| 89 | + candidates) is None | |
| 90 | + # le permclosed de Lévis n'est jamais retenu, même ville exacte | |
| 91 | + assert match_business({"name": "Chez Ashton", "city": "Lévis", | |
| 92 | + "address": "5430 Rue Wilfrid-Hallé"}, | |
| 93 | + candidates) is None | |
| 94 | + # ambiguïté (deux commerces distincts qui passent) -> None | |
| 95 | + cache2 = dict(_CACHE) | |
| 96 | + cache2["Business:biz3"] = dict(_CACHE["Business:biz1"], | |
| 97 | + encid="biz3", alias="ashton-restaurant", | |
| 98 | + name="Restaurant Ashton") | |
| 99 | + cands2 = parse_apollo_businesses(_page(cache2)) | |
| 100 | + assert match_business(resto, cands2) is None | |
| 101 | ||