SPB Git forge

spb/trouve-ka

Public ★ Pinned

Trouve-KA — moteur de recherche web indépendant, Québec-first. Crawler distribué, index OpenSearch, ranking bilingue, galerie d'images. En prod : www.trouve-ka.com

40commits 1branches 0releases
9.2 MBsize
maindefault branch
23 days agolast push
Python 51.9% TypeScript 32.7% JavaScript 4.8% CSS 4.1% Shell 2.9% HTML 1.8% SQL 1.5%

Robustesse long terme : auto-guérison, survie au reboot, backups croisés, rétention

scripts/ops : boot.sh (démarrage idempotent par rôle, LaunchDaemon RunAtLoad
sur les 7 nodes), watchdog.sh (M2M32, 2 min : santé API/web/PG/Redis/tunnels/
public + réparation locale et distante via clé ssh à commande forcée),
backup-pg.sh (pg_dump quotidien, rotation 7j+dimanches), backup-os.sh
(snapshots OpenSearch fs, rotation 7j), pull-backups.sh (réplication croisée
tirée entre M2M32 et M2M32b via clés tar-only), install-node.sh.

Résilience code : worker retry infini sur la connexion initiale (PG/OS),
boot.sh attend les tunnels avant le worker; scheduler : rétention quotidienne
(crawl_attempts 30 j, search_queries 180 j, par lots + VACUUM).

Sécurité : toutes les invocations ssh opérationnelles neutralisent
ControlMaster/agent (sinon le multiplexage d'une session humaine contourne
les clés restreintes). docs/RUNBOOK.md : pannes, restauration, déploiement.

Prouvé en réel : API tuée → ressuscitée en 40 s; reboot de m4ma → retour
complet sans intervention.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Simon-Pierre Boucher committed 1 mo ago (Aug 16, 2026) parent d034305

9 changed files +553 −2

added docs/RUNBOOK.md +82 −0
@@ -0,0 +1,82 @@
1 +# Trouve-KA — Runbook d'exploitation
2 +
3 +Author: Simon-Pierre Boucher — Contact: contact@spboucher.ai
4 +
5 +Architecture : voir `docs/architecture.md`. Résumé : **M2M32** (hub : Postgres+Redis
6 +docker, API/web/workers/scheduler/enrichment natifs pm2, ngrok), **M2M32b**
7 +(OpenSearch docker, heap 8 G), **m4mc** (embeddings+reranker), satellites de
8 +crawl (M2M32c, M4BP48, M1M32, m4ma). Tout le trafic inter-nodes passe par des
9 +**tunnels ssh système sur 127.0.0.1** (contournement du blocage « réseau
10 +local » macOS, qui filtre par binaire : python non signé bloqué, ssh/curl
11 +Apple exempts).
12 +
13 +## Auto-guérison (rien à faire dans la plupart des cas)
14 +
15 +| Mécanisme | Où | Cadence |
16 +|---|---|---|
17 +| `io.trouveka.boot` (boot.sh) | tous les nodes | au démarrage du node |
18 +| `io.trouveka.watchdog` (watchdog.sh) | M2M32 | toutes les 2 min |
19 +| Passe de flotte (heal tous les satellites) | M2M32 | ~30 min |
20 +| `io.trouveka.backup-pg` / `backup-os` / `pull` | M2M32 / M2M32b | quotidien 03:30–04:40 |
21 +| Rétention PG (crawl_attempts 30 j, search_queries 180 j) | scheduler tk-scheduler | quotidien |
22 +
23 +Prouvé en conditions réelles : API tuée → ressuscitée en 40 s; reboot complet
24 +de m4ma → tunnels+worker revenus sans intervention.
25 +
26 +Journaux : `~/trouveka-watchdog.log`, `~/trouveka-boot.log`,
27 +`~/trouveka-backups/backup.log`, `~/trouveka-launchd.log` sur chaque node.
28 +
29 +## Clés SSH dédiées (jamais de clé à accès complet sur un node)
30 +
31 +| Clé | Détenue par | Autorisée sur | Pouvoir exact |
32 +|---|---|---|---|
33 +| `trouveka_tunnel` (par node) | chaque node | M2M32, M2M32b | port-forwarding seulement |
34 +| `trouveka_heal` | M2M32 | tous les satellites | exécuter `boot.sh` uniquement |
35 +| `trouveka_pull` | M2M32 ↔ M2M32b | l'autre node | `tar` du dossier de backups uniquement |
36 +
37 +⚠️ Toute invocation ssh opérationnelle utilise `ControlMaster=no`,
38 +`IdentitiesOnly=yes`, `IdentityAgent=none` — sans quoi le multiplexage/agent
39 +d'une session humaine court-circuite les restrictions.
40 +
41 +## Pannes et remèdes
42 +
43 +- **Recherche en panne / site public muet** : attendre 2-4 min (watchdog).
44 + Sinon : `ssh M2M32 'sh ~/trouve-ka/scripts/ops/boot.sh'` puis
45 + `tail ~/trouveka-watchdog.log`.
46 +- **Un node ne crawle plus** : `ssh M2M32` puis
47 + `/usr/bin/ssh -i ~/.ssh/trouveka_heal -o IdentitiesOnly=yes -o IdentityAgent=none -o ControlMaster=no -o ControlPath=none simon-pierreboucher@<IP> heal`
48 + (IP dans `~/.trouveka-fleet`).
49 +- **Après reboot d'un node : Tailscale absent** (app GUI, exige une session).
50 + Le moteur n'en dépend PAS (tout passe par le LAN). Pour retrouver l'accès
51 + distant direct : ouvrir une session (écran/VNC) une fois, ou passer par
52 + `ssh -J M2M32 simon-pierreboucher@<IP LAN>`.
53 +- **Transactions Postgres zombies** (verrous, autovacuum bloqué) :
54 + `SELECT pg_terminate_backend(pid) FROM pg_stat_activity WHERE usename='trouveka' AND now()-xact_start > interval '5 minutes';`
55 + (les timeouts serveur limitent déjà ça à ≤5 min).
56 +
57 +## Restauration après sinistre
58 +
59 +- **Postgres perdu** (dumps : `~/trouveka-backups/pg/` sur M2M32, miroir sur
60 + M2M32b `~/trouveka-backups/pg-mirror/`) :
61 + `docker exec -i trouve-ka-postgres-1 pg_restore -U trouveka -d trouveka --clean --if-exists < trouveka-YYYYMMDD.dump`
62 +- **OpenSearch perdu** (snapshots : `~/trouveka-os-snapshots/` sur M2M32b,
63 + miroir sur M2M32 `~/trouveka-backups/os-mirror/`) : recréer le container
64 + avec `path.repo=/mnt/snapshots` + le volume, puis
65 + `PUT _snapshot/trouveka-fs` (même corps que backup-os.sh) et
66 + `POST _snapshot/trouveka-fs/daily-YYYYMMDD/_restore`.
67 + Si les deux copies sont perdues : le moteur se reconstruit par recrawl
68 + (Postgres garde URLs/domaines/priorités — c'est lui le capital).
69 +- **Un node de crawl perdu** : rien à restaurer — réinstaller le rôle :
70 + `SUDO_PW=... sh ~/trouve-ka/scripts/ops/install-node.sh crawler` + clé
71 + tunnel (voir mémoire projet).
72 +
73 +## Déploiement de code
74 +
75 +```
76 +rsync -az --delete --exclude .env --exclude .venv --exclude .venv-embed \
77 + --exclude node_modules --exclude .next --exclude __pycache__ --exclude .git \
78 + /Users/simon-pierreboucher/Desktop/trouve-ka/ <node>:trouve-ka/ # CHEMIN ABSOLU !
79 +ssh M2M32 'pm2 restart tk-api tk-crawler-1 tk-crawler-2 tk-enrichment tk-scheduler'
80 +# web : cd ~/trouve-ka/apps/web && pnpm build && pm2 restart tk-web
81 +# satellites : heal (boot.sh relance le worker avec le nouveau code après pkill)
82 +```
added scripts/ops/backup-os.sh +42 −0
@@ -0,0 +1,42 @@
1 +#!/bin/sh
2 +# Trouve-KA — snapshot quotidien OpenSearch (node search M2M32b)
3 +# Author: Simon-Pierre Boucher
4 +# Contact: contact@spboucher.ai
5 +#
6 +# Snapshots natifs OpenSearch (incrémentaux entre eux) dans un dossier hôte
7 +# monté dans le container (path.repo=/mnt/snapshots). Rotation 7 jours.
8 +# Restauration : POST _snapshot/trouveka-fs/<snap>/_restore
9 +# (index fermé ou renommé via rename_pattern au besoin).
10 +
11 +PATH="$HOME/.local/bin:/opt/homebrew/bin:/usr/local/bin:/usr/bin:/bin:/usr/sbin:/sbin"
12 +export PATH
13 +LOG="$HOME/trouveka-backups/backup.log"
14 +STAMP=$(date '+%Y%m%d')
15 +OS=http://localhost:9200
16 +
17 +mkdir -p "$HOME/trouveka-backups"
18 +log() { echo "$(date '+%F %T') [os] $*" >> "$LOG"; }
19 +
20 +# Dépôt de snapshots (idempotent)
21 +curl -s -m 20 -X PUT "$OS/_snapshot/trouveka-fs" -H 'Content-Type: application/json' \
22 + -d '{"type":"fs","settings":{"location":"/mnt/snapshots","compress":true}}' >/dev/null 2>&1
23 +
24 +RES=$(curl -s -m 590 -X PUT "$OS/_snapshot/trouveka-fs/daily-$STAMP?wait_for_completion=true" \
25 + -H 'Content-Type: application/json' \
26 + -d '{"indices":"trouveka-docs-v2","include_global_state":false}')
27 +case "$RES" in
28 + *'"state":"SUCCESS"'*) log "snapshot daily-$STAMP OK" ;;
29 + *already*exists*) log "snapshot daily-$STAMP déjà présent" ;;
30 + *) log "ÉCHEC snapshot : $(echo "$RES" | head -c 200)"; exit 1 ;;
31 +esac
32 +
33 +# Rotation : supprimer les snapshots > 7 jours (via l'API, jamais le filesystem)
34 +CUTOFF=$(date -v-7d '+%Y%m%d' 2>/dev/null || date -d '7 days ago' '+%Y%m%d')
35 +for snap in $(curl -s -m 30 "$OS/_snapshot/trouveka-fs/_all" | grep -o '"snapshot":"daily-[0-9]*"' | grep -o 'daily-[0-9]*'); do
36 + d=${snap#daily-}
37 + if [ "$d" -lt "$CUTOFF" ] 2>/dev/null; then
38 + curl -s -m 60 -X DELETE "$OS/_snapshot/trouveka-fs/$snap" >/dev/null
39 + log "rotation : $snap supprimé"
40 + fi
41 +done
42 +exit 0
added scripts/ops/backup-pg.sh +43 −0
@@ -0,0 +1,43 @@
1 +#!/bin/sh
2 +# Trouve-KA — sauvegarde quotidienne Postgres (hub M2M32)
3 +# Author: Simon-Pierre Boucher
4 +# Contact: contact@spboucher.ai
5 +#
6 +# pg_dump format custom (compressé, restaurable sélectivement) →
7 +# ~/trouveka-backups/pg/ + copie croisée sur M2M32b (un node ≠ un backup).
8 +# Rotation : 7 quotidiennes locales; les dimanches sont gardés 28 jours.
9 +# Restauration : pg_restore -d trouveka --clean --if-exists <fichier>
10 +
11 +PATH="$HOME/.local/bin:/opt/homebrew/bin:/usr/local/bin:/usr/bin:/bin:/usr/sbin:/sbin"
12 +export PATH
13 +DIR="$HOME/trouveka-backups/pg"
14 +LOG="$HOME/trouveka-backups/backup.log"
15 +STAMP=$(date '+%Y%m%d')
16 +OUT="$DIR/trouveka-$STAMP.dump"
17 +
18 +mkdir -p "$DIR"
19 +log() { echo "$(date '+%F %T') [pg] $*" >> "$LOG"; }
20 +
21 +if ! docker exec trouve-ka-postgres-1 pg_dump -U trouveka -d trouveka -Fc > "$OUT.tmp" 2>>"$LOG"; then
22 + log "ÉCHEC pg_dump"
23 + rm -f "$OUT.tmp"
24 + exit 1
25 +fi
26 +mv "$OUT.tmp" "$OUT"
27 +SIZE=$(du -h "$OUT" | cut -f1)
28 +log "dump OK : $OUT ($SIZE)"
29 +
30 +# Rotation : quotidiennes > 7 jours supprimées, sauf dimanches gardés 28 jours
31 +find "$DIR" -name 'trouveka-*.dump' -mtime +7 | while read -r f; do
32 + d=$(basename "$f" | sed 's/trouveka-\([0-9]*\)\.dump/\1/')
33 + dow=$(date -j -f '%Y%m%d' "$d" '+%u' 2>/dev/null || echo 1)
34 + if [ "$dow" = "7" ]; then
35 + find "$f" -mtime +28 -delete 2>/dev/null
36 + else
37 + rm -f "$f"
38 + fi
39 +done
40 +
41 +# La copie croisée est TIRÉE par M2M32b (pull-backups.sh + clé à commande
42 +# forcée tar) — ce script n'a aucun accès sortant à assurer.
43 +exit 0
added scripts/ops/boot.sh +119 −0
@@ -0,0 +1,119 @@
1 +#!/bin/sh
2 +# Trouve-KA — démarrage/réparation idempotent d'un node (exécutable à répétition)
3 +# Author: Simon-Pierre Boucher
4 +# Contact: contact@spboucher.ai
5 +#
6 +# Rôle du node lu dans ~/.trouveka-role : hub | search | embed | crawler
7 +# hub (M2M32) : PG+Redis (docker), tunnels OS/embed, pm2 tk-*, ngrok
8 +# search (M2M32b) : OpenSearch (docker), tunnels PG/Redis, worker crawl
9 +# embed (m4mc) : service embeddings/rerank, tunnels, worker crawl
10 +# crawler (autres) : tunnels PG/Redis/OS, worker crawl
11 +#
12 +# Conçu pour tourner sous LaunchDaemon (RunAtLoad) ET à la main ET via la clé
13 +# de guérison du watchdog. Ne touche à rien qui fonctionne déjà.
14 +# Règle d'or : les process python ne parlent qu'à 127.0.0.1 (tunnels ssh
15 +# système = seuls flux LAN, exempts du blocage « réseau local » macOS).
16 +
17 +PATH="$HOME/.local/bin:/opt/homebrew/bin:/usr/local/bin:/usr/bin:/bin:/usr/sbin:/sbin"
18 +export PATH
19 +ROLE=$(cat "$HOME/.trouveka-role" 2>/dev/null || echo crawler)
20 +LOG="$HOME/trouveka-boot.log"
21 +TK="$HOME/trouve-ka"
22 +SSH_OPTS="-i $HOME/.ssh/trouveka_tunnel -N -o ControlMaster=no -o ControlPath=none -o IdentitiesOnly=yes -o IdentityAgent=none -o ForwardAgent=no -o StrictHostKeyChecking=accept-new -o ServerAliveInterval=30 -o ServerAliveCountMax=3 -o ExitOnForwardFailure=yes"
23 +PG_HOST=192.168.2.90
24 +OS_HOST=192.168.2.77
25 +EMBED_HOST=192.168.2.75
26 +
27 +log() { echo "$(date '+%F %T') [$ROLE] $*" >> "$LOG"; }
28 +
29 +port_up() { nc -z -w 2 127.0.0.1 "$1" 2>/dev/null; }
30 +
31 +# Tunnel simple : ensure_tunnel <nom> <port_local> <hôte> <port_distant>
32 +ensure_tunnel() {
33 + port_up "$2" && return 0
34 + pkill -f "trouveka-tunloop-$1" 2>/dev/null
35 + sleep 1
36 + nohup sh -c "while true; do /usr/bin/ssh $SSH_OPTS -L 127.0.0.1:$2:localhost:$4 simon-pierreboucher@$3; sleep 5; done # trouveka-tunloop-$1" \
37 + >> "$HOME/trouveka-tunnels.log" 2>&1 &
38 + log "tunnel $1 (127.0.0.1:$2 → $3:$4) relancé"
39 +}
40 +
41 +# Tunnel double PG+Redis vers le hub
42 +ensure_pg_redis_tunnel() {
43 + port_up 15432 && port_up 16379 && return 0
44 + pkill -f "trouveka-tunloop-pg" 2>/dev/null
45 + sleep 1
46 + nohup sh -c "while true; do /usr/bin/ssh $SSH_OPTS -L 127.0.0.1:15432:localhost:5432 -L 127.0.0.1:16379:localhost:6379 simon-pierreboucher@$PG_HOST; sleep 5; done # trouveka-tunloop-pg" \
47 + >> "$HOME/trouveka-tunnels.log" 2>&1 &
48 + log "tunnel pg+redis relancé"
49 +}
50 +
51 +wait_port() { # $1 = port, $2 = secondes max
52 + t=0
53 + while [ "$t" -lt "${2:-60}" ]; do
54 + port_up "$1" && return 0
55 + sleep 2
56 + t=$((t+2))
57 + done
58 + log "attente port $1 : délai dépassé (${2:-60}s)"
59 + return 1
60 +}
61 +
62 +ensure_worker() {
63 + pgrep -f trouveka.crawler.worker >/dev/null && return 0
64 + [ -x "$TK/.venv/bin/python" ] || { log "worker : venv absent"; return 1; }
65 + # Au boot : laisser les tunnels s'établir avant de lancer le worker
66 + wait_port 15432 60
67 + cd "$TK" && nohup .venv/bin/python -m trouveka.crawler.worker >> "$HOME/trouveka-crawler.log" 2>&1 &
68 + log "worker crawl relancé"
69 +}
70 +
71 +ensure_colima() {
72 + docker info >/dev/null 2>&1 && return 0
73 + log "colima/docker arrêté : démarrage"
74 + colima start >> "$LOG" 2>&1
75 +}
76 +
77 +case "$ROLE" in
78 + hub)
79 + ensure_colima
80 + cd "$TK" && docker compose up -d postgres redis >> "$LOG" 2>&1
81 + ensure_tunnel search 9210 "$OS_HOST" 9200
82 + ensure_tunnel embed 8191 "$EMBED_HOST" 8091
83 + # pm2 : ressusciter la liste sauvegardée puis relancer tout process en erreur
84 + if ! pm2 pid tk-api >/dev/null 2>&1 || [ -z "$(pm2 pid tk-api 2>/dev/null)" ]; then
85 + pm2 resurrect >> "$LOG" 2>&1
86 + log "pm2 resurrect"
87 + fi
88 + ERRORED=$(pm2 jlist 2>/dev/null | grep -o '"name":"[^"]*","pm2_env":{"status":"errored"' | grep -o 'tk-[a-z0-9-]*')
89 + for name in $ERRORED; do pm2 restart "$name" >> "$LOG" 2>&1; log "pm2 restart $name (errored)"; done
90 + # ngrok : LaunchAgent gui absent après reboot sans session → filet nohup
91 + if ! pgrep -f "ngrok http" >/dev/null; then
92 + nohup ngrok http --url=www.trouve-ka.com 3000 >> "$HOME/trouve-ka-ngrok.log" 2>&1 &
93 + log "ngrok relancé"
94 + fi
95 + ;;
96 + search)
97 + ensure_colima
98 + docker start trouveka-opensearch >/dev/null 2>&1
99 + ensure_pg_redis_tunnel
100 + ensure_worker
101 + ;;
102 + embed)
103 + if ! pgrep -f "services.embedding.server" >/dev/null; then
104 + cd "$TK" && nohup .venv-embed/bin/uvicorn services.embedding.server:app --host 0.0.0.0 --port 8091 >> "$HOME/trouveka-embed.log" 2>&1 &
105 + log "service embeddings relancé"
106 + fi
107 + ensure_pg_redis_tunnel
108 + ensure_tunnel os 19200 "$OS_HOST" 9200
109 + ensure_worker
110 + ;;
111 + crawler|*)
112 + ensure_pg_redis_tunnel
113 + ensure_tunnel os 19200 "$OS_HOST" 9200
114 + ensure_worker
115 + ;;
116 +esac
117 +
118 +log "boot.sh terminé"
119 +exit 0
added scripts/ops/install-node.sh +57 −0
@@ -0,0 +1,57 @@
1 +#!/bin/sh
2 +# Trouve-KA — installation de la robustesse sur un node (rôle + LaunchDaemons)
3 +# Author: Simon-Pierre Boucher
4 +# Contact: contact@spboucher.ai
5 +#
6 +# Usage : SUDO_PW=... sh install-node.sh <hub|search|embed|crawler>
7 +# Idempotent. Installe :
8 +# - ~/.trouveka-role
9 +# - io.trouveka.boot (RunAtLoad : boot.sh au démarrage du node)
10 +# - hub : + io.trouveka.watchdog (2 min) + backup-pg (03:30) + pull (04:40)
11 +# - search : + backup-os (04:00) + pull (04:30)
12 +
13 +set -eu
14 +ROLE="${1:?role requis: hub|search|embed|crawler}"
15 +USER_NAME=$(whoami)
16 +HOME_DIR="$HOME"
17 +OPS="$HOME_DIR/trouve-ka/scripts/ops"
18 +
19 +echo "$ROLE" > "$HOME_DIR/.trouveka-role"
20 +
21 +sudo_cmd() { echo "${SUDO_PW:?SUDO_PW requis}" | sudo -S "$@" 2>/dev/null; }
22 +
23 +install_daemon() { # $1=label $2=program-args-xml $3=extra-xml
24 + PLIST="/tmp/io.trouveka.$1.plist"
25 + cat > "$PLIST" <<EOF
26 +<?xml version="1.0" encoding="UTF-8"?>
27 +<!DOCTYPE plist PUBLIC "-//Apple//DTD PLIST 1.0//EN" "http://www.apple.com/DTDs/PropertyList-1.0.dtd">
28 +<plist version="1.0"><dict>
29 + <key>Label</key><string>io.trouveka.$1</string>
30 + <key>ProgramArguments</key><array><string>/bin/sh</string><string>$2</string></array>
31 + <key>UserName</key><string>$USER_NAME</string>
32 + <key>StandardErrorPath</key><string>$HOME_DIR/trouveka-launchd.log</string>
33 + <key>StandardOutPath</key><string>$HOME_DIR/trouveka-launchd.log</string>
34 + $3
35 +</dict></plist>
36 +EOF
37 + sudo_cmd cp "$PLIST" "/Library/LaunchDaemons/io.trouveka.$1.plist"
38 + sudo_cmd launchctl bootout "system/io.trouveka.$1" 2>/dev/null || true
39 + sudo_cmd launchctl bootstrap system "/Library/LaunchDaemons/io.trouveka.$1.plist"
40 + echo "✓ daemon io.trouveka.$1"
41 +}
42 +
43 +install_daemon boot "$OPS/boot.sh" "<key>RunAtLoad</key><true/>"
44 +
45 +case "$ROLE" in
46 + hub)
47 + install_daemon watchdog "$OPS/watchdog.sh" "<key>StartInterval</key><integer>120</integer>"
48 + install_daemon backup-pg "$OPS/backup-pg.sh" "<key>StartCalendarInterval</key><dict><key>Hour</key><integer>3</integer><key>Minute</key><integer>30</integer></dict>"
49 + install_daemon pull "$OPS/pull-backups.sh" "<key>StartCalendarInterval</key><dict><key>Hour</key><integer>4</integer><key>Minute</key><integer>40</integer></dict>"
50 + ;;
51 + search)
52 + install_daemon backup-os "$OPS/backup-os.sh" "<key>StartCalendarInterval</key><dict><key>Hour</key><integer>4</integer><key>Minute</key><integer>0</integer></dict>"
53 + install_daemon pull "$OPS/pull-backups.sh" "<key>StartCalendarInterval</key><dict><key>Hour</key><integer>4</integer><key>Minute</key><integer>30</integer></dict>"
54 + ;;
55 +esac
56 +
57 +echo "✓ node installé (rôle: $ROLE)"
added scripts/ops/pull-backups.sh +39 −0
@@ -0,0 +1,39 @@
1 +#!/bin/sh
2 +# Trouve-KA — réplication croisée des sauvegardes (modèle pull)
3 +# Author: Simon-Pierre Boucher
4 +# Contact: contact@spboucher.ai
5 +#
6 +# Chaque node de données TIRE les sauvegardes de l'autre via une clé ssh à
7 +# commande forcée (le détenteur de la clé ne peut QUE recevoir le tar du
8 +# dossier de backups distant — aucun shell, aucun autre accès).
9 +# hub (M2M32) : tire les snapshots OpenSearch de M2M32b
10 +# search (M2M32b) : tire les dumps Postgres de M2M32
11 +# Ainsi la perte totale d'un node laisse toujours une copie ≤ 24 h sur l'autre.
12 +
13 +PATH="$HOME/.local/bin:/opt/homebrew/bin:/usr/local/bin:/usr/bin:/bin:/usr/sbin:/sbin"
14 +export PATH
15 +ROLE=$(cat "$HOME/.trouveka-role" 2>/dev/null || echo crawler)
16 +LOG="$HOME/trouveka-backups/backup.log"
17 +KEY="$HOME/.ssh/trouveka_pull"
18 +
19 +mkdir -p "$HOME/trouveka-backups"
20 +log() { echo "$(date '+%F %T') [pull-$ROLE] $*" >> "$LOG"; }
21 +
22 +pull() { # $1 = hôte distant, $2 = dossier local de destination
23 + mkdir -p "$2"
24 + if /usr/bin/ssh -i "$KEY" -o ControlMaster=no -o ControlPath=none -o IdentitiesOnly=yes -o IdentityAgent=none -o ForwardAgent=no -o ConnectTimeout=10 -o StrictHostKeyChecking=accept-new \
25 + "simon-pierreboucher@$1" pull 2>>"$LOG" | tar -xf - -C "$2" 2>>"$LOG"; then
26 + log "réplication depuis $1 OK ($(ls "$2" | wc -l | tr -d ' ') fichiers)"
27 + else
28 + log "ÉCHEC réplication depuis $1"
29 + fi
30 + # Rotation locale du miroir (garder 14 jours de fichiers)
31 + find "$2" -type f -mtime +14 -delete 2>/dev/null
32 +}
33 +
34 +case "$ROLE" in
35 + hub) pull 192.168.2.77 "$HOME/trouveka-backups/os-mirror" ;;
36 + search) pull 192.168.2.90 "$HOME/trouveka-backups/pg-mirror" ;;
37 + *) log "aucun rôle de réplication" ;;
38 +esac
39 +exit 0
added scripts/ops/watchdog.sh +111 −0
@@ -0,0 +1,111 @@
1 +#!/bin/sh
2 +# Trouve-KA — watchdog auto-guérisseur (tourne sur le hub M2M32 toutes les 2 min)
3 +# Author: Simon-Pierre Boucher
4 +# Contact: contact@spboucher.ai
5 +#
6 +# Vérifie chaque maillon et répare sans intervention humaine :
7 +# - API (health + search_ok), web, Postgres, Redis (localhost)
8 +# - tunnels OpenSearch (9210) et embeddings (8191)
9 +# - tunnel public ngrok (vue de l'extérieur)
10 +# - toutes les 30 min : passe de guérison sur toute la flotte satellite
11 +# (clé à commande forcée → exécute uniquement boot.sh distant)
12 +# Journal : ~/trouveka-watchdog.log (tronqué automatiquement).
13 +
14 +PATH="$HOME/.local/bin:/opt/homebrew/bin:/usr/local/bin:/usr/bin:/bin:/usr/sbin:/sbin"
15 +export PATH
16 +LOG="$HOME/trouveka-watchdog.log"
17 +touch "$LOG"
18 +STATE="$HOME/.trouveka-watchdog-count"
19 +HEAL_KEY="$HOME/.ssh/trouveka_heal"
20 +# Une IP LAN par ligne — générée à l'installation (scripts/ops/README)
21 +SATELLITES=$(cat "$HOME/.trouveka-fleet" 2>/dev/null)
22 +
23 +log() { echo "$(date '+%F %T') $*" >> "$LOG"; }
24 +
25 +heal_local() {
26 + log "HEAL local : $1"
27 + /bin/sh "$HOME/trouve-ka/scripts/ops/boot.sh"
28 +}
29 +
30 +heal_remote() { # $1 = ip
31 + log "HEAL distant : $1"
32 + /usr/bin/ssh -i "$HEAL_KEY" -o ControlMaster=no -o ControlPath=none -o IdentitiesOnly=yes -o IdentityAgent=none -o ForwardAgent=no -o ConnectTimeout=8 -o StrictHostKeyChecking=accept-new \
33 + "simon-pierreboucher@$1" heal >> "$LOG" 2>&1 || log "HEAL $1 : injoignable"
34 +}
35 +
36 +FAILURES=0
37 +
38 +# --- Postgres / Redis (docker local)
39 +nc -z -w 3 127.0.0.1 5432 || { heal_local "postgres injoignable"; FAILURES=$((FAILURES+1)); }
40 +nc -z -w 3 127.0.0.1 6379 || { heal_local "redis injoignable"; FAILURES=$((FAILURES+1)); }
41 +
42 +# --- Tunnels
43 +if ! curl -s -m 5 http://127.0.0.1:9210/_cluster/health | grep -q '"status"'; then
44 + heal_local "tunnel/OpenSearch 9210 muet"
45 + sleep 3
46 + curl -s -m 5 http://127.0.0.1:9210/_cluster/health | grep -q '"status"' || heal_remote 192.168.2.77
47 + FAILURES=$((FAILURES+1))
48 +fi
49 +if ! curl -s -m 5 http://127.0.0.1:8191/health | grep -q '"ok"'; then
50 + heal_local "tunnel/embeddings 8191 muet"
51 + sleep 3
52 + curl -s -m 5 http://127.0.0.1:8191/health | grep -q '"ok"' || heal_remote 192.168.2.75
53 + FAILURES=$((FAILURES+1))
54 +fi
55 +
56 +# --- API : santé + moteur de recherche réellement joignable
57 +HEALTH=$(curl -s -m 8 http://127.0.0.1:8080/api/health)
58 +case "$HEALTH" in
59 + *'"ok":true'*'"search_ok":true'*) : ;;
60 + *'"search_ok":false'*)
61 + log "API up mais search_ok=false"
62 + pm2 restart tk-api >> "$LOG" 2>&1
63 + FAILURES=$((FAILURES+1)) ;;
64 + *)
65 + log "API muette"
66 + pm2 restart tk-api >> "$LOG" 2>&1 || heal_local "API morte"
67 + FAILURES=$((FAILURES+1)) ;;
68 +esac
69 +
70 +# --- Web
71 +WEB=$(curl -s -m 8 -o /dev/null -w "%{http_code}" http://127.0.0.1:3000)
72 +if [ "$WEB" != "200" ]; then
73 + log "web local $WEB"
74 + pm2 restart tk-web >> "$LOG" 2>&1
75 + FAILURES=$((FAILURES+1))
76 +fi
77 +
78 +# --- Public (ngrok) : vue extérieure
79 +PUB=$(curl -s -m 12 -o /dev/null -w "%{http_code}" https://www.trouve-ka.com/)
80 +if [ "$PUB" = "000" ]; then
81 + log "public 000 : redémarrage ngrok"
82 + pkill -f "ngrok http" 2>/dev/null
83 + sleep 2
84 + nohup ngrok http --url=www.trouve-ka.com 3000 >> "$HOME/trouve-ka-ngrok.log" 2>&1 &
85 + FAILURES=$((FAILURES+1))
86 +fi
87 +
88 +# --- pm2 en erreur
89 +ERRORED=$(pm2 jlist 2>/dev/null | grep -o '"name":"[^"]*","pm2_env":{"status":"errored"' | grep -o 'tk-[a-z0-9-]*')
90 +for name in $ERRORED; do
91 + log "pm2 $name errored : restart"
92 + pm2 restart "$name" >> "$LOG" 2>&1
93 + FAILURES=$((FAILURES+1))
94 +done
95 +
96 +# --- Passe de flotte toutes les 15 exécutions (~30 min) : boot.sh partout
97 +COUNT=$(cat "$STATE" 2>/dev/null || echo 0)
98 +COUNT=$((COUNT+1))
99 +echo "$COUNT" > "$STATE"
100 +if [ $((COUNT % 15)) -eq 0 ]; then
101 + log "passe de flotte (#$COUNT)"
102 + for ip in $SATELLITES; do heal_remote "$ip"; done
103 +fi
104 +
105 +[ "$FAILURES" -eq 0 ] || log "cycle terminé : $FAILURES réparation(s)"
106 +
107 +# --- Rotation du journal (garder ~5000 lignes)
108 +if [ "$(wc -l < "$LOG" 2>/dev/null || echo 0)" -gt 10000 ]; then
109 + tail -5000 "$LOG" > "$LOG.tmp" && mv "$LOG.tmp" "$LOG"
110 +fi
111 +exit 0
modified services/crawler/worker.py +17 −2
@@ -66,9 +66,24 @@ class CrawlerWorker:
66 66 # ------------------------------------------------------------------ cycle de vie
67 67
68 68 async def start(self) -> None:
69 − await self.db.connect()
69 + # Connexion initiale résiliente : au boot d'un node, les tunnels ssh
70 + # vers PG/OpenSearch peuvent mettre quelques secondes — mourir ici
71 + # laisserait le node sans worker jusqu'à la passe du watchdog (§13).
72 + while True:
73 + try:
74 + await self.db.connect()
75 + break
76 + except Exception:
77 + log.warning("Postgres indisponible au démarrage, retry dans 10 s")
78 + await asyncio.sleep(10)
70 79 self.robots = RobotsCache(self.db, self.client, self.s.crawler_user_agent)
71 − await self.search.ensure_index()
80 + while True:
81 + try:
82 + await self.search.ensure_index()
83 + break
84 + except Exception:
85 + log.warning("OpenSearch indisponible au démarrage, retry dans 10 s")
86 + await asyncio.sleep(10)
72 87 log.info("worker démarré", extra={"ctx": {"worker_id": self.worker_id}})
73 88
74 89 loop = asyncio.get_running_loop()
modified services/scheduler/loop.py +43 −0
@@ -20,6 +20,13 @@ log = get_logger("scheduler")
20 20
21 21 STALE_RESET_INTERVAL = 60 # secondes
22 22 AUTHORITY_INTERVAL = 15 * 60 # secondes
23 +RETENTION_INTERVAL = 24 * 3600 # secondes — purge quotidienne
24 +
25 +# Rétention (soutenabilité long terme, §12 stockage économique) :
26 +# l'historique fin de crawl vieillit vite; les agrégats vivent dans domains.
27 +RETENTION_CRAWL_ATTEMPTS_DAYS = 30
28 +RETENTION_SEARCH_QUERIES_DAYS = 180
29 +RETENTION_DELETE_BATCH = 50_000
23 30
24 31
25 32 class SchedulerLoop:
@@ -47,6 +54,37 @@ class SchedulerLoop:
47 54 )
48 55 log.info("autorité de domaine recalculée")
49 56
57 + async def retention(self) -> None:
58 + """Purge quotidienne : borne la croissance des tables d'historique.
59 +
60 + Suppression par lots (jamais de DELETE géant qui verrouille la table)
61 + et VACUUM ANALYZE hors statement_timeout pour rendre l'espace."""
62 + for table, column, days in (
63 + ("crawl_attempts", "fetched_at", RETENTION_CRAWL_ATTEMPTS_DAYS),
64 + ("search_queries", "created_at", RETENTION_SEARCH_QUERIES_DAYS),
65 + ):
66 + total = 0
67 + while True:
68 + result = await self.db.pool.execute(
69 + f"""
70 + DELETE FROM {table} WHERE id IN (
71 + SELECT id FROM {table}
72 + WHERE {column} < now() - interval '{days} days'
73 + LIMIT {RETENTION_DELETE_BATCH}
74 + )
75 + """
76 + )
77 + deleted = int(result.split()[-1])
78 + total += deleted
79 + if deleted < RETENTION_DELETE_BATCH:
80 + break
81 + await asyncio.sleep(2) # laisser respirer les workers
82 + if total:
83 + async with self.db.pool.acquire() as conn:
84 + await conn.execute("SET statement_timeout = 0")
85 + await conn.execute(f"VACUUM (ANALYZE) {table}")
86 + log.info("rétention appliquée", extra={"ctx": {"table": table, "deleted": total}})
87 +
50 88 async def start(self) -> None:
51 89 await self.db.connect()
52 90 loop = asyncio.get_running_loop()
@@ -55,6 +93,7 @@ class SchedulerLoop:
55 93 log.info("scheduler démarré")
56 94
57 95 elapsed_authority = AUTHORITY_INTERVAL # premier calcul immédiat
96 + elapsed_retention = RETENTION_INTERVAL - 600 # première purge ~10 min après démarrage
58 97 while not self.stop_event.is_set():
59 98 try:
60 99 reset = await self.db.reset_stale_items(older_than_minutes=30)
@@ -63,6 +102,9 @@ class SchedulerLoop:
63 102 if elapsed_authority >= AUTHORITY_INTERVAL:
64 103 await self.recompute_authority()
65 104 elapsed_authority = 0
105 + if elapsed_retention >= RETENTION_INTERVAL:
106 + await self.retention()
107 + elapsed_retention = 0
66 108 except Exception:
67 109 log.exception("erreur scheduler (on continue)")
68 110 try:
@@ -70,6 +112,7 @@ class SchedulerLoop:
70 112 except TimeoutError:
71 113 pass
72 114 elapsed_authority += STALE_RESET_INTERVAL
115 + elapsed_retention += STALE_RESET_INTERVAL
73 116 await self.db.close()
74 117
75 118
76 119