Robustesse long terme : auto-guérison, survie au reboot, backups croisés, rétention
scripts/ops : boot.sh (démarrage idempotent par rôle, LaunchDaemon RunAtLoad sur les 7 nodes), watchdog.sh (M2M32, 2 min : santé API/web/PG/Redis/tunnels/ public + réparation locale et distante via clé ssh à commande forcée), backup-pg.sh (pg_dump quotidien, rotation 7j+dimanches), backup-os.sh (snapshots OpenSearch fs, rotation 7j), pull-backups.sh (réplication croisée tirée entre M2M32 et M2M32b via clés tar-only), install-node.sh. Résilience code : worker retry infini sur la connexion initiale (PG/OS), boot.sh attend les tunnels avant le worker; scheduler : rétention quotidienne (crawl_attempts 30 j, search_queries 180 j, par lots + VACUUM). Sécurité : toutes les invocations ssh opérationnelles neutralisent ControlMaster/agent (sinon le multiplexage d'une session humaine contourne les clés restreintes). docs/RUNBOOK.md : pannes, restauration, déploiement. Prouvé en réel : API tuée → ressuscitée en 40 s; reboot de m4ma → retour complet sans intervention. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
9 changed files +553 −2
added
docs/RUNBOOK.md
+82 −0
@@ -0,0 +1,82 @@ | ||
| 1 | +# Trouve-KA — Runbook d'exploitation | |
| 2 | + | |
| 3 | +Author: Simon-Pierre Boucher — Contact: contact@spboucher.ai | |
| 4 | + | |
| 5 | +Architecture : voir `docs/architecture.md`. Résumé : **M2M32** (hub : Postgres+Redis | |
| 6 | +docker, API/web/workers/scheduler/enrichment natifs pm2, ngrok), **M2M32b** | |
| 7 | +(OpenSearch docker, heap 8 G), **m4mc** (embeddings+reranker), satellites de | |
| 8 | +crawl (M2M32c, M4BP48, M1M32, m4ma). Tout le trafic inter-nodes passe par des | |
| 9 | +**tunnels ssh système sur 127.0.0.1** (contournement du blocage « réseau | |
| 10 | +local » macOS, qui filtre par binaire : python non signé bloqué, ssh/curl | |
| 11 | +Apple exempts). | |
| 12 | + | |
| 13 | +## Auto-guérison (rien à faire dans la plupart des cas) | |
| 14 | + | |
| 15 | +| Mécanisme | Où | Cadence | | |
| 16 | +|---|---|---| | |
| 17 | +| `io.trouveka.boot` (boot.sh) | tous les nodes | au démarrage du node | | |
| 18 | +| `io.trouveka.watchdog` (watchdog.sh) | M2M32 | toutes les 2 min | | |
| 19 | +| Passe de flotte (heal tous les satellites) | M2M32 | ~30 min | | |
| 20 | +| `io.trouveka.backup-pg` / `backup-os` / `pull` | M2M32 / M2M32b | quotidien 03:30–04:40 | | |
| 21 | +| Rétention PG (crawl_attempts 30 j, search_queries 180 j) | scheduler tk-scheduler | quotidien | | |
| 22 | + | |
| 23 | +Prouvé en conditions réelles : API tuée → ressuscitée en 40 s; reboot complet | |
| 24 | +de m4ma → tunnels+worker revenus sans intervention. | |
| 25 | + | |
| 26 | +Journaux : `~/trouveka-watchdog.log`, `~/trouveka-boot.log`, | |
| 27 | +`~/trouveka-backups/backup.log`, `~/trouveka-launchd.log` sur chaque node. | |
| 28 | + | |
| 29 | +## Clés SSH dédiées (jamais de clé à accès complet sur un node) | |
| 30 | + | |
| 31 | +| Clé | Détenue par | Autorisée sur | Pouvoir exact | | |
| 32 | +|---|---|---|---| | |
| 33 | +| `trouveka_tunnel` (par node) | chaque node | M2M32, M2M32b | port-forwarding seulement | | |
| 34 | +| `trouveka_heal` | M2M32 | tous les satellites | exécuter `boot.sh` uniquement | | |
| 35 | +| `trouveka_pull` | M2M32 ↔ M2M32b | l'autre node | `tar` du dossier de backups uniquement | | |
| 36 | + | |
| 37 | +⚠️ Toute invocation ssh opérationnelle utilise `ControlMaster=no`, | |
| 38 | +`IdentitiesOnly=yes`, `IdentityAgent=none` — sans quoi le multiplexage/agent | |
| 39 | +d'une session humaine court-circuite les restrictions. | |
| 40 | + | |
| 41 | +## Pannes et remèdes | |
| 42 | + | |
| 43 | +- **Recherche en panne / site public muet** : attendre 2-4 min (watchdog). | |
| 44 | + Sinon : `ssh M2M32 'sh ~/trouve-ka/scripts/ops/boot.sh'` puis | |
| 45 | + `tail ~/trouveka-watchdog.log`. | |
| 46 | +- **Un node ne crawle plus** : `ssh M2M32` puis | |
| 47 | + `/usr/bin/ssh -i ~/.ssh/trouveka_heal -o IdentitiesOnly=yes -o IdentityAgent=none -o ControlMaster=no -o ControlPath=none simon-pierreboucher@<IP> heal` | |
| 48 | + (IP dans `~/.trouveka-fleet`). | |
| 49 | +- **Après reboot d'un node : Tailscale absent** (app GUI, exige une session). | |
| 50 | + Le moteur n'en dépend PAS (tout passe par le LAN). Pour retrouver l'accès | |
| 51 | + distant direct : ouvrir une session (écran/VNC) une fois, ou passer par | |
| 52 | + `ssh -J M2M32 simon-pierreboucher@<IP LAN>`. | |
| 53 | +- **Transactions Postgres zombies** (verrous, autovacuum bloqué) : | |
| 54 | + `SELECT pg_terminate_backend(pid) FROM pg_stat_activity WHERE usename='trouveka' AND now()-xact_start > interval '5 minutes';` | |
| 55 | + (les timeouts serveur limitent déjà ça à ≤5 min). | |
| 56 | + | |
| 57 | +## Restauration après sinistre | |
| 58 | + | |
| 59 | +- **Postgres perdu** (dumps : `~/trouveka-backups/pg/` sur M2M32, miroir sur | |
| 60 | + M2M32b `~/trouveka-backups/pg-mirror/`) : | |
| 61 | + `docker exec -i trouve-ka-postgres-1 pg_restore -U trouveka -d trouveka --clean --if-exists < trouveka-YYYYMMDD.dump` | |
| 62 | +- **OpenSearch perdu** (snapshots : `~/trouveka-os-snapshots/` sur M2M32b, | |
| 63 | + miroir sur M2M32 `~/trouveka-backups/os-mirror/`) : recréer le container | |
| 64 | + avec `path.repo=/mnt/snapshots` + le volume, puis | |
| 65 | + `PUT _snapshot/trouveka-fs` (même corps que backup-os.sh) et | |
| 66 | + `POST _snapshot/trouveka-fs/daily-YYYYMMDD/_restore`. | |
| 67 | + Si les deux copies sont perdues : le moteur se reconstruit par recrawl | |
| 68 | + (Postgres garde URLs/domaines/priorités — c'est lui le capital). | |
| 69 | +- **Un node de crawl perdu** : rien à restaurer — réinstaller le rôle : | |
| 70 | + `SUDO_PW=... sh ~/trouve-ka/scripts/ops/install-node.sh crawler` + clé | |
| 71 | + tunnel (voir mémoire projet). | |
| 72 | + | |
| 73 | +## Déploiement de code | |
| 74 | + | |
| 75 | +``` | |
| 76 | +rsync -az --delete --exclude .env --exclude .venv --exclude .venv-embed \ | |
| 77 | + --exclude node_modules --exclude .next --exclude __pycache__ --exclude .git \ | |
| 78 | + /Users/simon-pierreboucher/Desktop/trouve-ka/ <node>:trouve-ka/ # CHEMIN ABSOLU ! | |
| 79 | +ssh M2M32 'pm2 restart tk-api tk-crawler-1 tk-crawler-2 tk-enrichment tk-scheduler' | |
| 80 | +# web : cd ~/trouve-ka/apps/web && pnpm build && pm2 restart tk-web | |
| 81 | +# satellites : heal (boot.sh relance le worker avec le nouveau code après pkill) | |
| 82 | +``` | |
added
scripts/ops/backup-os.sh
+42 −0
@@ -0,0 +1,42 @@ | ||
| 1 | +#!/bin/sh | |
| 2 | +# Trouve-KA — snapshot quotidien OpenSearch (node search M2M32b) | |
| 3 | +# Author: Simon-Pierre Boucher | |
| 4 | +# Contact: contact@spboucher.ai | |
| 5 | +# | |
| 6 | +# Snapshots natifs OpenSearch (incrémentaux entre eux) dans un dossier hôte | |
| 7 | +# monté dans le container (path.repo=/mnt/snapshots). Rotation 7 jours. | |
| 8 | +# Restauration : POST _snapshot/trouveka-fs/<snap>/_restore | |
| 9 | +# (index fermé ou renommé via rename_pattern au besoin). | |
| 10 | + | |
| 11 | +PATH="$HOME/.local/bin:/opt/homebrew/bin:/usr/local/bin:/usr/bin:/bin:/usr/sbin:/sbin" | |
| 12 | +export PATH | |
| 13 | +LOG="$HOME/trouveka-backups/backup.log" | |
| 14 | +STAMP=$(date '+%Y%m%d') | |
| 15 | +OS=http://localhost:9200 | |
| 16 | + | |
| 17 | +mkdir -p "$HOME/trouveka-backups" | |
| 18 | +log() { echo "$(date '+%F %T') [os] $*" >> "$LOG"; } | |
| 19 | + | |
| 20 | +# Dépôt de snapshots (idempotent) | |
| 21 | +curl -s -m 20 -X PUT "$OS/_snapshot/trouveka-fs" -H 'Content-Type: application/json' \ | |
| 22 | + -d '{"type":"fs","settings":{"location":"/mnt/snapshots","compress":true}}' >/dev/null 2>&1 | |
| 23 | + | |
| 24 | +RES=$(curl -s -m 590 -X PUT "$OS/_snapshot/trouveka-fs/daily-$STAMP?wait_for_completion=true" \ | |
| 25 | + -H 'Content-Type: application/json' \ | |
| 26 | + -d '{"indices":"trouveka-docs-v2","include_global_state":false}') | |
| 27 | +case "$RES" in | |
| 28 | + *'"state":"SUCCESS"'*) log "snapshot daily-$STAMP OK" ;; | |
| 29 | + *already*exists*) log "snapshot daily-$STAMP déjà présent" ;; | |
| 30 | + *) log "ÉCHEC snapshot : $(echo "$RES" | head -c 200)"; exit 1 ;; | |
| 31 | +esac | |
| 32 | + | |
| 33 | +# Rotation : supprimer les snapshots > 7 jours (via l'API, jamais le filesystem) | |
| 34 | +CUTOFF=$(date -v-7d '+%Y%m%d' 2>/dev/null || date -d '7 days ago' '+%Y%m%d') | |
| 35 | +for snap in $(curl -s -m 30 "$OS/_snapshot/trouveka-fs/_all" | grep -o '"snapshot":"daily-[0-9]*"' | grep -o 'daily-[0-9]*'); do | |
| 36 | + d=${snap#daily-} | |
| 37 | + if [ "$d" -lt "$CUTOFF" ] 2>/dev/null; then | |
| 38 | + curl -s -m 60 -X DELETE "$OS/_snapshot/trouveka-fs/$snap" >/dev/null | |
| 39 | + log "rotation : $snap supprimé" | |
| 40 | + fi | |
| 41 | +done | |
| 42 | +exit 0 | |
added
scripts/ops/backup-pg.sh
+43 −0
@@ -0,0 +1,43 @@ | ||
| 1 | +#!/bin/sh | |
| 2 | +# Trouve-KA — sauvegarde quotidienne Postgres (hub M2M32) | |
| 3 | +# Author: Simon-Pierre Boucher | |
| 4 | +# Contact: contact@spboucher.ai | |
| 5 | +# | |
| 6 | +# pg_dump format custom (compressé, restaurable sélectivement) → | |
| 7 | +# ~/trouveka-backups/pg/ + copie croisée sur M2M32b (un node ≠ un backup). | |
| 8 | +# Rotation : 7 quotidiennes locales; les dimanches sont gardés 28 jours. | |
| 9 | +# Restauration : pg_restore -d trouveka --clean --if-exists <fichier> | |
| 10 | + | |
| 11 | +PATH="$HOME/.local/bin:/opt/homebrew/bin:/usr/local/bin:/usr/bin:/bin:/usr/sbin:/sbin" | |
| 12 | +export PATH | |
| 13 | +DIR="$HOME/trouveka-backups/pg" | |
| 14 | +LOG="$HOME/trouveka-backups/backup.log" | |
| 15 | +STAMP=$(date '+%Y%m%d') | |
| 16 | +OUT="$DIR/trouveka-$STAMP.dump" | |
| 17 | + | |
| 18 | +mkdir -p "$DIR" | |
| 19 | +log() { echo "$(date '+%F %T') [pg] $*" >> "$LOG"; } | |
| 20 | + | |
| 21 | +if ! docker exec trouve-ka-postgres-1 pg_dump -U trouveka -d trouveka -Fc > "$OUT.tmp" 2>>"$LOG"; then | |
| 22 | + log "ÉCHEC pg_dump" | |
| 23 | + rm -f "$OUT.tmp" | |
| 24 | + exit 1 | |
| 25 | +fi | |
| 26 | +mv "$OUT.tmp" "$OUT" | |
| 27 | +SIZE=$(du -h "$OUT" | cut -f1) | |
| 28 | +log "dump OK : $OUT ($SIZE)" | |
| 29 | + | |
| 30 | +# Rotation : quotidiennes > 7 jours supprimées, sauf dimanches gardés 28 jours | |
| 31 | +find "$DIR" -name 'trouveka-*.dump' -mtime +7 | while read -r f; do | |
| 32 | + d=$(basename "$f" | sed 's/trouveka-\([0-9]*\)\.dump/\1/') | |
| 33 | + dow=$(date -j -f '%Y%m%d' "$d" '+%u' 2>/dev/null || echo 1) | |
| 34 | + if [ "$dow" = "7" ]; then | |
| 35 | + find "$f" -mtime +28 -delete 2>/dev/null | |
| 36 | + else | |
| 37 | + rm -f "$f" | |
| 38 | + fi | |
| 39 | +done | |
| 40 | + | |
| 41 | +# La copie croisée est TIRÉE par M2M32b (pull-backups.sh + clé à commande | |
| 42 | +# forcée tar) — ce script n'a aucun accès sortant à assurer. | |
| 43 | +exit 0 | |
added
scripts/ops/boot.sh
+119 −0
@@ -0,0 +1,119 @@ | ||
| 1 | +#!/bin/sh | |
| 2 | +# Trouve-KA — démarrage/réparation idempotent d'un node (exécutable à répétition) | |
| 3 | +# Author: Simon-Pierre Boucher | |
| 4 | +# Contact: contact@spboucher.ai | |
| 5 | +# | |
| 6 | +# Rôle du node lu dans ~/.trouveka-role : hub | search | embed | crawler | |
| 7 | +# hub (M2M32) : PG+Redis (docker), tunnels OS/embed, pm2 tk-*, ngrok | |
| 8 | +# search (M2M32b) : OpenSearch (docker), tunnels PG/Redis, worker crawl | |
| 9 | +# embed (m4mc) : service embeddings/rerank, tunnels, worker crawl | |
| 10 | +# crawler (autres) : tunnels PG/Redis/OS, worker crawl | |
| 11 | +# | |
| 12 | +# Conçu pour tourner sous LaunchDaemon (RunAtLoad) ET à la main ET via la clé | |
| 13 | +# de guérison du watchdog. Ne touche à rien qui fonctionne déjà. | |
| 14 | +# Règle d'or : les process python ne parlent qu'à 127.0.0.1 (tunnels ssh | |
| 15 | +# système = seuls flux LAN, exempts du blocage « réseau local » macOS). | |
| 16 | + | |
| 17 | +PATH="$HOME/.local/bin:/opt/homebrew/bin:/usr/local/bin:/usr/bin:/bin:/usr/sbin:/sbin" | |
| 18 | +export PATH | |
| 19 | +ROLE=$(cat "$HOME/.trouveka-role" 2>/dev/null || echo crawler) | |
| 20 | +LOG="$HOME/trouveka-boot.log" | |
| 21 | +TK="$HOME/trouve-ka" | |
| 22 | +SSH_OPTS="-i $HOME/.ssh/trouveka_tunnel -N -o ControlMaster=no -o ControlPath=none -o IdentitiesOnly=yes -o IdentityAgent=none -o ForwardAgent=no -o StrictHostKeyChecking=accept-new -o ServerAliveInterval=30 -o ServerAliveCountMax=3 -o ExitOnForwardFailure=yes" | |
| 23 | +PG_HOST=192.168.2.90 | |
| 24 | +OS_HOST=192.168.2.77 | |
| 25 | +EMBED_HOST=192.168.2.75 | |
| 26 | + | |
| 27 | +log() { echo "$(date '+%F %T') [$ROLE] $*" >> "$LOG"; } | |
| 28 | + | |
| 29 | +port_up() { nc -z -w 2 127.0.0.1 "$1" 2>/dev/null; } | |
| 30 | + | |
| 31 | +# Tunnel simple : ensure_tunnel <nom> <port_local> <hôte> <port_distant> | |
| 32 | +ensure_tunnel() { | |
| 33 | + port_up "$2" && return 0 | |
| 34 | + pkill -f "trouveka-tunloop-$1" 2>/dev/null | |
| 35 | + sleep 1 | |
| 36 | + nohup sh -c "while true; do /usr/bin/ssh $SSH_OPTS -L 127.0.0.1:$2:localhost:$4 simon-pierreboucher@$3; sleep 5; done # trouveka-tunloop-$1" \ | |
| 37 | + >> "$HOME/trouveka-tunnels.log" 2>&1 & | |
| 38 | + log "tunnel $1 (127.0.0.1:$2 → $3:$4) relancé" | |
| 39 | +} | |
| 40 | + | |
| 41 | +# Tunnel double PG+Redis vers le hub | |
| 42 | +ensure_pg_redis_tunnel() { | |
| 43 | + port_up 15432 && port_up 16379 && return 0 | |
| 44 | + pkill -f "trouveka-tunloop-pg" 2>/dev/null | |
| 45 | + sleep 1 | |
| 46 | + nohup sh -c "while true; do /usr/bin/ssh $SSH_OPTS -L 127.0.0.1:15432:localhost:5432 -L 127.0.0.1:16379:localhost:6379 simon-pierreboucher@$PG_HOST; sleep 5; done # trouveka-tunloop-pg" \ | |
| 47 | + >> "$HOME/trouveka-tunnels.log" 2>&1 & | |
| 48 | + log "tunnel pg+redis relancé" | |
| 49 | +} | |
| 50 | + | |
| 51 | +wait_port() { # $1 = port, $2 = secondes max | |
| 52 | + t=0 | |
| 53 | + while [ "$t" -lt "${2:-60}" ]; do | |
| 54 | + port_up "$1" && return 0 | |
| 55 | + sleep 2 | |
| 56 | + t=$((t+2)) | |
| 57 | + done | |
| 58 | + log "attente port $1 : délai dépassé (${2:-60}s)" | |
| 59 | + return 1 | |
| 60 | +} | |
| 61 | + | |
| 62 | +ensure_worker() { | |
| 63 | + pgrep -f trouveka.crawler.worker >/dev/null && return 0 | |
| 64 | + [ -x "$TK/.venv/bin/python" ] || { log "worker : venv absent"; return 1; } | |
| 65 | + # Au boot : laisser les tunnels s'établir avant de lancer le worker | |
| 66 | + wait_port 15432 60 | |
| 67 | + cd "$TK" && nohup .venv/bin/python -m trouveka.crawler.worker >> "$HOME/trouveka-crawler.log" 2>&1 & | |
| 68 | + log "worker crawl relancé" | |
| 69 | +} | |
| 70 | + | |
| 71 | +ensure_colima() { | |
| 72 | + docker info >/dev/null 2>&1 && return 0 | |
| 73 | + log "colima/docker arrêté : démarrage" | |
| 74 | + colima start >> "$LOG" 2>&1 | |
| 75 | +} | |
| 76 | + | |
| 77 | +case "$ROLE" in | |
| 78 | + hub) | |
| 79 | + ensure_colima | |
| 80 | + cd "$TK" && docker compose up -d postgres redis >> "$LOG" 2>&1 | |
| 81 | + ensure_tunnel search 9210 "$OS_HOST" 9200 | |
| 82 | + ensure_tunnel embed 8191 "$EMBED_HOST" 8091 | |
| 83 | + # pm2 : ressusciter la liste sauvegardée puis relancer tout process en erreur | |
| 84 | + if ! pm2 pid tk-api >/dev/null 2>&1 || [ -z "$(pm2 pid tk-api 2>/dev/null)" ]; then | |
| 85 | + pm2 resurrect >> "$LOG" 2>&1 | |
| 86 | + log "pm2 resurrect" | |
| 87 | + fi | |
| 88 | + ERRORED=$(pm2 jlist 2>/dev/null | grep -o '"name":"[^"]*","pm2_env":{"status":"errored"' | grep -o 'tk-[a-z0-9-]*') | |
| 89 | + for name in $ERRORED; do pm2 restart "$name" >> "$LOG" 2>&1; log "pm2 restart $name (errored)"; done | |
| 90 | + # ngrok : LaunchAgent gui absent après reboot sans session → filet nohup | |
| 91 | + if ! pgrep -f "ngrok http" >/dev/null; then | |
| 92 | + nohup ngrok http --url=www.trouve-ka.com 3000 >> "$HOME/trouve-ka-ngrok.log" 2>&1 & | |
| 93 | + log "ngrok relancé" | |
| 94 | + fi | |
| 95 | + ;; | |
| 96 | + search) | |
| 97 | + ensure_colima | |
| 98 | + docker start trouveka-opensearch >/dev/null 2>&1 | |
| 99 | + ensure_pg_redis_tunnel | |
| 100 | + ensure_worker | |
| 101 | + ;; | |
| 102 | + embed) | |
| 103 | + if ! pgrep -f "services.embedding.server" >/dev/null; then | |
| 104 | + cd "$TK" && nohup .venv-embed/bin/uvicorn services.embedding.server:app --host 0.0.0.0 --port 8091 >> "$HOME/trouveka-embed.log" 2>&1 & | |
| 105 | + log "service embeddings relancé" | |
| 106 | + fi | |
| 107 | + ensure_pg_redis_tunnel | |
| 108 | + ensure_tunnel os 19200 "$OS_HOST" 9200 | |
| 109 | + ensure_worker | |
| 110 | + ;; | |
| 111 | + crawler|*) | |
| 112 | + ensure_pg_redis_tunnel | |
| 113 | + ensure_tunnel os 19200 "$OS_HOST" 9200 | |
| 114 | + ensure_worker | |
| 115 | + ;; | |
| 116 | +esac | |
| 117 | + | |
| 118 | +log "boot.sh terminé" | |
| 119 | +exit 0 | |
added
scripts/ops/install-node.sh
+57 −0
@@ -0,0 +1,57 @@ | ||
| 1 | +#!/bin/sh | |
| 2 | +# Trouve-KA — installation de la robustesse sur un node (rôle + LaunchDaemons) | |
| 3 | +# Author: Simon-Pierre Boucher | |
| 4 | +# Contact: contact@spboucher.ai | |
| 5 | +# | |
| 6 | +# Usage : SUDO_PW=... sh install-node.sh <hub|search|embed|crawler> | |
| 7 | +# Idempotent. Installe : | |
| 8 | +# - ~/.trouveka-role | |
| 9 | +# - io.trouveka.boot (RunAtLoad : boot.sh au démarrage du node) | |
| 10 | +# - hub : + io.trouveka.watchdog (2 min) + backup-pg (03:30) + pull (04:40) | |
| 11 | +# - search : + backup-os (04:00) + pull (04:30) | |
| 12 | + | |
| 13 | +set -eu | |
| 14 | +ROLE="${1:?role requis: hub|search|embed|crawler}" | |
| 15 | +USER_NAME=$(whoami) | |
| 16 | +HOME_DIR="$HOME" | |
| 17 | +OPS="$HOME_DIR/trouve-ka/scripts/ops" | |
| 18 | + | |
| 19 | +echo "$ROLE" > "$HOME_DIR/.trouveka-role" | |
| 20 | + | |
| 21 | +sudo_cmd() { echo "${SUDO_PW:?SUDO_PW requis}" | sudo -S "$@" 2>/dev/null; } | |
| 22 | + | |
| 23 | +install_daemon() { # $1=label $2=program-args-xml $3=extra-xml | |
| 24 | + PLIST="/tmp/io.trouveka.$1.plist" | |
| 25 | + cat > "$PLIST" <<EOF | |
| 26 | +<?xml version="1.0" encoding="UTF-8"?> | |
| 27 | +<!DOCTYPE plist PUBLIC "-//Apple//DTD PLIST 1.0//EN" "http://www.apple.com/DTDs/PropertyList-1.0.dtd"> | |
| 28 | +<plist version="1.0"><dict> | |
| 29 | + <key>Label</key><string>io.trouveka.$1</string> | |
| 30 | + <key>ProgramArguments</key><array><string>/bin/sh</string><string>$2</string></array> | |
| 31 | + <key>UserName</key><string>$USER_NAME</string> | |
| 32 | + <key>StandardErrorPath</key><string>$HOME_DIR/trouveka-launchd.log</string> | |
| 33 | + <key>StandardOutPath</key><string>$HOME_DIR/trouveka-launchd.log</string> | |
| 34 | + $3 | |
| 35 | +</dict></plist> | |
| 36 | +EOF | |
| 37 | + sudo_cmd cp "$PLIST" "/Library/LaunchDaemons/io.trouveka.$1.plist" | |
| 38 | + sudo_cmd launchctl bootout "system/io.trouveka.$1" 2>/dev/null || true | |
| 39 | + sudo_cmd launchctl bootstrap system "/Library/LaunchDaemons/io.trouveka.$1.plist" | |
| 40 | + echo "✓ daemon io.trouveka.$1" | |
| 41 | +} | |
| 42 | + | |
| 43 | +install_daemon boot "$OPS/boot.sh" "<key>RunAtLoad</key><true/>" | |
| 44 | + | |
| 45 | +case "$ROLE" in | |
| 46 | + hub) | |
| 47 | + install_daemon watchdog "$OPS/watchdog.sh" "<key>StartInterval</key><integer>120</integer>" | |
| 48 | + install_daemon backup-pg "$OPS/backup-pg.sh" "<key>StartCalendarInterval</key><dict><key>Hour</key><integer>3</integer><key>Minute</key><integer>30</integer></dict>" | |
| 49 | + install_daemon pull "$OPS/pull-backups.sh" "<key>StartCalendarInterval</key><dict><key>Hour</key><integer>4</integer><key>Minute</key><integer>40</integer></dict>" | |
| 50 | + ;; | |
| 51 | + search) | |
| 52 | + install_daemon backup-os "$OPS/backup-os.sh" "<key>StartCalendarInterval</key><dict><key>Hour</key><integer>4</integer><key>Minute</key><integer>0</integer></dict>" | |
| 53 | + install_daemon pull "$OPS/pull-backups.sh" "<key>StartCalendarInterval</key><dict><key>Hour</key><integer>4</integer><key>Minute</key><integer>30</integer></dict>" | |
| 54 | + ;; | |
| 55 | +esac | |
| 56 | + | |
| 57 | +echo "✓ node installé (rôle: $ROLE)" | |
added
scripts/ops/pull-backups.sh
+39 −0
@@ -0,0 +1,39 @@ | ||
| 1 | +#!/bin/sh | |
| 2 | +# Trouve-KA — réplication croisée des sauvegardes (modèle pull) | |
| 3 | +# Author: Simon-Pierre Boucher | |
| 4 | +# Contact: contact@spboucher.ai | |
| 5 | +# | |
| 6 | +# Chaque node de données TIRE les sauvegardes de l'autre via une clé ssh à | |
| 7 | +# commande forcée (le détenteur de la clé ne peut QUE recevoir le tar du | |
| 8 | +# dossier de backups distant — aucun shell, aucun autre accès). | |
| 9 | +# hub (M2M32) : tire les snapshots OpenSearch de M2M32b | |
| 10 | +# search (M2M32b) : tire les dumps Postgres de M2M32 | |
| 11 | +# Ainsi la perte totale d'un node laisse toujours une copie ≤ 24 h sur l'autre. | |
| 12 | + | |
| 13 | +PATH="$HOME/.local/bin:/opt/homebrew/bin:/usr/local/bin:/usr/bin:/bin:/usr/sbin:/sbin" | |
| 14 | +export PATH | |
| 15 | +ROLE=$(cat "$HOME/.trouveka-role" 2>/dev/null || echo crawler) | |
| 16 | +LOG="$HOME/trouveka-backups/backup.log" | |
| 17 | +KEY="$HOME/.ssh/trouveka_pull" | |
| 18 | + | |
| 19 | +mkdir -p "$HOME/trouveka-backups" | |
| 20 | +log() { echo "$(date '+%F %T') [pull-$ROLE] $*" >> "$LOG"; } | |
| 21 | + | |
| 22 | +pull() { # $1 = hôte distant, $2 = dossier local de destination | |
| 23 | + mkdir -p "$2" | |
| 24 | + if /usr/bin/ssh -i "$KEY" -o ControlMaster=no -o ControlPath=none -o IdentitiesOnly=yes -o IdentityAgent=none -o ForwardAgent=no -o ConnectTimeout=10 -o StrictHostKeyChecking=accept-new \ | |
| 25 | + "simon-pierreboucher@$1" pull 2>>"$LOG" | tar -xf - -C "$2" 2>>"$LOG"; then | |
| 26 | + log "réplication depuis $1 OK ($(ls "$2" | wc -l | tr -d ' ') fichiers)" | |
| 27 | + else | |
| 28 | + log "ÉCHEC réplication depuis $1" | |
| 29 | + fi | |
| 30 | + # Rotation locale du miroir (garder 14 jours de fichiers) | |
| 31 | + find "$2" -type f -mtime +14 -delete 2>/dev/null | |
| 32 | +} | |
| 33 | + | |
| 34 | +case "$ROLE" in | |
| 35 | + hub) pull 192.168.2.77 "$HOME/trouveka-backups/os-mirror" ;; | |
| 36 | + search) pull 192.168.2.90 "$HOME/trouveka-backups/pg-mirror" ;; | |
| 37 | + *) log "aucun rôle de réplication" ;; | |
| 38 | +esac | |
| 39 | +exit 0 | |
added
scripts/ops/watchdog.sh
+111 −0
@@ -0,0 +1,111 @@ | ||
| 1 | +#!/bin/sh | |
| 2 | +# Trouve-KA — watchdog auto-guérisseur (tourne sur le hub M2M32 toutes les 2 min) | |
| 3 | +# Author: Simon-Pierre Boucher | |
| 4 | +# Contact: contact@spboucher.ai | |
| 5 | +# | |
| 6 | +# Vérifie chaque maillon et répare sans intervention humaine : | |
| 7 | +# - API (health + search_ok), web, Postgres, Redis (localhost) | |
| 8 | +# - tunnels OpenSearch (9210) et embeddings (8191) | |
| 9 | +# - tunnel public ngrok (vue de l'extérieur) | |
| 10 | +# - toutes les 30 min : passe de guérison sur toute la flotte satellite | |
| 11 | +# (clé à commande forcée → exécute uniquement boot.sh distant) | |
| 12 | +# Journal : ~/trouveka-watchdog.log (tronqué automatiquement). | |
| 13 | + | |
| 14 | +PATH="$HOME/.local/bin:/opt/homebrew/bin:/usr/local/bin:/usr/bin:/bin:/usr/sbin:/sbin" | |
| 15 | +export PATH | |
| 16 | +LOG="$HOME/trouveka-watchdog.log" | |
| 17 | +touch "$LOG" | |
| 18 | +STATE="$HOME/.trouveka-watchdog-count" | |
| 19 | +HEAL_KEY="$HOME/.ssh/trouveka_heal" | |
| 20 | +# Une IP LAN par ligne — générée à l'installation (scripts/ops/README) | |
| 21 | +SATELLITES=$(cat "$HOME/.trouveka-fleet" 2>/dev/null) | |
| 22 | + | |
| 23 | +log() { echo "$(date '+%F %T') $*" >> "$LOG"; } | |
| 24 | + | |
| 25 | +heal_local() { | |
| 26 | + log "HEAL local : $1" | |
| 27 | + /bin/sh "$HOME/trouve-ka/scripts/ops/boot.sh" | |
| 28 | +} | |
| 29 | + | |
| 30 | +heal_remote() { # $1 = ip | |
| 31 | + log "HEAL distant : $1" | |
| 32 | + /usr/bin/ssh -i "$HEAL_KEY" -o ControlMaster=no -o ControlPath=none -o IdentitiesOnly=yes -o IdentityAgent=none -o ForwardAgent=no -o ConnectTimeout=8 -o StrictHostKeyChecking=accept-new \ | |
| 33 | + "simon-pierreboucher@$1" heal >> "$LOG" 2>&1 || log "HEAL $1 : injoignable" | |
| 34 | +} | |
| 35 | + | |
| 36 | +FAILURES=0 | |
| 37 | + | |
| 38 | +# --- Postgres / Redis (docker local) | |
| 39 | +nc -z -w 3 127.0.0.1 5432 || { heal_local "postgres injoignable"; FAILURES=$((FAILURES+1)); } | |
| 40 | +nc -z -w 3 127.0.0.1 6379 || { heal_local "redis injoignable"; FAILURES=$((FAILURES+1)); } | |
| 41 | + | |
| 42 | +# --- Tunnels | |
| 43 | +if ! curl -s -m 5 http://127.0.0.1:9210/_cluster/health | grep -q '"status"'; then | |
| 44 | + heal_local "tunnel/OpenSearch 9210 muet" | |
| 45 | + sleep 3 | |
| 46 | + curl -s -m 5 http://127.0.0.1:9210/_cluster/health | grep -q '"status"' || heal_remote 192.168.2.77 | |
| 47 | + FAILURES=$((FAILURES+1)) | |
| 48 | +fi | |
| 49 | +if ! curl -s -m 5 http://127.0.0.1:8191/health | grep -q '"ok"'; then | |
| 50 | + heal_local "tunnel/embeddings 8191 muet" | |
| 51 | + sleep 3 | |
| 52 | + curl -s -m 5 http://127.0.0.1:8191/health | grep -q '"ok"' || heal_remote 192.168.2.75 | |
| 53 | + FAILURES=$((FAILURES+1)) | |
| 54 | +fi | |
| 55 | + | |
| 56 | +# --- API : santé + moteur de recherche réellement joignable | |
| 57 | +HEALTH=$(curl -s -m 8 http://127.0.0.1:8080/api/health) | |
| 58 | +case "$HEALTH" in | |
| 59 | + *'"ok":true'*'"search_ok":true'*) : ;; | |
| 60 | + *'"search_ok":false'*) | |
| 61 | + log "API up mais search_ok=false" | |
| 62 | + pm2 restart tk-api >> "$LOG" 2>&1 | |
| 63 | + FAILURES=$((FAILURES+1)) ;; | |
| 64 | + *) | |
| 65 | + log "API muette" | |
| 66 | + pm2 restart tk-api >> "$LOG" 2>&1 || heal_local "API morte" | |
| 67 | + FAILURES=$((FAILURES+1)) ;; | |
| 68 | +esac | |
| 69 | + | |
| 70 | +# --- Web | |
| 71 | +WEB=$(curl -s -m 8 -o /dev/null -w "%{http_code}" http://127.0.0.1:3000) | |
| 72 | +if [ "$WEB" != "200" ]; then | |
| 73 | + log "web local $WEB" | |
| 74 | + pm2 restart tk-web >> "$LOG" 2>&1 | |
| 75 | + FAILURES=$((FAILURES+1)) | |
| 76 | +fi | |
| 77 | + | |
| 78 | +# --- Public (ngrok) : vue extérieure | |
| 79 | +PUB=$(curl -s -m 12 -o /dev/null -w "%{http_code}" https://www.trouve-ka.com/) | |
| 80 | +if [ "$PUB" = "000" ]; then | |
| 81 | + log "public 000 : redémarrage ngrok" | |
| 82 | + pkill -f "ngrok http" 2>/dev/null | |
| 83 | + sleep 2 | |
| 84 | + nohup ngrok http --url=www.trouve-ka.com 3000 >> "$HOME/trouve-ka-ngrok.log" 2>&1 & | |
| 85 | + FAILURES=$((FAILURES+1)) | |
| 86 | +fi | |
| 87 | + | |
| 88 | +# --- pm2 en erreur | |
| 89 | +ERRORED=$(pm2 jlist 2>/dev/null | grep -o '"name":"[^"]*","pm2_env":{"status":"errored"' | grep -o 'tk-[a-z0-9-]*') | |
| 90 | +for name in $ERRORED; do | |
| 91 | + log "pm2 $name errored : restart" | |
| 92 | + pm2 restart "$name" >> "$LOG" 2>&1 | |
| 93 | + FAILURES=$((FAILURES+1)) | |
| 94 | +done | |
| 95 | + | |
| 96 | +# --- Passe de flotte toutes les 15 exécutions (~30 min) : boot.sh partout | |
| 97 | +COUNT=$(cat "$STATE" 2>/dev/null || echo 0) | |
| 98 | +COUNT=$((COUNT+1)) | |
| 99 | +echo "$COUNT" > "$STATE" | |
| 100 | +if [ $((COUNT % 15)) -eq 0 ]; then | |
| 101 | + log "passe de flotte (#$COUNT)" | |
| 102 | + for ip in $SATELLITES; do heal_remote "$ip"; done | |
| 103 | +fi | |
| 104 | + | |
| 105 | +[ "$FAILURES" -eq 0 ] || log "cycle terminé : $FAILURES réparation(s)" | |
| 106 | + | |
| 107 | +# --- Rotation du journal (garder ~5000 lignes) | |
| 108 | +if [ "$(wc -l < "$LOG" 2>/dev/null || echo 0)" -gt 10000 ]; then | |
| 109 | + tail -5000 "$LOG" > "$LOG.tmp" && mv "$LOG.tmp" "$LOG" | |
| 110 | +fi | |
| 111 | +exit 0 | |
modified
services/crawler/worker.py
+17 −2
@@ -66,9 +66,24 @@ class CrawlerWorker: | ||
| 66 | 66 | # ------------------------------------------------------------------ cycle de vie |
| 67 | 67 | |
| 68 | 68 | async def start(self) -> None: |
| 69 | − await self.db.connect() | |
| 69 | + # Connexion initiale résiliente : au boot d'un node, les tunnels ssh | |
| 70 | + # vers PG/OpenSearch peuvent mettre quelques secondes — mourir ici | |
| 71 | + # laisserait le node sans worker jusqu'à la passe du watchdog (§13). | |
| 72 | + while True: | |
| 73 | + try: | |
| 74 | + await self.db.connect() | |
| 75 | + break | |
| 76 | + except Exception: | |
| 77 | + log.warning("Postgres indisponible au démarrage, retry dans 10 s") | |
| 78 | + await asyncio.sleep(10) | |
| 70 | 79 | self.robots = RobotsCache(self.db, self.client, self.s.crawler_user_agent) |
| 71 | − await self.search.ensure_index() | |
| 80 | + while True: | |
| 81 | + try: | |
| 82 | + await self.search.ensure_index() | |
| 83 | + break | |
| 84 | + except Exception: | |
| 85 | + log.warning("OpenSearch indisponible au démarrage, retry dans 10 s") | |
| 86 | + await asyncio.sleep(10) | |
| 72 | 87 | log.info("worker démarré", extra={"ctx": {"worker_id": self.worker_id}}) |
| 73 | 88 | |
| 74 | 89 | loop = asyncio.get_running_loop() |
modified
services/scheduler/loop.py
+43 −0
@@ -20,6 +20,13 @@ log = get_logger("scheduler") | ||
| 20 | 20 | |
| 21 | 21 | STALE_RESET_INTERVAL = 60 # secondes |
| 22 | 22 | AUTHORITY_INTERVAL = 15 * 60 # secondes |
| 23 | +RETENTION_INTERVAL = 24 * 3600 # secondes — purge quotidienne | |
| 24 | + | |
| 25 | +# Rétention (soutenabilité long terme, §12 stockage économique) : | |
| 26 | +# l'historique fin de crawl vieillit vite; les agrégats vivent dans domains. | |
| 27 | +RETENTION_CRAWL_ATTEMPTS_DAYS = 30 | |
| 28 | +RETENTION_SEARCH_QUERIES_DAYS = 180 | |
| 29 | +RETENTION_DELETE_BATCH = 50_000 | |
| 23 | 30 | |
| 24 | 31 | |
| 25 | 32 | class SchedulerLoop: |
@@ -47,6 +54,37 @@ class SchedulerLoop: | ||
| 47 | 54 | ) |
| 48 | 55 | log.info("autorité de domaine recalculée") |
| 49 | 56 | |
| 57 | + async def retention(self) -> None: | |
| 58 | + """Purge quotidienne : borne la croissance des tables d'historique. | |
| 59 | + | |
| 60 | + Suppression par lots (jamais de DELETE géant qui verrouille la table) | |
| 61 | + et VACUUM ANALYZE hors statement_timeout pour rendre l'espace.""" | |
| 62 | + for table, column, days in ( | |
| 63 | + ("crawl_attempts", "fetched_at", RETENTION_CRAWL_ATTEMPTS_DAYS), | |
| 64 | + ("search_queries", "created_at", RETENTION_SEARCH_QUERIES_DAYS), | |
| 65 | + ): | |
| 66 | + total = 0 | |
| 67 | + while True: | |
| 68 | + result = await self.db.pool.execute( | |
| 69 | + f""" | |
| 70 | + DELETE FROM {table} WHERE id IN ( | |
| 71 | + SELECT id FROM {table} | |
| 72 | + WHERE {column} < now() - interval '{days} days' | |
| 73 | + LIMIT {RETENTION_DELETE_BATCH} | |
| 74 | + ) | |
| 75 | + """ | |
| 76 | + ) | |
| 77 | + deleted = int(result.split()[-1]) | |
| 78 | + total += deleted | |
| 79 | + if deleted < RETENTION_DELETE_BATCH: | |
| 80 | + break | |
| 81 | + await asyncio.sleep(2) # laisser respirer les workers | |
| 82 | + if total: | |
| 83 | + async with self.db.pool.acquire() as conn: | |
| 84 | + await conn.execute("SET statement_timeout = 0") | |
| 85 | + await conn.execute(f"VACUUM (ANALYZE) {table}") | |
| 86 | + log.info("rétention appliquée", extra={"ctx": {"table": table, "deleted": total}}) | |
| 87 | + | |
| 50 | 88 | async def start(self) -> None: |
| 51 | 89 | await self.db.connect() |
| 52 | 90 | loop = asyncio.get_running_loop() |
@@ -55,6 +93,7 @@ class SchedulerLoop: | ||
| 55 | 93 | log.info("scheduler démarré") |
| 56 | 94 | |
| 57 | 95 | elapsed_authority = AUTHORITY_INTERVAL # premier calcul immédiat |
| 96 | + elapsed_retention = RETENTION_INTERVAL - 600 # première purge ~10 min après démarrage | |
| 58 | 97 | while not self.stop_event.is_set(): |
| 59 | 98 | try: |
| 60 | 99 | reset = await self.db.reset_stale_items(older_than_minutes=30) |
@@ -63,6 +102,9 @@ class SchedulerLoop: | ||
| 63 | 102 | if elapsed_authority >= AUTHORITY_INTERVAL: |
| 64 | 103 | await self.recompute_authority() |
| 65 | 104 | elapsed_authority = 0 |
| 105 | + if elapsed_retention >= RETENTION_INTERVAL: | |
| 106 | + await self.retention() | |
| 107 | + elapsed_retention = 0 | |
| 66 | 108 | except Exception: |
| 67 | 109 | log.exception("erreur scheduler (on continue)") |
| 68 | 110 | try: |
@@ -70,6 +112,7 @@ class SchedulerLoop: | ||
| 70 | 112 | except TimeoutError: |
| 71 | 113 | pass |
| 72 | 114 | elapsed_authority += STALE_RESET_INTERVAL |
| 115 | + elapsed_retention += STALE_RESET_INTERVAL | |
| 73 | 116 | await self.db.close() |
| 74 | 117 | |
| 75 | 118 | |
| 76 | 119 | |