-
Thu, Sep 10, 2026 5
-
UI proxy: keep Content-Encoding so compressed Next.js responses render in browsers
…
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
-
Harvester: correct MLX weight estimates from HF metadata, flag installed mirrors; docs: managed-firewall WireGuard relay
…
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
-
Harvester: prefer the Apple Silicon runtime, MoE-aware KV fallback, size-constrained starter slots; README perf table
…
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
-
Harvester: skip repositories without weight files instead of aborting the scan
…
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
-
LLM API v0.1.0 — private OpenAI-compatible local model server for Apple Silicon
…
FastAPI gateway + SQLite registry, MLX-LM / mlx-vlm and llama.cpp workers, load-on-demand with memory policy and eviction, OpenAI endpoints (chat/completions/embeddings/rerank, streaming), management API, Model Harvester, benchmarks, API keys, Next.js 16 console, tests and docs. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
-