SPB Git

spb/zyquo-local Public MIT

Native macOS AI chat that runs LLMs 100% locally on Apple Silicon with MLX — no cloud, no API keys.

Swift 97.2% Shell 1.8% Makefile 1%

phase0: build recipe resolved — CLT + standalone Metal toolchain, SDK pin

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
simon-pierre boucher committed 11 days ago (Jul 30, 2026)

Showing 4 changed files with +428 and −0

added .gitignore +5 −0
@@ -0,0 +1,5 @@
1 +.build/
2 +dist/
3 +*.icns.tmp/
4 +.DS_Store
5 +Package.resolved
added ZyquoLocal-CLAUDE.md +325 −0
@@ -0,0 +1,325 @@
1 +# CLAUDE.md — Zyquo Local
2 +
3 +## Project Identity
4 +
5 +**Zyquo Local** is the on-device sibling of **Zyquo Cloud**: a legendary, native macOS AI chat client written in **Swift + SwiftUI**, built **entirely without Xcode** where possible (Swift Package Manager + command-line toolchain), that runs large language models **100% locally on Apple Silicon using MLX**. No API keys, no network calls for inference, no data ever leaving the Mac.
6 +
7 +The core promise: the user opens the app, **browses Hugging Face directly inside the interface, downloads MLX models with one click**, and chats with them — with the same legendary design language and polish as Zyquo Cloud. Zyquo Local must feel like the definitive local-LLM app for the Mac: faster, cleaner, and more beautiful than LM Studio or Ollama frontends.
8 +
9 +**Naming conventions (use these consistently everywhere):**
10 +- Display name / product name: `Zyquo Local`
11 +- App bundle: `Zyquo Local.app`
12 +- Bundle identifier: `com.zyquo.local`
13 +- Executable / SPM target: `ZyquoLocal` (no space)
14 +- Data folder: `~/Library/Application Support/ZyquoLocal/`
15 +- Models storage: `~/Library/Application Support/ZyquoLocal/Models/`
16 +- Repo module prefix in file headers: `Zyquo Local`
17 +- **Platform: Apple Silicon (arm64) ONLY.** MLX requires Apple Silicon. At launch, detect Intel Macs and show a clear, polite unsupported-hardware screen.
18 +
19 +---
20 +
21 +## 📋 MANDATORY FILE HEADER — EVERY CODE FILE
22 +
23 +**Every single code file you write** (all `.swift` files, plus `Makefile`, shell scripts, `Package.swift`, verification scripts — anything containing code) **MUST begin with this header comment**, adapted to the file's comment syntax:
24 +
25 +```swift
26 +//
27 +// <FileName>.swift
28 +// Zyquo Local
29 +//
30 +// Author: Simon-Pierre Boucher
31 +// Mail: contact@spboucher.ai
32 +//
33 +```
34 +
35 +For shell scripts / Makefiles:
36 +
37 +```bash
38 +#
39 +# <filename>
40 +# Zyquo Local
41 +#
42 +# Author: Simon-Pierre Boucher
43 +# Mail: contact@spboucher.ai
44 +#
45 +```
46 +
47 +No exceptions. If you ever create or refactor a file and the header is missing, add it. Before declaring the project done, run a sweep over the repository to verify every code file carries the header.
48 +
49 +---
50 +
51 +## 🧭 METHODOLOGY — WORK METHODICALLY, KEEP EVERYTHING COHERENT
52 +
53 +You must execute this project **strictly in phase order (0 → 8)**. Do not jump ahead, do not interleave phases, do not start the UI before the inference engine compiles and generates text, and do not write engine code before Phase 0 research is complete.
54 +
55 +**Working rules:**
56 +
57 +1. **One phase at a time.** At the start of each phase, write a short plan (checklist) into `docs/PLAN.md`; check items off as you complete them. At the end of each phase, perform a **phase checkpoint**: build the project (`swift build`), run what's runnable, fix all warnings/errors, and write a 3–5 line phase summary in `docs/PLAN.md` before moving on.
58 +2. **Phase gates:** Phase 0 is complete only when `docs/MLX-RESEARCH.md` and `docs/MODELS.md` are complete (see Phase 0). Phase 2 is complete only when a minimal CLI proof-of-concept loads one small MLX model and streams generated tokens to stdout. Phase 3 is complete only when a model can be searched, downloaded with live progress, resumed after interruption, and deleted. Phase 4 spec is the contract for all UI in Phase 6. Phase 7 is complete only when the model verification table is fully green. Phase 8 is complete only when `spctl` says "Notarized Developer ID".
59 +3. **Single source of truth, everywhere:**
60 + - Curated model data → only from `ModelCatalog` (generated from `docs/MODELS.md`). Live Hugging Face results come from `HubService` only.
61 + - Colors, fonts, spacing, radii → only from `ZyquoTheme` design tokens. Zero raw hex values or magic numbers in views.
62 + - All inference behavior → only in the `Engine/` layer, never leaking into ViewModels or Views.
63 + - Product naming → only per the conventions above. Never `Zyquo` alone, never `ZyquoLocal` in user-facing text.
64 +4. **Coherence sweeps:** after Phases 3, 6, and 8, do a dedicated consistency pass over the whole codebase: naming conventions uniform (types `UpperCamelCase`, one term per concept — always `LocalModel`, never a mix of `Model`/`LLM`/`LocalModel`), no dead code, file headers present, folder structure matches Phase 2 exactly.
65 +5. **Compile early, compile often.** Never accumulate more than one file of unbuilt changes. If the build breaks, fixing it is the immediate priority.
66 +6. **Commit discipline:** one logical unit per commit, message prefixed by phase (e.g., `phase3: resumable downloads with progress`). Never commit model weights or large binaries.
67 +7. **When research and reality disagree** (e.g., a model in `docs/MODELS.md` fails to load in Phase 7), update `docs/MODELS.md` AND `ModelCatalog` together — the two must never drift apart.
68 +
69 +---
70 +
71 +## ⚠️ PHASE 0 — MANDATORY INTENSIVE WEB RESEARCH (DO THIS FIRST, BEFORE WRITING ANY CODE)
72 +
73 +Do NOT rely on your training data — MLX Swift APIs, package structure, and the Hugging Face model landscape change fast. Research the CURRENT state of everything below using official sources (github.com/ml-explore/mlx-swift, github.com/ml-explore/mlx-swift-examples, huggingface.co/docs, huggingface.co/mlx-community, github.com/huggingface/swift-transformers). Produce TWO research documents before writing any code.
74 +
75 +### 0.A — `docs/MLX-RESEARCH.md` — the MLX Swift stack
76 +
77 +Document precisely:
78 +
79 +1. **Current MLX Swift packages and how to depend on them via SPM**: `mlx-swift` (core: MLX, MLXNN, MLXRandom, MLXFast…) and the LLM layer (`MLXLLM` / `MLXLMCommon` / `MLXVLM` — verify their current home: mlx-swift-examples repo or a dedicated package), exact repository URLs, current stable versions/tags, and minimum macOS requirement.
80 +2. **The exact current API** for: loading a model from a local directory (`ModelContainer` / `loadModelContainer` / factory APIs — verify names), tokenization, applying chat templates, streaming token-by-token generation (callback or `AsyncSequence`), generation parameters supported (temperature, topP, repetitionPenalty, maxTokens, seed…), KV cache handling for multi-turn chat, and how to unload/free a model.
81 +3. **Model directory format** MLX expects on disk: `config.json`, `*.safetensors` (possibly sharded + `model.safetensors.index.json`), `tokenizer.json` / `tokenizer_config.json`, chat template location. Which files are strictly required.
82 +4. **Supported architectures** in the current MLX Swift LLM layer (Llama, Qwen2/2.5/3, Mistral, Gemma/2/3, Phi-3/4, DeepSeek/distills, SmolLM, OpenELM, etc. — verify the actual list) so the app can warn before downloading an unsupported model.
83 +5. **⚠️ The no-Xcode build question — resolve this definitively:** MLX contains Metal kernels. Verify whether `swift build` works with only Command Line Tools, or whether the Metal shader compiler requires full Xcode (`xcrun -sdk macosx metal`). Test it. Document the finding and the working recipe in `docs/BUILD.md`. If full Xcode (or its Metal toolchain component) turns out to be strictly required for compiling MLX's shaders, the rule adapts to: **build via `swift build` / `xcodebuild` from the command line only — never the Xcode IDE, no `.xcodeproj` authored by hand** — and the Makefile must automate everything end-to-end regardless.
84 +6. **swift-transformers / Hub API in Swift**: what `huggingface/swift-transformers` offers for Hub downloads and tokenizers, and whether to use it or implement `HubService` directly on the HTTP API (document both paths, pick one, justify in the doc).
85 +7. **Memory model**: unified memory implications, MLX GPU memory/cache limits (`MLX.GPU.set(cacheLimit:)` etc.), and how to estimate RAM needed for a model (≈ weights size + KV cache + overhead).
86 +
87 +### 0.B — `docs/MODELS.md` — the curated model catalog + Hub integration
88 +
89 +1. **Hugging Face Hub HTTP API**, fully documented: model search (`GET /api/models?author=mlx-community&search=…&sort=downloads`), model info + file listing (`/api/models/{repo_id}` with `siblings`), file download URLs (`https://huggingface.co/{repo}/resolve/main/{file}`), LFS redirects, `HEAD` for file sizes, rate limits, optional user HF token for gated models (Llama, Gemma) via `Authorization: Bearer`.
90 +2. **The `mlx-community` organization** — the primary source of ready-to-run MLX models. Document the naming scheme (`Model-Name-4bit`, `-8bit`, `-bf16`) and what the quantization suffixes mean for quality/RAM.
91 +3. **A curated "Featured" catalog of 20–30 excellent models** across sizes, verified to exist RIGHT NOW on the Hub with exact repo IDs and download sizes. Cover: tiny (≤3B: Qwen small, Llama 3.2 small, SmolLM, Gemma small), mid (7–14B: Qwen, Mistral/Ministral, Llama, Phi, Gemma), large (27–70B+ for 64GB+ Macs), plus coding models and reasoning models (DeepSeek-R1 distills, QwQ-class). For each: repo ID, params, quant, disk size, min recommended RAM, category tags, one-line description.
92 +4. **RAM recommendation table**: which model sizes fit comfortably on 8 / 16 / 24 / 32 / 48 / 64 / 128 GB Macs — this feeds the in-app compatibility badges.
93 +
94 +---
95 +
96 +## PHASE 1 — Project Setup
97 +
98 +- **Toolchain:** Swift Package Manager. `Package.swift` with executable target `ZyquoLocal`, depending on the MLX packages identified in Phase 0. Build with `swift build -c release` per the recipe in `docs/BUILD.md`.
99 +- **App bundle:** `Makefile` that: (1) builds release, (2) assembles `Zyquo Local.app` (`Contents/MacOS/ZyquoLocal`, `Info.plist`, `Resources/AppIcon.icns`, and any Metal library/bundle resources MLX requires — verify resource bundling for SPM-built apps), (3) signs (Phase 8; ad-hoc for `make dev`).
100 +- **Info.plist:** `CFBundleDisplayName` = `Zyquo Local`, bundle ID `com.zyquo.local`, `LSMinimumSystemVersion` per MLX requirements (macOS 14.0 if MLX requires it — set from Phase 0 findings), `NSHighResolutionCapable`, `LSApplicationCategoryType` (`public.app-category.productivity`), `LSArchitecturePriority` arm64.
101 +- **Entry point:** `@main` SwiftUI `App`; proper activation when launched from terminal.
102 +- **Dependencies:** only MLX packages + (if chosen) swift-transformers + optionally Apple's `swift-markdown`. Nothing else.
103 +
104 +---
105 +
106 +## PHASE 2 — Architecture + Inference Proof-of-Concept
107 +
108 +```
109 +Sources/ZyquoLocal/
110 +├── App/ # @main, windows, menu bar extra, Apple Silicon gate
111 +├── DesignSystem/ # ZyquoTheme — same token system as Zyquo Cloud, Local palette
112 +├── Models/ # Conversation, Message, LocalModel, DownloadTask, Persona…
113 +├── Engine/
114 +│ ├── InferenceEngine.swift # actor: load/unload, warmup, streaming generate
115 +│ ├── GenerationParams.swift # temp, topP, repetitionPenalty, maxTokens, seed
116 +│ ├── ChatSession.swift # multi-turn history → prompt via chat template, KV cache reuse
117 +│ └── MemoryAdvisor.swift # RAM estimation, fits/tight/won't-fit verdicts
118 +├── Hub/
119 +│ ├── HubService.swift # HF search, model info, file listing
120 +│ ├── DownloadManager.swift # queued, resumable, per-file progress, checksum
121 +│ └── ModelStore.swift # on-disk library: scan, validate, size, delete
122 +├── Services/
123 +│ ├── PersistenceService.swift # conversations as JSON in Application Support
124 +│ └── ModelCatalog.swift # curated Featured catalog from docs/MODELS.md
125 +├── ViewModels/
126 +└── Views/
127 +```
128 +
129 +- **`InferenceEngine` is an actor.** One model loaded at a time (v1). Loading states: `unloaded → loading(progress) → ready → generating`. Generation exposed as `AsyncThrowingStream<GenerationEvent>` where events include `.token(String)`, `.stats(tokensPerSec, ...)`, `.finished(reason)`. Cancellation must actually stop the generation loop.
130 +- **`ChatSession`** builds the prompt from conversation history using the model's own chat template (from `tokenizer_config.json`), truncates oldest turns when exceeding the context window (keep the system prompt), and reuses KV cache across turns when the API allows.
131 +- **Stats are first-class:** measure time-to-first-token, tokens/sec, generated token count, and peak memory for every response.
132 +- **PHASE GATE:** before any UI, ship a tiny CLI mode (`ZyquoLocal --poc <model-dir> "<prompt>"`) that loads a small model (e.g., a ≤1B mlx-community model) and streams tokens to stdout with final stats. This proves the whole stack.
133 +
134 +---
135 +
136 +## PHASE 3 — HUGGING FACE INTEGRATION: BROWSE & DOWNLOAD IN-APP (THE HEART OF THE APP)
137 +
138 +This must be flawless — it is the feature that defines Zyquo Local.
139 +
140 +**HubService:**
141 +- Search the Hub live (default scope: `mlx-community`, toggle "All of Hugging Face" with an MLX-compatibility filter based on `config.json` architectures from Phase 0.A.4)
142 +- Sort by downloads / likes / recency; filter by size class and quantization
143 +- Fetch full file listings with per-file sizes; compute total download size before starting
144 +- Optional **HF token** field in Settings for gated models (Llama, Gemma) — stored in the app's config, masked in UI, sent only to huggingface.co
145 +
146 +**DownloadManager:**
147 +- Download all required files of a repo (per Phase 0.A.3's required-file list) into `Models/{org}/{repo}/`
148 +- **Resumable** (HTTP Range) across app restarts; **pause / resume / cancel** per model
149 +- Live progress: per-file and overall bytes, speed (MB/s), ETA; `URLSession` background-friendly configuration, 2 concurrent file downloads max
150 +- Integrity: verify final file sizes against the Hub listing (and checksums where available); atomic completion (download to `.partial`, move into place, mark model valid only when all files verified)
151 +- Disk space pre-check before starting; clear error if insufficient
152 +
153 +**ModelStore:**
154 +- Scans the Models folder at launch; validates each model directory (required files present); reports size on disk
155 +- Delete model (with confirmation showing reclaimed space); reveal in Finder
156 +- Tracks last-used date and per-model default generation params
157 +
158 +---
159 +
160 +## PHASE 4 — DESIGN SYSTEM & UI SPECIFICATION (LIGHT THEME, PIXEL-PERFECT)
161 +
162 +Zyquo Local shares the **same design DNA and token system as Zyquo Cloud** (`ZyquoTheme`), with a distinct **"Local" identity**: where Cloud is sky-blue and airy, Local is **grounded, warm, on-device** — an emerald-graphite story evoking silicon and privacy.
163 +
164 +### 4.1 — Light theme specification
165 +
166 +| Token | Value (light) | Usage |
167 +|---|---|---|
168 +| `background` | `#FAFBFA` (warm neutral off-white, faint green undertone) | Main canvas |
169 +| `surface` | `#FFFFFF` | Cards, assistant bubbles, input bar |
170 +| `surfaceSecondary` | `#F2F5F3` | Hover, code blocks |
171 +| `sidebar` | `NSVisualEffectView` `.sidebar` material | Sidebar |
172 +| `accent` | `#0E9F6E` (refined emerald — "on-device" green) | Primary actions, selection, links, send |
173 +| `accentSubtle` | `#E7F6F0` | Selected rows, user bubble tint |
174 +| `textPrimary` | `#1A1E1C` | Body text |
175 +| `textSecondary` | `#6B7472` | Metadata |
176 +| `textTertiary` | `#9EA8A5` | Placeholders |
177 +| `border` | `#E4E9E6` | 0.5pt hairlines |
178 +| `success` / `warning` / `danger` | `#2FA36B` / `#D9822B` / `#D64545` | Status |
179 +
180 +Same global rules as Zyquo Cloud: no pure black on pure white, 0.5pt hairlines, ultra-soft shadows (`black.opacity(0.06)`, radius 12, y 2) only on floating panels, dark theme derived from tokens (deep graphite-green, not flat gray), light theme is the flagship.
181 +
182 +**Typography** (identical scale to Zyquo Cloud): `title` 20pt semibold; `body` 13.5pt regular, line-height 1.45; `bodyEmphasis` 13.5pt medium; `caption` 11pt; `code` 12.5pt SF Mono; user-adjustable chat size 12–18pt. **Spacing** 4/8/12/16/20/24/32; **radii** 6/10/14; message column max 760pt centered.
183 +
184 +### 4.2 — Layout & screens (exact spec)
185 +
186 +**Main window**`NavigationSplitView`, min 980×640, default 1240×800:
187 +
188 +- **Sidebar (260pt, translucent):** "Zyquo Local" wordmark (icon glyph + name); two top-level sections: **Chats** (search, New Chat button, conversations grouped Pinned/Today/Yesterday/Previous 7 Days/Older, each row = title + model badge + relative time, hover pin/delete) and **Library** (entry to the Model Library screen, with a live badge showing active downloads count + a mini overall progress ring). Footer: settings gear + currently loaded model chip with a colored RAM dot.
189 +- **Chat area:**
190 + - *Header (52pt):* editable conversation title; centered **model chip** (model name + quant badge, e.g. "Qwen2.5-7B · 4bit"; click → model switcher popover listing downloaded models with RAM verdict badges; switching triggers unload/load with inline progress in the chip); right: performance toggle (shows live tokens/sec while generating), export, info popover (system prompt, params, context usage bar showing tokens used / context window).
191 + - *Transcript:* identical structure to Zyquo Cloud (user right in `accentSubtle` bubbles, assistant left on `surface`, 16pt rhythm, hover timestamps, jump-to-bottom pill, full Markdown + syntax-highlighted code blocks with copy, collapsible "Thinking…" section for reasoning models like R1-distills — parse `<think>` tags). Under each assistant message, a subtle `caption` stats line: `⚡ 42.3 tok/s · 512 tokens · 1.2s to first token`.
192 + - *Model-loading state:* when a model is loading, the chat area shows an elegant centered loading card (model name, animated progress, RAM being allocated) — input disabled with a clear hint.
193 + - *Input bar:* identical floating card as Zyquo Cloud (radius 14, soft shadow); attach button for **text files** (txt, md, code, csv, json → injected into the message); parameters quick-toggle; circular emerald send button (⌘↩); stop button during generation.
194 +- **Empty state (no model downloaded yet):** a beautiful onboarding hero — icon, "Download your first model", 3–4 recommended starter models as cards (name, size, RAM fit for THIS Mac, one-line description, Download button). This is many users' first screen: it must be stunning.
195 +
196 +**Model Library screen** (pushed in the detail column, or ⌘L):
197 +- **Two tabs: "Installed" and "Discover".**
198 +- *Installed:* grid/list of downloaded models — name, quant badge, params, disk size, last used, RAM verdict badge (green "Fits" / orange "Tight" / red "Too large" for this machine via `MemoryAdvisor`), actions: Load, chat shortcut, per-model default params, Reveal in Finder, Delete (confirmation with reclaimed GB). Header shows total disk used by models.
199 +- *Discover:* search field (live Hub search), scope segmented control (Featured / mlx-community / All MLX-compatible), filters (size class, quantization), sort (downloads/likes/newest). Result cards: model name, org, params + quant, **download size**, downloads count, RAM verdict for this Mac, short description; primary **Download** button → card flips into live progress state (progress bar, MB/s, ETA, pause/cancel). The **Featured** tab renders the curated catalog from `docs/MODELS.md` with editorial one-liners — it must feel hand-picked, like an App Store front page.
200 +- *Downloads drawer:* a bottom bar (or popover from the sidebar badge) listing all active/queued downloads with individual controls.
201 +
202 +**Settings** (native tabs, 720×520): 1. **General** (default model on launch, keep model loaded in background toggle) 2. **Models & Storage** (models folder location + change, total usage, HF token field masked, auto-verify downloads) 3. **Inference** (default generation params with explanations, GPU cache limit, context length cap) 4. **Appearance** (Light/Dark/System, accent choices: emerald default + graphite, sky, amber, rose; font size slider with live preview) 5. **Shortcuts** 6. **Advanced** (reveal data folder, export/import conversations).
203 +
204 +**Quick Chat panel** (⌥Space): same Spotlight-style floating panel as Zyquo Cloud, using the currently loaded model; if none is loaded, offers one-click load of the last-used model.
205 +
206 +**Compare mode:** 2 columns (v1: two models — note both must fit in RAM together; `MemoryAdvisor` gates this), same prompt broadcast, independent streaming and stats — a spectacular way to visually compare local models.
207 +
208 +### 4.3 — Motion & micro-interactions
209 +Same standard as Zyquo Cloud: smooth streaming (no jitter), blinking caret at stream tail, 150ms fade+rise on sent messages, 80ms hover eases, 0.97 press scale, `.snappy` popovers, 60fps always (lazy transcript rendering). Plus Local-specific: download progress animates fluidly (no jumpy bars), the model chip morphs smoothly between unloaded/loading/ready states, and tokens/sec ticker updates at 4Hz max (no flicker).
210 +
211 +### 4.4 — Design quality gate
212 +Before declaring the project done, review every screen: consistent token usage, aligned baselines, no clipped text, correct dark mode, clean font-size scaling, ALL states designed (no models yet, model loading, downloading, download failed, out of RAM, generating, error). If a screen looks "developer-made" rather than "designed", iterate until it doesn't.
213 +
214 +---
215 +
216 +## PHASE 5 — APP ICON: ULTRA-LEGENDARY "LOCAL" ICON, DESIGNED IN SVG
217 +
218 +Designed in SVG first (`assets/icon/zyquo-local.svg`), then converted to `.icns`. It must be the visual sibling of the Zyquo Cloud icon — same squircle, same Z-monogram DNA, same premium quality — but telling the **on-device** story instead of the sky.
219 +
220 +**Creative direction:**
221 +- **Concept — the Z on silicon.** Two directions to explore (render both, keep the best):
222 + 1. *Z-chip:* the bold Z monogram seated at the center of a subtly stylized **silicon die** — a minimal square chip outline with fine traces/pins radiating from its edges, the Z glowing like an active core. Reads as "the intelligence lives on THIS chip".
223 + 2. *Z-core:* the Z itself drawn as a luminous circuit path — its strokes are clean conductor traces with rounded corners and 2–3 tiny node dots at the bends, glowing emerald on a deep graphite field. One shape, letterform + circuitry.
224 +- **Canvas:** macOS Big Sur–style rounded **squircle** (Apple curvature).
225 +- **Palette (mirrors the app):** deep graphite-to-near-black vertical gradient background with a faint green undertone (`#1E2622 → #0F1412` territory), the Z/traces in **luminous emerald** (`#17C787 → #0E9F6E` gradient) with a restrained outer glow, plus near-white micro-highlights on trace nodes. Where Cloud is day-sky and airy, Local is dark-silicon and glowing — side by side in the Dock, the pair must be instantly recognizable as siblings.
226 +- **Precision & iteration:** clean paths, `viewBox="0 0 1024 1024"`, optical centering, glow effects that don't turn to mud when downscaled. Render at 1024/512/256/128/64/32/16, LOOK at each, refine; bake a simplified variant (drop traces, keep glowing Z) for 16/32px if needed.
227 +
228 +**Pipeline (Makefile):** `zyquo-local.svg` → PNGs (16→1024, incl. `@2x`) via `rsvg-convert` or a small CoreGraphics rasterizer → `AppIcon.iconset``iconutil -c icns`. SVG stays in the repo as source of truth. Derive the monochrome **menu bar template icon** (18×18pt, `isTemplate = true`) and the in-app empty-state/wordmark glyph from the same SVG.
229 +
230 +---
231 +
232 +## PHASE 6 — Features (This is where Zyquo Local becomes LEGENDARY)
233 +
234 +### Core chat
235 +- Multi-conversation sidebar: search, pin, rename, delete, folders/tags
236 +- Full **token-streaming** with stop button; per-message stats line (tok/s, tokens, TTFT)
237 +- **Model switcher** per conversation (among downloaded models), with load progress and RAM verdicts; conversation remembers its model
238 +- Message actions: copy, edit & resend, regenerate (optionally with another model), delete, quote-reply
239 +- **Reasoning display** for thinking models (R1 distills, QwQ-class): parse `<think>…</think>` into the collapsible section
240 +- System prompt per conversation + global default; context usage bar (tokens used vs context window) with automatic oldest-turn truncation
241 +- Per-conversation generation params: temperature, top_p, repetition penalty, max tokens, seed — with sensible defaults and inline explanations
242 +
243 +### Model management (the differentiator)
244 +- Full in-app Hub browse/search/download (Phase 3) with the Featured curated catalog
245 +- `MemoryAdvisor` verdicts everywhere a model is shown (Fits / Tight / Too large **for this specific Mac**, based on physical RAM detected via `sysctl hw.memsize`)
246 +- One model loaded at a time; explicit Load/Unload; optional "keep loaded" setting; unload frees memory verifiably
247 +- Live memory readout while a model is loaded (app footprint) in the footer chip
248 +
249 +### Attachments & productivity
250 +- Drag & drop / attach **text files** (txt, md, code, csv, json) into messages
251 +- **Prompt Library**: ship ≥50 quality templates + user templates with `{{input}}` variables
252 +- **Personas**: system prompt + preferred model + params
253 +- **Quick Chat** panel (⌥Space)
254 +- **Compare mode** (2 local models side-by-side, RAM-gated)
255 +- **Export** conversation → Markdown and PDF; full-text search across conversations
256 +- Auto-generated conversation titles using the loaded model itself (short, cheap prompt after first exchange)
257 +
258 +### macOS-native polish
259 +- Shortcuts: ⌘N new chat, ⌘K model switcher/command palette, ⌘L Library, ⌘F search, ⌘↩ send, ⌘⇧E export, ⌥Space Quick Chat
260 +- **Menu bar extra** (toggleable): quick chat + download progress at a glance
261 +- Launch fast; the app itself must stay light — all heaviness lives in explicit model loading
262 +
263 +---
264 +
265 +## PHASE 7 — LOCAL MODEL VERIFICATION (MANDATORY)
266 +
267 +No API keys here — verification means proving that **downloading and running real models works end-to-end on this machine**:
268 +
269 +1. Build a verification harness (`zyquo-verify` CLI target or script) that:
270 + - Downloads at least **5 models from the Featured catalog spanning architectures and sizes** (e.g., one ≤1B, one ~3B, one ~7–8B, one reasoning distill, one coding model — pick what fits this Mac's RAM)
271 + - For each: validates the downloaded files, loads the model, runs a deterministic generation (`temperature 0`, prompt `"Reply with exactly: OK"`), verifies non-empty coherent output, runs a **multi-turn** exchange (context carry-over works), runs a **streaming cancellation** test, unloads, and confirms memory is released
272 + - Records tok/s and TTFT per model
273 +2. Additionally, dry-verify the ENTIRE Featured catalog against the live Hub: every repo ID exists, files are listed, sizes match `docs/MODELS.md`.
274 +3. Produce a results table: model → download ✅/❌ → load ✅/❌ → generate ✅/❌ → tok/s. **Fix every failure** (wrong repo ID, unsupported arch, template issue) and iterate until green. Remove or replace catalog entries that are genuinely broken.
275 +4. Downloaded test models may be deleted afterward to reclaim disk, keeping the smallest one for ongoing dev.
276 +
277 +---
278 +
279 +## PHASE 8 — SIGNING & NOTARIZATION (REAL, NOT AD-HOC)
280 +
281 +The user has an existing, working signing/notarization setup for another project. **Before doing anything, read and inspect the folder:**
282 +
283 +```
284 +/Users/simon-pierreboucher/Desktop/other/OTHER/zyquo-term
285 +```
286 +
287 +Locate everything related to signing and notarization there: scripts, Makefile targets, the **Developer ID Application identity name**, **Team ID**, **notarytool keychain profile or Apple ID + app-specific password**, entitlements files, any `.env`/config. **Reuse the exact same identity, Team ID, and notarytool credentials/profile for Zyquo Local.** Never invent placeholders, never print secrets, never commit them.
288 +
289 +Then implement `make release`:
290 +
291 +1. Build release binary (arm64 only — MLX), assemble `Zyquo Local.app` including any MLX resource bundles/metallibs
292 +2. `entitlements.plist` with **Hardened Runtime**; only what's needed (network client for Hugging Face downloads; add JIT/unsigned-memory entitlements ONLY if MLX demonstrably requires them — test without first)
293 +3. `codesign --force --options runtime --timestamp --entitlements entitlements.plist --sign "Developer ID Application: <identity from zyquo-term>" "Zyquo Local.app"` — sign nested frameworks/dylibs/metallibs first, then the app
294 +4. `ditto -c -k --keepParent``xcrun notarytool submit "Zyquo Local.zip" --keychain-profile "<profile from zyquo-term>" --wait`
295 +5. `xcrun stapler staple "Zyquo Local.app"`; verify `spctl -a -vv` says "accepted, source=Notarized Developer ID" and `stapler validate` passes
296 +6. Optional signed+stapled DMG via `hdiutil`
297 +7. On failure: `notarytool log`, fix (nested code signatures are the classic culprit with bundled Metal libraries), resubmit until it passes.
298 +
299 +Keep `make dev` with ad-hoc signing for iteration.
300 +
301 +---
302 +
303 +## Engineering Standards
304 +
305 +- Swift 5.9+ (Swift 6 mode if the toolchain allows); zero compiler warnings
306 +- `InferenceEngine` as an actor; all Hub/network types `Codable` structs — no dictionary spelunking
307 +- Robust error surfaces: human-readable messages for every failure class (no disk space, download interrupted → auto-resume offer, gated model → HF token hint, unsupported architecture, out-of-memory during load → suggest smaller quant)
308 +- Cancellation everywhere: downloads, model loading, generation
309 +- All UI strings centralized; design tokens only — no hardcoded colors/sizes in views
310 +- `README.md` + `docs/BUILD.md` with the exact no-Xcode(-IDE) build recipe
311 +- Commit in logical increments with clear messages; never commit model weights
312 +
313 +## Definition of Done
314 +
315 +- `make release` produces a **Developer ID–signed, notarized, stapled** `Zyquo Local.app` (verified by `spctl`)
316 +- In-app Hugging Face browsing + one-click resumable downloads work flawlessly, with the Featured curated catalog live-verified
317 +- Chat with downloaded MLX models works: streaming, multi-turn with context management, stop, stats (tok/s, TTFT), reasoning display
318 +- `MemoryAdvisor` verdicts are accurate for this machine; loading/unloading verifiably frees memory
319 +- The silicon-themed SVG icon exists, is striking at all sizes, embedded as `.icns` + menu bar template icon; clearly the sibling of the Zyquo Cloud icon
320 +- The emerald light theme matches the Phase 4 spec exactly and passes the design quality gate; dark theme derived and correct
321 +- Naming coherent everywhere: `Zyquo Local` user-facing, `com.zyquo.local`, `ZyquoLocal` target/data folder
322 +- Phase 7 verification table fully green; `docs/MODELS.md` and `ModelCatalog` in perfect sync
323 +- **Every code file starts with the mandatory Author/Mail header** (verified by a repository-wide sweep)
324 +- `docs/PLAN.md` shows every phase completed with its checkpoint summary
325 +- Zyquo Local feels like a polished, legendary native Mac app — the best local-LLM experience on macOS
added docs/BUILD.md +63 −0
@@ -0,0 +1,63 @@
1 +<!--
2 + BUILD.md
3 + Zyquo Local
4 +
5 + Author: Simon-Pierre Boucher
6 + Mail: contact@spboucher.ai
7 +-->
8 +
9 +# Building Zyquo Local Without the Xcode IDE
10 +
11 +Zyquo Local is built entirely from the command line with SwiftPM. The Xcode IDE is never used and no `.xcodeproj` exists. Two toolchain pieces are required beyond a plain macOS install:
12 +
13 +1. **Command Line Tools** (`xcode-select --install`) — provides Swift 6.4 and the macOS SDKs.
14 +2. **Apple Metal Toolchain** — MLX compiles its GPU kernels (`.metal` files) at build time, and the `metal` compiler is **not** included in the Command Line Tools.
15 +
16 +## The Metal toolchain finding (Phase 0, resolved empirically 2026-07-30)
17 +
18 +- `swift build` on a CLT-only machine fails inside `Cmlx` with
19 + `error: unable to spawn process 'metal' (No such file or directory)`
20 + mlx-swift 0.31.6 declares its Metal sources as SPM resources and SwiftPM's
21 + `CompileMetalFile` rule shells out to `metal`.
22 +- SwiftPM resolves `metal` **via `PATH`** (verified with a logging stub), so a
23 + full Xcode install is NOT required. A standalone `Metal.xctoolchain` on
24 + `PATH` is sufficient.
25 +- This machine uses the Metal Toolchain component copied from cluster node
26 + M4M36 (`com.apple.MobileAsset.MetalToolchain-v17.5.188.0`, Apple metal
27 + version 32023.883) into:
28 +
29 + ```
30 + ~/Developer/Metal.xctoolchain
31 + ```
32 +
33 + On a machine with Xcode 26+, the same component is installed by
34 + `xcodebuild -downloadComponent MetalToolchain`.
35 +
36 +## SDK pin for SwiftUI on macOS 27 CLT
37 +
38 +The macOS 27 SDKs shipped with the CLT expose SwiftUI's `@State` / `@Bindable`
39 +(and related Observation wrappers) as Xcode-only macro plugins, which a plain
40 +`swift build` cannot expand. As with Zyquo Term (see its ADR-0001), builds pin:
41 +
42 +```sh
43 +export SDKROOT=/Library/Developer/CommandLineTools/SDKs/MacOSX26.5.sdk
44 +```
45 +
46 +The Makefile and all scripts set both `SDKROOT` and the Metal `PATH`
47 +automatically.
48 +
49 +## Canonical build recipe
50 +
51 +```sh
52 +# everything below is wrapped by `make build` / `make dev` / `make release`
53 +export SDKROOT=/Library/Developer/CommandLineTools/SDKs/MacOSX26.5.sdk
54 +export PATH="$HOME/Developer/Metal.xctoolchain/usr/bin:$PATH"
55 +
56 +swift build -c release # release binary
57 +make dev # debug app bundle, ad-hoc signed, and launch
58 +make release # signed + notarized + stapled Zyquo Local.app
59 +```
60 +
61 +Verified: a scratch package depending on mlx-swift 0.31.6 builds in seconds
62 +with this recipe and executes on the GPU (`Device(gpu, 0)`), confirming the
63 +runtime `mlx.metallib` resource is produced and loaded correctly.
added docs/PLAN.md +35 −0
@@ -0,0 +1,35 @@
1 +<!--
2 + PLAN.md
3 + Zyquo Local
4 +
5 + Author: Simon-Pierre Boucher
6 + Mail: contact@spboucher.ai
7 +-->
8 +
9 +# Zyquo Local — Build Plan
10 +
11 +Machine: Apple M5 Max, 18 cores, 48 GB RAM, macOS 27.0 (beta), Swift 6.4 (Command Line Tools only, no Xcode IDE installed).
12 +
13 +## Phase 0 — Intensive research (in progress)
14 +
15 +- [ ] Research current mlx-swift packages, versions, SPM coordinates, min macOS
16 +- [ ] Research current MLXLLM/MLXLMCommon API surface (load, tokenize, chat template, streaming, params, KV cache, unload)
17 +- [ ] Document model directory format (required files)
18 +- [ ] Document supported architectures list
19 +- [ ] Resolve the no-Xcode build question EMPIRICALLY (scratch SPM package + `swift build` with CLT only)
20 +- [ ] Decide swift-transformers vs direct HTTP HubService
21 +- [ ] Document memory model + GPU cache limits + RAM estimation
22 +- [ ] Write `docs/MLX-RESEARCH.md`
23 +- [ ] Research HF Hub HTTP API (search, model info, resolve URLs, LFS, tokens)
24 +- [ ] Curate 20–30 Featured models, live-verified on the Hub with sizes
25 +- [ ] RAM recommendation table
26 +- [ ] Write `docs/MODELS.md`
27 +
28 +## Phase 1 — Project setup — pending
29 +## Phase 2 — Architecture + inference PoC — pending
30 +## Phase 3 — Hub browse & download — pending
31 +## Phase 4 — Design system & UI spec — pending
32 +## Phase 5 — App icon — pending
33 +## Phase 6 — Features / full UI — pending
34 +## Phase 7 — Local model verification — pending
35 +## Phase 8 — Signing & notarization — pending
36