SPB Git

spb/forge Public MIT

Forge — LLM training from scratch in pure C++20 + Metal on Apple Silicon.

C++ 61.2% C 23% Python 7.6% TeX 7.2% CMake 1.1%

History of src/train/trainer.cpp · clear filter

  1. Add elapsed_s column to log.csv for structured metrics consumers
    Forge Studio (the GUI companion) tails log.csv as its metrics channel;
    wall-clock elapsed enables its time-axis mode and honest ETA math.
    
    Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
    simon-pierre boucher committed 5 days ago (Aug 5, 2026) · 1 file changed +6 −3
  2. Add wave 3: DeepSeek-V3-style MoE routing, all config-selected
    - moe_scoring "softmax"|"sigmoid" (new sigmoid op, CPU+Metal+backward)
    - aux-loss-free balancing (V3 "noaux"): top-k selection ranks score+bias
      while gate values stay biasless; per-expert load counted on-GPU each
      forward and the balance bias nudged ±moe_bias_gamma after every
      optimizer step; bias is checkpointed as a grad-free parameter
    - moe_norm_topk (renormalize kept gates or keep raw sigmoid scores),
      routed_scaling_factor (V3: 2.5), moe_d_ff (per-expert width),
      first_k_dense (dense MLPs for the first k layers)
    - parity tests: biased/no-renorm topk, sigmoid, expert_counts, and a
      full V3-style model (sigmoid/noaux/scaled/first-dense) — CPU==GPU
    
    Remaining wave-3 items (MLA, MuonClip QK-clip, MTP) documented in
    ARCHITECTURES.md.
    
    Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
    simon-pierre boucher committed 5 days ago (Aug 5, 2026) · 1 file changed +2
  3. Add configurable training modes: Muon optimizer and WSD schedule
    - optimizer "adamw" | "muon": Newton-Schulz orthogonalized momentum on 2-D
      hidden matrices (composed from the existing matmul kernels on Metal),
      AdamW kept for embeddings/head/1-D params; muon_lr follows the lr schedule
    - schedule "cosine" | "wsd": warmup-stable-decay with 1-sqrt cooldown,
      extendable runs, wsd_decay_frac
    - gpt-50m base config + Muon+WSD and MobileLLM-style deep-and-thin variants
    
    Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
    simon-pierre boucher committed 5 days ago (Aug 5, 2026) · 1 file changed +33 −5
  4. Forge: LLM training from scratch in C++20 + Metal on Apple Silicon
    A complete transformer training stack with no ML dependencies: tensors,
    autograd, hand-written Metal kernels, flash attention (forward and backward),
    AdamW, BPE tokenizer, checkpointing and generation. Architecture is fully
    config-driven — the same binary trains 12M to 205M parameter models.
    
    Every Metal kernel is validated against a CPU reference (85 parity checks,
    <=1e-4, most bit-exact), gradients against central finite differences, and
    each optimization was accepted only after the training loss trajectory stayed
    numerically unchanged.
    
    Measured findings (M5 Max, documented in RESEARCH.md and paper/forge.tex):
    
    - `constant constexpr` for MSL tile constants declares an address-space
      variable, not a compile-time constant. Loops stop unrolling and every
      matrix accumulator spills: 0.82 -> 10.21 TFLOPS once switched to enums.
    - That defect is invisible in the AIR at every -O level, because unrolling
      happens in the driver back end. Benchmark; do not read the IR.
    - Register pressure, not bandwidth, dominates attention backward. Guided by
      measured spill counts, three restructurings took it 107 -> 7.05 ms (15.2x).
    - On M5, mpp::tensor_ops::matmul2d reaches 51.5 TFLOPS with f16 operands vs
      10.6 for a tuned simdgroup_matrix kernel (4.9x), verified numerically.
      f16 on the simdgroup path alone is worth only +18-22%.
    - Concurrent dispatch for the optimizer sweep: +22% on the 100M config.
    
    Trained the 12.2M config for one epoch over 19.14M TinyStories tokens:
    loss 8.40 -> 2.99, validation 3.009, perplexity 20.27, ~38.2k tokens/sec.
    Simon-Pierre Boucher committed 10 days ago (Jul 31, 2026) · 1 file changed +157