From df52514de88a97e920ca74a794d97cfbcdd95fb3 Mon Sep 17 00:00:00 2001 From: Davide Angelocola Date: Sat, 3 Oct 2026 18:02:07 +0200 Subject: [PATCH] local: LocalClefTypeSafeClient.loadOnGpu; no GPU option for Laya Clef-flash's 4-bit graph on WebGPU gives the same answers as the CPU on every launch and runs about 3.5x faster, so it gets a public loadOnGpu (macOS on Apple Silicon: the WebGPU backend ships only in ONNX Runtime's macOS native library). Laya stays CPU-only in the public API: on WebGPU its choice probabilities drifted between process launches (0.49-0.65 for an input the CPU answers 0.65), the failure mode that removed int8 Laya. Timings live in the how-to, not the javadoc. Co-Authored-By: Claude Sonnet 5 --- CHANGELOG.md | 1 + CLAUDE.md | 5 +++-- docs/explanation.md | 6 ++++++ docs/how-to.md | 3 ++- docs/reference.md | 1 + .../dfa1/typesafe/local/LocalClefTypeSafeClient.java | 11 +++++++++++ 6 files changed, 24 insertions(+), 3 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index f4d88cc..75fc994 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -7,6 +7,7 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0 ## [Unreleased] +- **`local`: `LocalClefTypeSafeClient.loadOnGpu(...)`** — Clef-flash on ONNX Runtime's WebGPU backend (macOS on Apple Silicon), same answers as the CPU; Laya gets no GPU option, since its WebGPU answers drift between launches. (#19) - **docs: run Clef-flash on a Mac with MLX** — a how-to pointing the regular client at mlx-community's local Clef-flash server: about 0.55 s per request on an M5, much closer to Jev than Laya. (#14) - **`local`: `LocalClefTypeSafeClient`** — Cloudflare's Clef-flash (Qwen3.5-9B + joint schema head) on Ollaya's ONNX graph over the upstream bf16 weights; one forward pass per request; `scripts/clef/quantize_q4.py` converts it to 4-bit weights (7.7 GB, 2–4 s per request on an M5 GPU). (#14) - **`local`: `LocalLayaTypeSafeClient`, `LocalQwenTypeSafeClient`** — new module, one client class per model, evaluating in-process on ONNX Runtime from a local model directory (Laya fp32 from onnx-community, or Qwen2.5 as a prompted-LLM baseline), with a `Local engines` workflow publishing agreement-with-Jev and throughput tables. (#14) diff --git a/CLAUDE.md b/CLAUDE.md index bc4693d..24ca92c 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -96,8 +96,9 @@ local — one public TypeSafeClient per model (io.github.dfa1.typesafe.local src/test/resources). Tests needing model files are @Tag("model"), excluded by the module's own excludedGroups (acceptance,model); opt in with -DexcludedGroups=acceptance -Dengine=laya. JevComparison (test scope, main) replays 104 requests against cached real-Jev answers - (src/test/resources/jev); LocalTypeSafeClientBenchmark is JMH. Package-private WebGPU switch - (LayaEngine/QwenEngine.load(dir, gpu)) is experimental. LayaEngine reads onnx-community's layout + (src/test/resources/jev); LocalTypeSafeClientBenchmark is JMH. WebGPU (Metal, macOS-only + native lib): public only as LocalClefTypeSafeClient.loadOnGpu (stable, same answers as CPU); Laya/Qwen keep + the package-private load(dir, gpu) for the harness — Laya's WebGPU answers drift between launches. LayaEngine reads onnx-community's layout (onnx/model.onnx, config.json "laya", bool marker_mask). The `Local engines` workflow runs tests, comparison and JMH on Linux/macOS and writes tables to the job summary. diff --git a/docs/explanation.md b/docs/explanation.md index f4523de..2293d02 100644 --- a/docs/explanation.md +++ b/docs/explanation.md @@ -43,6 +43,12 @@ of 32), the same format onnx-community uses for its q4 models. That brought it t GPU, with probabilities within about 0.03 of the fp32 graph. The rest of the cost is the model's size: every token passes through 9B weights. +Only Clef-flash has a GPU option (`loadOnGpu`, WebGPU on Metal, macOS only). Its 4-bit graph gives the same answers +on the GPU as on the CPU, run after run, and runs about 3.5× faster there. Laya doesn't get one: on WebGPU its +choice probabilities changed from one process launch to the next (0.65 on the CPU, anywhere from 0.49 to 0.65 on the +GPU for the same input) — the same kind of silent, hardware-dependent drift that removed int8 Laya — and it saved +only tens of milliseconds anyway. + Several things were measured and rejected: - **Parallel sessions.** On a CPU, one ONNX session already uses every core, so splitting the same work across diff --git a/docs/how-to.md b/docs/how-to.md index 0e70ce7..81e21bb 100644 --- a/docs/how-to.md +++ b/docs/how-to.md @@ -685,7 +685,8 @@ cd local && uv run scripts/clef/quantize_q4.py # writes ~/.cache/typesafe-loca ``` `LocalClefTypeSafeClient.load(Path.of(..., "clef-flash-q4"))` then needs 7.7 GB and, on the same M5, answers in 7–13 s -for 1–3 questions on the CPU, or 2–4 s on the GPU (WebGPU, still experimental here): batch work, not interactive use. +for 1–3 questions on the CPU, or 2–4 s on the GPU with `LocalClefTypeSafeClient.loadOnGpu(...)` (macOS on Apple +Silicon only; same answers as the CPU): batch work, not interactive use. On a Mac, the same model runs about 5× faster through MLX; see [Run Clef-flash on a Mac with MLX](#run-clef-flash-on-a-mac-with-mlx). diff --git a/docs/reference.md b/docs/reference.md index 107fccf..1194a3e 100644 --- a/docs/reference.md +++ b/docs/reference.md @@ -330,6 +330,7 @@ public final class LocalQwenTypeSafeClient implements TypeSafeClient { } public final class LocalClefTypeSafeClient implements TypeSafeClient { public static LocalClefTypeSafeClient load(Path dir); // *.safetensors + tokenizer.json (Cloudflare) + flash/model-*.onnx + public static LocalClefTypeSafeClient loadOnGpu(Path dir); // same, on WebGPU (macOS on Apple Silicon only) } ``` diff --git a/local/src/main/java/io/github/dfa1/typesafe/local/LocalClefTypeSafeClient.java b/local/src/main/java/io/github/dfa1/typesafe/local/LocalClefTypeSafeClient.java index 2a0162e..0670c2b 100644 --- a/local/src/main/java/io/github/dfa1/typesafe/local/LocalClefTypeSafeClient.java +++ b/local/src/main/java/io/github/dfa1/typesafe/local/LocalClefTypeSafeClient.java @@ -27,4 +27,15 @@ private LocalClefTypeSafeClient(Engine engine) { public static LocalClefTypeSafeClient load(Path dir) { return new LocalClefTypeSafeClient(ClefEngine.load(dir)); } + + /** + * {@link #load}, but evaluating on the GPU through ONNX Runtime's WebGPU backend (Metal), with the same answers as + * on the CPU. Worth it for the 4-bit graph; see the how-to for measurements. + * + *

macOS on Apple Silicon only: the WebGPU backend ships only in ONNX Runtime's macOS native library, so + * elsewhere this throws {@link IllegalStateException}. + */ + public static LocalClefTypeSafeClient loadOnGpu(Path dir) { + return new LocalClefTypeSafeClient(ClefEngine.load(dir, true)); + } }