Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
1 change: 1 addition & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -7,6 +7,7 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0

## [Unreleased]

- **`local`: `LocalClefTypeSafeClient.loadOnGpu(...)`** — Clef-flash on ONNX Runtime's WebGPU backend (macOS on Apple Silicon), same answers as the CPU; Laya gets no GPU option, since its WebGPU answers drift between launches. (#19)
- **docs: run Clef-flash on a Mac with MLX** — a how-to pointing the regular client at mlx-community's local Clef-flash server: about 0.55 s per request on an M5, much closer to Jev than Laya. (#14)
- **`local`: `LocalClefTypeSafeClient`** — Cloudflare's Clef-flash (Qwen3.5-9B + joint schema head) on Ollaya's ONNX graph over the upstream bf16 weights; one forward pass per request; `scripts/clef/quantize_q4.py` converts it to 4-bit weights (7.7 GB, 2–4 s per request on an M5 GPU). (#14)
- **`local`: `LocalLayaTypeSafeClient`, `LocalQwenTypeSafeClient`** — new module, one client class per model, evaluating in-process on ONNX Runtime from a local model directory (Laya fp32 from onnx-community, or Qwen2.5 as a prompted-LLM baseline), with a `Local engines` workflow publishing agreement-with-Jev and throughput tables. (#14)
Expand Down
5 changes: 3 additions & 2 deletions CLAUDE.md
Original file line number Diff line number Diff line change
Expand Up @@ -96,8 +96,9 @@ local — one public TypeSafeClient per model (io.github.dfa1.typesafe.local
src/test/resources). Tests needing model files are @Tag("model"), excluded by the module's
own excludedGroups (acceptance,model); opt in with -DexcludedGroups=acceptance -Dengine=laya.
JevComparison (test scope, main) replays 104 requests against cached real-Jev answers
(src/test/resources/jev); LocalTypeSafeClientBenchmark is JMH. Package-private WebGPU switch
(LayaEngine/QwenEngine.load(dir, gpu)) is experimental. LayaEngine reads onnx-community's layout
(src/test/resources/jev); LocalTypeSafeClientBenchmark is JMH. WebGPU (Metal, macOS-only
native lib): public only as LocalClefTypeSafeClient.loadOnGpu (stable, same answers as CPU); Laya/Qwen keep
the package-private load(dir, gpu) for the harness — Laya's WebGPU answers drift between launches. LayaEngine reads onnx-community's layout
(onnx/model.onnx, config.json "laya", bool marker_mask).
The `Local engines` workflow runs tests, comparison and JMH on Linux/macOS and writes tables to
the job summary.
Expand Down
6 changes: 6 additions & 0 deletions docs/explanation.md
Original file line number Diff line number Diff line change
Expand Up @@ -43,6 +43,12 @@ of 32), the same format onnx-community uses for its q4 models. That brought it t
GPU, with probabilities within about 0.03 of the fp32 graph. The rest of the cost is the model's size: every token
passes through 9B weights.

Only Clef-flash has a GPU option (`loadOnGpu`, WebGPU on Metal, macOS only). Its 4-bit graph gives the same answers
on the GPU as on the CPU, run after run, and runs about 3.5× faster there. Laya doesn't get one: on WebGPU its
choice probabilities changed from one process launch to the next (0.65 on the CPU, anywhere from 0.49 to 0.65 on the
GPU for the same input) — the same kind of silent, hardware-dependent drift that removed int8 Laya — and it saved
only tens of milliseconds anyway.

Several things were measured and rejected:

- **Parallel sessions.** On a CPU, one ONNX session already uses every core, so splitting the same work across
Expand Down
3 changes: 2 additions & 1 deletion docs/how-to.md
Original file line number Diff line number Diff line change
Expand Up @@ -685,7 +685,8 @@ cd local && uv run scripts/clef/quantize_q4.py # writes ~/.cache/typesafe-loca
```

`LocalClefTypeSafeClient.load(Path.of(..., "clef-flash-q4"))` then needs 7.7 GB and, on the same M5, answers in 7–13 s
for 1–3 questions on the CPU, or 2–4 s on the GPU (WebGPU, still experimental here): batch work, not interactive use.
for 1–3 questions on the CPU, or 2–4 s on the GPU with `LocalClefTypeSafeClient.loadOnGpu(...)` (macOS on Apple
Silicon only; same answers as the CPU): batch work, not interactive use.
On a Mac, the same model runs about 5× faster through MLX; see
[Run Clef-flash on a Mac with MLX](#run-clef-flash-on-a-mac-with-mlx).

Expand Down
1 change: 1 addition & 0 deletions docs/reference.md
Original file line number Diff line number Diff line change
Expand Up @@ -330,6 +330,7 @@ public final class LocalQwenTypeSafeClient implements TypeSafeClient {
}
public final class LocalClefTypeSafeClient implements TypeSafeClient {
public static LocalClefTypeSafeClient load(Path dir); // *.safetensors + tokenizer.json (Cloudflare) + flash/model-*.onnx
public static LocalClefTypeSafeClient loadOnGpu(Path dir); // same, on WebGPU (macOS on Apple Silicon only)
}
```

Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -27,4 +27,15 @@ private LocalClefTypeSafeClient(Engine engine) {
public static LocalClefTypeSafeClient load(Path dir) {
return new LocalClefTypeSafeClient(ClefEngine.load(dir));
}

/**
* {@link #load}, but evaluating on the GPU through ONNX Runtime's WebGPU backend (Metal), with the same answers as
* on the CPU. Worth it for the 4-bit graph; see the how-to for measurements.
*
* <p>macOS on Apple Silicon only: the WebGPU backend ships only in ONNX Runtime's macOS native library, so
* elsewhere this throws {@link IllegalStateException}.
*/
public static LocalClefTypeSafeClient loadOnGpu(Path dir) {
return new LocalClefTypeSafeClient(ClefEngine.load(dir, true));
}
}
Loading