Cross-platform evaluation, hardware profiling, and benchmarking rig for local LLMs on Ollama, Microsoft Foundry Local, ONNX Runtime GenAI and Prism.
Tailored for macOS Apple Silicon (M1/M2/M3/M4 Metal & Unified Memory) and Linux / WSL2 (NVIDIA GeForce RTX CUDA).
BenchRig is an automated benchmarking and profiling suite that measures real-world code generation precision, logical reasoning, prompt prefill / generation speeds, context saturation, and hardware telemetry across local language models served via Ollama (llama.cpp) and Microsoft Foundry Server Runtime (Foundry Local / ONNX Runtime GenAI).
Unlike generic perplexity benchmarks, this suite focuses on practical developer workloads:
- Executes code in isolated sandboxes and checks deterministic unit test assertions.
- Direct Cross-Runtime & Engine Comparison: Benchmarks
llama.cppagainst ONNX Runtime GenAI (via Microsoft Foundry Local, directly, or through a Prism server) side-by-side on identical hardware. - Analyzes reasoning models (e.g. DeepSeek-R1) by inspecting
<think>token patterns and extracting final answers. - Measures true Time to First Token (TTFT) via high-precision streaming probes.
- Monitors hardware saturation in real-time (Metal buffer memory & GPU utilization on Apple Silicon; VRAM, power draw, and temperatures on NVIDIA).
flowchart TD
subgraph CLI ["Benchmark CLI (benchrig/cli.py)"]
A["--check / --runtime (ollama|foundry|onnx-gpu|prism|all)\n--models / --suite / --runs"] --> B[BenchmarkRunner]
end
subgraph Hardware ["Cross-Platform Hardware Abstraction (benchrig/core/hardware.py)"]
B -->|Initialize| HW[HardwareProvider Factory]
HW -->|macOS Darwin| M1["DarwinAppleSiliconProvider\n- sysctl UMA RAM\n- vm_stat memory\n- ioreg GPU load\n- Ollama /api/ps Metal VRAM"]
HW -->|Linux / WSL2| NV["LinuxNvidiaProvider\n- nvidia-smi VRAM\n- /proc/meminfo\n- GPU Power & Temp"]
end
subgraph Runtimes ["Unified Runtime Clients (benchrig/core/client.py)"]
B --> BaseClient["BaseRuntimeClient (base class)"]
BaseClient --> Ollama["OllamaClient (llama.cpp)\n- http://localhost:11434"]
BaseClient --> Foundry["FoundryClient (ONNX Runtime GenAI)\n- Foundry daemon, port from ~/.foundry/daemon.json"]
Foundry --> Prism["PrismClient (Prism server: ONNX GenAI + Ollama)\n- http://127.0.0.1:5272/v1"]
end
subgraph Execution ["Test Execution Engine (benchrig/core/runner.py)"]
Runtimes --> S1["Speed Suite\n(Decode & Prefill TTFT)"]
Runtimes --> S2["Coding Suite\n(Isolated Sandbox Runner)"]
Runtimes --> S3["Reasoning Suite\n(<think> Parser & Verifier)"]
Runtimes --> S4["Polish NLP Suite\n(Declension & Grammar)"]
Runtimes --> S5["Context Scaling\n(512 to 8192 tokens)"]
S2 --> Sandbox["Sandboxed Python Subprocess\ncore/sandbox.py"]
S3 --> ReasonParser["Reasoning Answer Extractor\ncore/reasoning_parser.py"]
end
subgraph Reporting ["Reporting & Output (benchrig/reporting/)"]
B --> Leaderboard["Rich Terminal Leaderboard\n(Dedicated 'Runtime' Column)"]
B --> MarkdownReport["Markdown Report Generator\n(Cross-Engine Delta Section)"]
B --> JSONHistory["JSON History Dumps\nresults/runs/*.json"]
end
| Feature | Description |
|---|---|
| π Multi-Runtime Engine Support | Evaluate models across Ollama (llama.cpp), Microsoft Foundry Local, direct ONNX Runtime GenAI and a Prism server with unified CLI and scoring. See Runtimes. |
| β± High-Precision Timing | Measures generation tokens/sec, prefill tokens/sec, and Time to First Token (TTFT) via nanosecond-precision streaming. |
| π» Automated Sandboxed Coding | Automatically extracts code blocks from LLM responses, wraps them with test harnesses, and executes them in isolated subprocesses against test assertions. |
π§ Reasoning & <think> Parser |
Detects whether models generate chain-of-thought blocks (<think>...</think>), calculates thinking token volume, and extracts final answers. |
| π Context Scaling (512 - 8k) | Progressively loads larger contexts (512, 1024, 2048, 4096, 8192 tokens) to assess TTFT degradation and memory growth. |
| π Hardware Telemetry | Samples GPU/UMA memory usage, GPU utilization %, temperatures, and power draw during execution without requiring root on macOS. |
| π Leaderboard & Markdown Reports | Produces Rich terminal tables with category medals (π₯, π₯, π₯), runtime indicators, and comparative Markdown summaries in results/LATEST_SUMMARY.md. |
Comprehensive step-by-step tutorials and engineering deep dives are available in docs/:
- Tutorial 1: Quickstart Guide β Zero to benchmark in 5 minutes across macOS and Linux/WSL2.
- Tutorial 2: Foundry GPU Setup (WSL2 / Linux) β Complete guide for NVIDIA GPU acceleration on Microsoft Foundry Local (Cache Injection & Direct ONNX GenAI).
- Tutorial 3: Fair 1:1 Cross-Engine Benchmarking β Standardizing parameters, cached baseline evaluation (
--baseline), and scorecards. - Tutorial 4: Authoring Custom Benchmark Scenarios β Designing deterministic coding challenges, reasoning puzzles, and test harnesses.
- System Architecture & Design | CLI Reference & Options | Configuration Reference
- Foundry Local, CUDA & TensorRT Guide | Hardware Telemetry
- Benchmark Suites | Developer & Contributing Guide
pipx install benchrig # or: pip install benchrig
pip install "benchrig[onnx-gpu]" # + direct ONNX Runtime GenAI (CUDA) engine
pip install "benchrig[charts]" # + matplotlib charts
benchrig --version
benchrig --check # diagnose runtimes and accelerators
benchrig --models qwen2.5-coder:7b --suite codingDefaults (config.yaml and the scenario suites) are bundled in the package. Put a ./config.yaml or ./scenarios/ in your
working directory (or pass --config / --scenarios-dir, or set BENCHRIG_CONFIG) to override them. Direct ONNX models are
looked up in $BENCHRIG_MODEL_DIRS, then ./models, then ~/.benchrig/models.
We provide an automated setup script that verifies your Apple Silicon chip, checks Python, creates a virtual environment, installs dependencies, and tests your Ollama connection:
git clone https://github.com/senssei/benchrig.git
cd benchrig
# Run the automated setup script:
chmod +x setup_mac.sh
./setup_mac.shgit clone https://github.com/senssei/benchrig.git
cd benchrig
# Create virtual environment & install in editable mode
python3 -m venv .venv
source .venv/bin/activate
pip install -e ".[dev]"
# Run environment diagnostic check
benchrig --checkVerifies Ollama, MS Foundry, direct ONNX and Prism connectivity, lists installed models across runtimes, and displays detected hardware:
benchrig --checkCompare models running under Ollama (llama.cpp) and MS Foundry (ONNX Runtime GenAI) head-to-head on the same hardware:
# Compare all installed models across both runtimes:
benchrig --runtime all --models installed
# Compare specific models across engines:
benchrig --models ollama:qwen2.5-coder:7b,foundry:phi-4
# Benchmark exclusively on Microsoft Foundry Server:
benchrig --runtime foundry --models phi-4,qwen2.5-coder-7b
# Benchmark ONNX models through a Prism server (`pip install prism-local && prism serve`):
benchrig --runtime prism --models prism:phi-4-miniEach result records the engine that served the model and, for Prism, the device it ran on (cuda or cpu). Prism's default
--device auto runs ONNX models on CUDA even when the model name says generic-cpu; start prism serve --device cpu|cuda
when the comparison depends on it. See Runtimes.
benchrig --models qwen2.5-coder:7b,llama3.1:8bAvailable suites: speed, coding, reasoning, polish, context, all.
# Test only coding with automated unit tests (2 runs per model):
benchrig --models qwen2.5-coder:7b --suite coding --runs 2
# Test context scaling up to 8k tokens:
benchrig --models llama3.1:8b --suite contextDownloads missing standard models from the respective Ollama, MS Foundry or Prism (prism pull) catalog:
benchrig --pull-recommended
# Or specify runtime:
benchrig --runtime foundry --pull-recommended
benchrig --runtime prism --pull-recommendedEdit config.yaml to customize execution parameters, endpoints, and scoring weights:
ollama:
base_url: "http://localhost:11434"
timeout_sec: 180
warmup: true # Pre-warms weights into memory before timing
unload_after_test: true # Keeps memory clean between model runs
default_num_ctx: 4096
foundry:
base_url: "http://localhost:5272/v1" # Microsoft Foundry Local REST endpoint
auto_detect_port: true # Automatically queries active port via CLI
cli_path: "foundry" # Optional Foundry Local CLI executable
timeout_sec: 180
warmup: true
unload_after_test: true
prism:
base_url: "http://127.0.0.1:5272/v1" # Prism server (`prism serve`); key via $PRISM_API_KEY
timeout_sec: 180
warmup: true
unload_after_test: false # Prism loads models on demand
benchmark:
default_runs: 1
default_runtime: "ollama" # "ollama", "foundry", "onnx-gpu", "prism" or "all"
composite_weights:
coding: 0.40 # 40% automated unit tests pass rate
reasoning: 0.30 # 30% reasoning & logic ground truth
performance: 0.30 # 30% normalized decode speed & memory efficiency
hardware:
sample_interval_sec: 0.15
# "auto" computes 88% of detected VRAM / Unified Memory
vram_warning_threshold_mb: "auto"The ollama-coder / foundry-coder agent skills and the stdio MCP servers that offload coding tasks to local models now live in
their own repository: senssei/local-coders.
Test cases are stored as clean JSON files inside the scenarios/ directory:
- scenarios/coding.json: Python coding tasks paired with test assertion arrays evaluated in sandbox subprocesses.
- scenarios/reasoning.json: Multi-step math and logic puzzles with expected ground truth strings and regex patterns.
- scenarios/speed.json: Raw generation and prefill throughput prompts.
- scenarios/context_scaling.json: Context scaling tests up to 8k tokens.
- scenarios/polish.json: Multilingual tests checking Polish language morphology and syntax.
Detailed architecture, configuration guides, benchmark specifications, and operational manuals are available in the docs/ directory:
| Document | Description |
|---|---|
| π System Architecture | Deep dive into the Hardware Abstraction Layer (HAL), Runtime Abstraction Layer (RAL), sandbox isolation, and reporting pipeline. |
| π WSL2 CUDA & TensorRT Guide | Complete guide to configuring Microsoft Foundry Local with NVIDIA CUDA and TensorRT acceleration on WSL2. |
| π§ Runtimes | The four runtimes (Ollama, Foundry Local, direct ONNX, Prism): how BenchRig connects to each, model prefixes, and how to read the engine and device in results. |
| βοΈ Configuration Reference | Full reference for config.yaml, environment variables (OLLAMA_HOST, FOUNDRY_BASE_URL), and dynamic port discovery. |
| π§ͺ Benchmark Suites Mechanics | Evaluation methodology for Speed, Coding, Reasoning, Polish NLP, and Context Scaling suites. |
| π» CLI Usage & Recipes | Command-line parameters, scenario filtering, cross-engine flags, and automation scripts. |
| π Hardware Telemetry & Profiling | Real-time GPU VRAM, compute load, Apple Silicon UMA memory, power draw, and temperature sampling. |
| π Scenario Authoring Guide | Schema reference and instructions for creating custom coding, reasoning, and context scaling scenarios. |
| π©βπ» Developer & Contributing Guide | Guide for adding runtime clients (BaseRuntimeClient), running test suites, and adhering to sandbox security. |
benchrig/
βββ benchrig/ # The installable package (import `benchrig`, command `benchrig`)
β βββ cli.py # CLI entrypoint
β βββ data/config.yaml # Bundled default configuration (endpoints, weights, thresholds)
β βββ data/scenarios/ # Bundled test scenario definitions (JSON)
β βββ core/ # Runtime clients, hardware providers, runner, sandbox, reasoning parser
β βββ reporting/ # Rich terminal UI and Markdown report generator
βββ pyproject.toml # Packaging metadata, ruff and pytest configuration
βββ CHANGELOG.md CONTRIBUTING.md SECURITY.md
βββ setup_mac.sh # Quickstart installer for macOS Apple Silicon
βββ LICENSE # MIT License
βββ README.md # Project documentation
βββ AGENTS.md # Guidelines for AI agents working on this repo
βββ .github/workflows/ci.yml # CI: ruff lint/format check + pytest (Python 3.10-3.13, wheel smoke test)
βββ docs/ # Comprehensive documentation guides (guides + 4 tutorials)
β βββ README.md # Documentation & tutorials index
β βββ architecture.md
β βββ benchmark-suites.md
β βββ cli.md
β βββ configuration.md
β βββ development.md
β βββ foundry-wsl-cuda-tensorrt.md
β βββ hardware-telemetry.md
β βββ scenarios.md
β βββ tutorials/ # Hands-on step-by-step tutorials
β βββ quickstart.md
β βββ foundry-gpu-setup.md
β βββ cross-engine-benchmarking.md
β βββ custom-scenarios.md
βββ examples/
β βββ calculator.py # Sample module for test generation benchmarks
β βββ run_onnx_gpu.py # Standalone direct ONNX GenAI CUDA runner
βββ results/
β βββ 1TO1_COMPARISON_REPORT.md # Cross-engine comparative report
β βββ LATEST_SUMMARY.md # Latest benchmark Markdown report
β βββ latest.json # Latest scorecard JSON
β βββ runs/ # Historical benchmark runs
βββ tests/ # Offline unit test suite (no network or GPU required)
βββ test_benchmark_cli.py
βββ test_foundry_runtime.py
βββ test_hardware.py
βββ test_onnx_client.py
βββ test_packaging.py
βββ test_runner_suites.py
βββ test_sandbox.py
βββ test_token_savings.py
This project is licensed under the MIT License - see the LICENSE file for details.