A reproducible, neutral benchmark for choosing a long-context serving stack on NVIDIA GB10. It compares vLLM, SGLang, and direct TensorRT-LLM serving with the same prompts, client-side metric formulas, cache protocols, and concurrency matrix.
The checked-in cross-engine values are preliminary one-repetition evidence. They are useful for portfolio review and experiment planning, but are not a final universal ranking. Run the three-repetition matrix before making production decisions.
- Engine-neutral long-context workload construction and exact token-count validation.
- Cold unique-prefix versus warm shared-prefix cache experiments.
- TTFT, TPOT, ITL, E2E latency, request/output throughput, cache evidence, and host telemetry.
- Sequential Docker orchestration with pinned model, tokenizer, dataset, and image evidence.
- Decision-oriented reporting that separates observations from hypotheses and profiler-backed conclusions.
For the supplied warm shared-prefix comparison at C1/C2/C4, TensorRT-LLM has the lowest reported TTFT and E2E latency and the highest reported request throughput. The supplied data uses one repetition per configuration; the cold C4 TensorRT-LLM ITL P95 anomaly is retained as a profiling target rather than explained causally.
See results analysis, limitations, and the compact summary artifacts for the evidence status and units.
| Dimension | Default control |
|---|---|
| Model | openai/gpt-oss-20b |
| Hardware target | One NVIDIA GB10 / DGX Spark system |
| Prompt / output | Exactly 120,000 input tokens / 512 generated tokens |
| Samples | 100 canonical records |
| Modes | cold, warm_shared |
| Concurrency | 1, 2, 4 |
| Repetitions | 3 for the final matrix |
| Memory target | 0.80 engine-specific fraction |
| Client | One neutral OpenAI-compatible streaming client |
Scheduler, batching, kernel, cache-block, and memory-allocation semantics differ across runtimes. The controls are matched by intent and documented as such; they are not claimed to be mechanically identical.
Install the project and developer tools:
python3 -m pip install -e ".[dev]"Run a GPU/Docker smoke test across all three engines:
./scripts/smoke_test.sh --overwrite --cooldown-seconds 5Run the full three-engine matrix:
./bench run --engines all --modes cold,warm_shared \
--concurrency 1,2,4 --repetitions 3The smoke and full commands perform real inference and require Docker, NVIDIA Container Toolkit, model/data access, and sufficient disk. Ordinary CI runs tests, dry-run orchestration, package checks, and manifest verification only; it does not run GPU inference.
Generate accepted-run tables from the reporting pipeline:
./bench report
python3 scripts/generate_charts.py --input results/report/summary.csvThe report writes an autogenerated marker and results/report/provenance.json containing the benchmark commit, experiment-lock hash, accepted runs, and accepted repetitions. The supplied portfolio charts are stored under assets/charts/ with their compact source CSVs and preliminary-results labeling. The chart generator remains available for regenerating SVG decision views from accepted report summaries.
Do not hand-edit public numeric tables. The intended evidence flow is:
accepted run JSON + request timings
-> ./bench report
-> summary CSV + provenance
-> scripts/generate_charts.py
-> charts and portfolio report
- Methodology and fairness contract
- Architecture
- Results and analysis
- Profiling plan
- Fairness and limitations
- Production recommendations
- Reproducibility and implementation details
- Result schema
- Design-to-implementation mapping
TensorRT-LLM uses the direct trtllm-serve backend with the same prepared prompts and neutral client. Triton integration is intentionally deferred and documented as planned work; it is not represented as a completed experiment.
GitHub Actions runs on Python 3.10, 3.11, and 3.12. It installs .[dev], runs pytest, validates the CLI and three-engine dry-run plan, and verifies PROJECT_MANIFEST.sha256. See .github/workflows/ci.yml.
The repository excludes model weights, Hugging Face caches, TensorRT engines, server logs, raw telemetry, and request JSONL from version control. Large evidence bundles belong in GitHub Releases with their lock file, image digests, environment manifest, and benchmark commit.
- vLLM integration
- SGLang integration
- TensorRT-LLM direct integration
- 120K cold and warm shared-prefix workloads
- Cache evidence and result provenance
- CI and packaging validation
- Three repetitions per matrix cell
- Nsight Systems profiling
- TensorRT-LLM through Triton
- Context scaling: 8K, 32K, 64K, 120K
- SLA-constrained throughput
Maintainer: Venkata Rami Reddy Kallu. See CITATION.cff for citation metadata. The project is licensed under MIT.