Skip to content

[Klaud Cold] Update qwen3.5-fp8-h200-sglang-agentic-hicache-mtp SGLang image to v0.5.19-cu130 / 将 qwen3.5-fp8-h200-sglang-agentic-hicache-mtp 的 SGLang 镜像更新至 v0.5.19-cu130 - #2966

Merged
adibarra merged 3 commits into
mainfrom
klaud/auto-8e051edc9c729447-0a3fae3324b49107
Sep 10, 2026
Merged

[Klaud Cold] Update qwen3.5-fp8-h200-sglang-agentic-hicache-mtp SGLang image to v0.5.19-cu130 / 将 qwen3.5-fp8-h200-sglang-agentic-hicache-mtp 的 SGLang 镜像更新至 v0.5.19-cu130#2966
adibarra merged 3 commits into
mainfrom
klaud/auto-8e051edc9c729447-0a3fae3324b49107

Conversation

@Klaud-Cold

@Klaud-Cold Klaud-Cold commented Sep 10, 2026

Copy link
Copy Markdown
Collaborator

Move the qwen3.5-fp8-h200-sglang-agentic-hicache-mtp master image from the dev nightly lmsysorg/sglang:nightly-dev-cu13-20260907-30705c00 to the current SGLang release lmsysorg/sglang:v0.5.19-cu130 (Docker Hub digest sha256:d6e7288627be8b02be88e4bba38e73f6d50e2826869f753c13a4c4385ab3eda9, image label ai.sglang.build.commit = sgl-project/sglang@0bcd822, tag v0.5.19). Model, TP8/EP1 topology, EAGLE MTP settings, DRAM HiCache offload, the concurrency list and benchmarks/single_node/agentic/qwen3.5_fp8_h200_mtp.sh are unchanged.

Baseline

  • Published date: 2026-09-08 (GET /api/v1/benchmarks?model=Qwen-3.5-397B-A17B&date=2026-09-08&exact=true, GET /api/v1/workflow-info?date=2026-09-08&benchmarkType=agentic_traces, GET /api/v1/evaluations?model=Qwen-3.5-397B-A17B)
  • Old image: lmsysorg/sglang:nightly-dev-cu13-20260907-30705c00 (digest sha256:19b8fa1223cc339c1eae7a5b703f1a8c2543b5b119155bf3d7efaef18f77f007, image label build commit sgl-project/sglang@30705c00, main as of 2026-09-07)
  • Workload / topology: single-node H200 (cluster:h200-dgxc), Qwen/Qwen3.5-397B-A17B-FP8, SGLang FP8 with fp8_e4m3 KV, TP8 EP1, EAGLE MTP (3 steps, top-k 1, 4 draft tokens, golden acceptance length 3.39), DRAM HiCache offload, Agentic Traces (agentic-coding, with-subagents 256k corpus), concurrency 2 / 4 / 8 / 10 / 12 / 16 / 20 / 24, recipe fingerprint 8b0a0360a358c71d6f9c553603f93792cab395d86dbf8ac73b5b1be735e6378d
  • Producer: run 34173546821 (head 9eb6d08a30711f971c6a5e8577183720c96851dc, changelog PR #2868); benchmark result IDs 441163, 441151, 441164, 441160, 441162, 441154, 441159, 441155. The logical curve_workflow_run_id 2408 is a snapshot ID, not the producer.
Conc Output tok/s/GPU Input tok/s/GPU Median TTFT (s) Median TPOT (ms) Median E2E (s) GPU cache hit
2 24.08 2158.7 0.541 3.26 1.85 0.944
4 31.85 2945.2 0.517 3.56 2.04 0.944
8 48.46 4891.9 0.604 4.37 2.25 0.944
10 59.54 5922.7 0.694 5.15 2.81 0.944
12 71.28 7638.9 0.629 5.92 2.61 0.993
16 86.74 9771.9 0.772 7.29 3.65 0.947
20 99.68 11578.5 0.847 9.51 4.13 0.948
24 110.16 13411.3 0.926 11.55 5.06 0.945
  • Published eval: gsm8k at concurrency 24, em_strict 0.9773 / em_flexible 0.9765 (evaluation ID 10031, same producer run). The public evaluations feed labels this row disagg: true, which does not match the single-node TP8 recipe; the benchmark rows above carry the correct disagg: false identity.

qwen3.5-fp8-h200-sglang-agentic-hicache-mtp 的主配置镜像从开发 nightly lmsysorg/sglang:nightly-dev-cu13-20260907-30705c00 切换到当前 SGLang 发布版 lmsysorg/sglang:v0.5.19-cu130(Docker Hub 摘要 sha256:d6e7288627be8b02be88e4bba38e73f6d50e2826869f753c13a4c4385ab3eda9,镜像标签 ai.sglang.build.commit = sgl-project/sglang@0bcd822,标签 v0.5.19)。模型、TP8/EP1 拓扑、EAGLE MTP 设置、DRAM HiCache 卸载、并发列表以及 benchmarks/single_node/agentic/qwen3.5_fp8_h200_mtp.sh 均保持不变。

基线

  • 发布日期: 2026-09-08(GET /api/v1/benchmarks?model=Qwen-3.5-397B-A17B&date=2026-09-08&exact=trueGET /api/v1/workflow-info?date=2026-09-08&benchmarkType=agentic_tracesGET /api/v1/evaluations?model=Qwen-3.5-397B-A17B
  • 旧镜像: lmsysorg/sglang:nightly-dev-cu13-20260907-30705c00(摘要 sha256:19b8fa1223cc339c1eae7a5b703f1a8c2543b5b119155bf3d7efaef18f77f007,镜像标签构建提交 sgl-project/sglang@30705c00,即 2026-09-07 的 main)
  • 工作负载 / 拓扑: 单节点 H200(cluster:h200-dgxc),Qwen/Qwen3.5-397B-A17B-FP8,SGLang FP8 与 fp8_e4m3 KV,TP8 EP1,EAGLE MTP(3 步、top-k 1、4 个草稿 token、黄金接受长度 3.39),DRAM HiCache 卸载,Agentic Traces(agentic-coding,含子代理的 256k 语料),并发 2 / 4 / 8 / 10 / 12 / 16 / 20 / 24,配方指纹 8b0a0360a358c71d6f9c553603f93792cab395d86dbf8ac73b5b1be735e6378d
  • 数据来源: 运行 34173546821(head 9eb6d08a30711f971c6a5e8577183720c96851dc,changelog PR #2868);基准结果 ID 441163、441151、441164、441160、441162、441154、441159、441155。逻辑 curve_workflow_run_id 2408 是快照 ID,不是生产运行。
并发 输出 tok/s/GPU 输入 tok/s/GPU TTFT 中位数 (s) TPOT 中位数 (ms) 端到端中位数 (s) GPU 缓存命中率
2 24.08 2158.7 0.541 3.26 1.85 0.944
4 31.85 2945.2 0.517 3.56 2.04 0.944
8 48.46 4891.9 0.604 4.37 2.25 0.944
10 59.54 5922.7 0.694 5.15 2.81 0.944
12 71.28 7638.9 0.629 5.92 2.61 0.993
16 86.74 9771.9 0.772 7.29 3.65 0.947
20 99.68 11578.5 0.847 9.51 4.13 0.948
24 110.16 13411.3 0.926 11.55 5.06 0.945
  • 已发布评测: 并发 24 的 gsm8k,em_strict 0.9773 / em_flexible 0.9765(评测 ID 10031,同一数据来源运行)。公共评测接口将该行标记为 disagg: true,与单节点 TP8 配方不符;上表基准行的 disagg: false 身份是正确的。

🤖 Generated with Claude Code


Note

Low Risk
Benchmark config and changelog only; no application or infra logic changes beyond the SGLang container image pin.

Overview
Pins the qwen3.5-fp8-h200-sglang-agentic-hicache-mtp master config from the cu13 dev nightly lmsysorg/sglang:nightly-dev-cu13-20260907-30705c00 to the release image lmsysorg/sglang:v0.5.19-cu130.

Adds a perf-changelog.yaml entry for that config key documenting the image swap (digests/build commits) and stating that model, TP8/EP1 DRAM HiCache topology, EAGLE MTP settings, concurrency sweep, and benchmarks/single_node/agentic/qwen3.5_fp8_h200_mtp.sh are unchanged.

Reviewed by Cursor Bugbot for commit cf5902f. Bugbot is set up for automated code reviews on this repo. Configure here.

@github-actions

Copy link
Copy Markdown
Contributor

Thanks for the contribution! Please reach out to respective companies' CODEOWNER to fill in the latest PR_REVIEW_CHECKLIST.md before pinging core maintainer on Slack for review. In order for the signoff PR check bot to trigger, you must follow the PR_REVIEW_CHECKLIST.md template correctly, including the phrase As a PR reviewer and CODEOWNER, I have reviewed this and have.

For PR verification, add the full-sweep-fail-fast label (strongly recommended) to this PR — the benchmark sweep only runs on labeled PRs. Use full-sweep-enabled only if you need matrix jobs to keep running past a failure.

PR authors are responsible for ensuring that after merging, all GitHub Action jobs fully pass. A lot of the time, failures are just flakes and simply re-running the failed jobs will fix it. See GitHub's docs on re-running failed jobs


感谢你的贡献!请联系相应公司的 CODEOWNER 填写最新的 PR_REVIEW_CHECKLIST.md,然后再在 Slack 上联系核心维护者进行审阅。为了触发 signoff PR 检查机器人,你必须正确遵循 PR_REVIEW_CHECKLIST.md 模板,包括保留英文语句 As a PR reviewer and CODEOWNER, I have reviewed this and have

如需进行 PR 验证,请为此 PR 添加 full-sweep-fail-fast 标签(强烈推荐)— 基准测试 sweep 仅在带有标签的 PR 上运行。仅当需要矩阵任务在失败后继续运行时才使用 full-sweep-enabled

PR 作者有责任确保合并后所有 GitHub Action 任务完全通过。 很多时候失败只是偶发抖动(flake),重新运行失败的任务即可解决。参见 GitHub 关于重新运行失败任务的文档

@Klaud-Cold

Klaud-Cold commented Sep 10, 2026

Copy link
Copy Markdown
Collaborator Author

Initial attempt

  • Image: lmsysorg/sglang:v0.5.19-cu130 (digest sha256:d6e7288627be8b02be88e4bba38e73f6d50e2826869f753c13a4c4385ab3eda9, image label build commit 0bcd822, pushed 2026-09-04). Head 83ce61be139f9e2c3d33321e7011ed96c3f30918.
  • Change: only the qwen3.5-fp8-h200-sglang-agentic-hicache-mtp image line in configs/nvidia-master.yaml. Model, TP8/EP1, EAGLE MTP shape, golden AL 3.39, DRAM HiCache, concurrency list and the shared launch script are untouched.
  • Engine source delta: the old nightly is main at 30705c00 (2026-09-07). The release branch forked from main at 6c72b49 (2026-09-01) and carries 14 cherry-picks, all AMD/ROCm, DP-attention, MXFP8 or CI fixes outside this Hopper FP8 path (compare). The release therefore lacks the 288 main commits the nightly had; the ones on this recipe's path are the Qwen3.5 GDN prefill projection layout optimization ([Performance] Optimize Qwen3.5 GDN prefill projection layouts sgl-project/sglang#36267), the mamba radix-cache SSM state indexing fix (#37836), HiCache write-through pending alignment (#37278), EAGLE draft-extend logits pruning (#35546), CUDA-graph pool sizing from warmup (#36911) and the FlashInfer sliding-window D2H sync removal (#32218). Moving to the release is a deliberate stability trade; the smoke and full sweep decide whether it holds.
  • Flag and env check at v0.5.19: every option the script passes (--enable-hierarchical-cache, --hicache-size, --hicache-io-backend kernel, --hicache-mem-layout page_first, --hicache-write-policy write_through_selective, --speculative-algorithm EAGLE with steps/topk/draft tokens, --mamba-ssm-dtype, --tokenizer-worker-num, --scheduler-recv-interval, --enable-flashinfer-allreduce-fusion, --kv-cache-dtype fp8_e4m3, --page-size 64) is defined in python/sglang/srt/server_args.py at the tag, and SGLANG_SIMULATE_ACC_LEN / _METHOD / _TOKEN_MODE are declared in environ.py at both commits with the same defaults. server_args.py differs between the two commits only in prefill context-parallel, unified-cache and UNO options that this recipe does not use.
  • Runtime patching: the launch script's conditional sed on multi_tokenizer_mixin.py is a no-op on this image because v0.5.19 already emits cached_tokens_details in the BatchStrOutput branch (lines 313-315 at the tag, identical to the nightly). No patch executes; the image runs as shipped.
  • Stack compatibility: the release image has the same CUDA 13.0.3, FlashInfer 0.6.18, sgl-kernel 0.4.6.post1 and NCCL 2.28.3 as the nightly (image config env), so H200 driver requirements are unchanged. PR [Klaud Cold] Update dsr1-fp8-h200-sglang-mtp SGLang image to v0.5.19-cu130 / 将 dsr1-fp8-h200-sglang-mtp 的 SGLang 镜像升级至 v0.5.19-cu130 #2955 already completed a green full sweep with v0.5.19-cu130 on cluster:h200-dgxc.
  • Smoke plan: e2e-tests.yml on main with ref=83ce61be…, test-config --config-files configs/nvidia-master.yaml --config-keys qwen3.5-fp8-h200-sglang-agentic-hicache-mtp --trim-conc, fail-fast, Klaud background priority. Generator output: one point, TP8/EP1 DRAM HiCache MTP at concurrency 2 on cluster:h200-dgxc; test-config mode selects no agentic eval, so the gsm8k check happens in the final sweep.
  • Run: 34477898041 (dispatched 2026-09-10 12:38 UTC, e2e-tests.yml on main, ref 83ce61be…).
  • Result: targeted smoke passed (run 34477898041, job on h200-dgxc-slurm_00, server ready at 12:50 UTC, replay 12:51-13:51 UTC). 644 trace records, 622 profiled, 22 warmup-dropped, 0 errors; power valid. The job log shows the grep -q cached_tokens_details guard succeeding, so the sed -i branch never ran. Concurrency-2 comparison against the published 2026-09-08 point (same TP8/EP1 DRAM HiCache MTP shape, same 256k with-subagents corpus):
Metric (conc 2) Baseline (nightly 30705c00) Smoke (v0.5.19-cu130) Delta
Output tok/s/GPU 24.08 24.15 +0.3%
Input tok/s/GPU 2158.7 2172.6 +0.6%
Median TTFT (s) 0.541 0.574 +6.2% (slower)
Median TPOT (ms) 3.26 3.47 +6.4% (slower)
Median E2E (s) 1.849 1.881 +1.7% (slower)
Mean TTFT (s) 0.761 0.780 +2.5% (slower)
GPU cache hit rate 0.944 0.965 +2.1 pt

Throughput is flat; per-token latency at this lowest-concurrency point is a few percent slower on the release. Smoke is compatibility evidence only, not a full curve. gsm8k: N/A in test-config mode (no agentic eval selected); it runs in the final sweep.

  • Next: done; the changelog entry, full-sweep-enabled transition and exact-head full sweep are recorded in the Final full sweep comment (passed, 0 repairs used).

初始尝试

  • 镜像: lmsysorg/sglang:v0.5.19-cu130(摘要 sha256:d6e7288627be8b02be88e4bba38e73f6d50e2826869f753c13a4c4385ab3eda9,镜像标签构建提交 0bcd822,推送于 2026-09-04)。Head 83ce61be139f9e2c3d33321e7011ed96c3f30918
  • 改动: 仅修改 configs/nvidia-master.yamlqwen3.5-fp8-h200-sglang-agentic-hicache-mtp 的镜像行。模型、TP8/EP1、EAGLE MTP 形态、黄金 AL 3.39、DRAM HiCache、并发列表和共享启动脚本均未改动。
  • 引擎源码差异: 旧 nightly 对应 main 的 30705c00(2026-09-07)。发布分支于 2026-09-01 从 main 的 6c72b49 分出,并包含 14 个 cherry-pick,全部为 AMD/ROCm、DP-attention、MXFP8 或 CI 修复,不在本 Hopper FP8 路径上(对比)。因此发布版缺少 nightly 中的 288 个 main 提交;其中与本配方路径相关的有:Qwen3.5 GDN prefill 投影布局优化([Performance] Optimize Qwen3.5 GDN prefill projection layouts sgl-project/sglang#36267)、mamba radix cache SSM 状态索引修复(#37836)、HiCache write-through pending 对齐(#37278)、EAGLE draft-extend logits 裁剪(#35546)、基于 warmup 的 CUDA graph 池大小(#36911)以及 FlashInfer 滑窗 D2H 同步移除(#32218)。切换到发布版是有意的稳定性取舍;由 smoke 和完整 sweep 决定是否成立。
  • v0.5.19 的参数与环境变量检查: 脚本传入的所有选项(--enable-hierarchical-cache--hicache-size--hicache-io-backend kernel--hicache-mem-layout page_first--hicache-write-policy write_through_selective--speculative-algorithm EAGLE 及步数/topk/草稿 token、--mamba-ssm-dtype--tokenizer-worker-num--scheduler-recv-interval--enable-flashinfer-allreduce-fusion--kv-cache-dtype fp8_e4m3--page-size 64)在该标签的 python/sglang/srt/server_args.py 中均有定义;SGLANG_SIMULATE_ACC_LEN / _METHOD / _TOKEN_MODE 在两个提交的 environ.py 中均已声明且默认值相同。两个提交间 server_args.py 的差异仅涉及 prefill context-parallel、unified-cache 和 UNO 选项,本配方均未使用。
  • 运行时补丁: 启动脚本对 multi_tokenizer_mixin.py 的条件 sed 在本镜像上为空操作,因为 v0.5.19 的 BatchStrOutput 分支已经输出 cached_tokens_details(标签处第 313-315 行,与 nightly 相同)。不会执行任何补丁;镜像按原样运行。
  • 技术栈兼容性: 发布版镜像与 nightly 使用相同的 CUDA 13.0.3、FlashInfer 0.6.18、sgl-kernel 0.4.6.post1 和 NCCL 2.28.3(镜像配置环境变量),因此 H200 驱动要求不变。PR [Klaud Cold] Update dsr1-fp8-h200-sglang-mtp SGLang image to v0.5.19-cu130 / 将 dsr1-fp8-h200-sglang-mtp 的 SGLang 镜像升级至 v0.5.19-cu130 #2955 已在 cluster:h200-dgxc 上用 v0.5.19-cu130 完成绿色完整 sweep。
  • Smoke 计划:main 上运行 e2e-tests.ymlref=83ce61be…test-config --config-files configs/nvidia-master.yaml --config-keys qwen3.5-fp8-h200-sglang-agentic-hicache-mtp --trim-conc,fail-fast,Klaud 后台优先级。生成器输出:cluster:h200-dgxc 上 TP8/EP1 DRAM HiCache MTP 并发 2 的单个点;test-config 模式不选择 agentic 评测,gsm8k 检查将在最终 sweep 中进行。
  • 运行: 34477898041(2026-09-10 12:38 UTC 派发,main 上的 e2e-tests.yml,ref 83ce61be…)。
  • 结果: 定向 smoke 通过(运行 34477898041,作业位于 h200-dgxc-slurm_00,服务器于 12:50 UTC 就绪,回放 12:51-13:51 UTC)。644 条 trace 记录,622 条计入,22 条 warmup 丢弃,0 条错误;功耗有效。作业日志显示 grep -q cached_tokens_details 守卫成功,因此 sed -i 分支从未执行。并发 2 与 2026-09-08 已发布点(相同 TP8/EP1 DRAM HiCache MTP 形态、相同 256k 含子代理语料)的对比:
指标(并发 2) 基线(nightly 30705c00) Smoke(v0.5.19-cu130) 差异
输出 tok/s/GPU 24.08 24.15 +0.3%
输入 tok/s/GPU 2158.7 2172.6 +0.6%
TTFT 中位数 (s) 0.541 0.574 +6.2%(变慢)
TPOT 中位数 (ms) 3.26 3.47 +6.4%(变慢)
端到端中位数 (s) 1.849 1.881 +1.7%(变慢)
TTFT 均值 (s) 0.761 0.780 +2.5%(变慢)
GPU 缓存命中率 0.944 0.965 +2.1 个百分点

吞吐持平;该最低并发点上发布版的逐 token 延迟慢几个百分点。Smoke 仅为兼容性证据,不是完整曲线。gsm8k:test-config 模式下为 N/A(未选择 agentic 评测);将在最终 sweep 中运行。

  • 下一步: 已完成;changelog 条目、full-sweep-enabled 切换以及精确 head 的完整 sweep 记录在最终完整 sweep 评论中(已通过,使用 0 次修复)。

Klaud-Cold added a commit that referenced this pull request Sep 10, 2026
…pdate to v0.5.19-cu130

Append the perf-changelog entry for moving the H200 Qwen3.5 FP8 SGLang AgentX HiCache MTP recipe to lmsysorg/sglang:v0.5.19-cu130 (PR #2966).

为将 H200 Qwen3.5 FP8 SGLang AgentX HiCache MTP 配方切换到 lmsysorg/sglang:v0.5.19-cu130(PR #2966)追加 perf-changelog 条目。

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@Klaud-Cold

Klaud-Cold commented Sep 10, 2026

Copy link
Copy Markdown
Collaborator Author

Final full sweep

  • Head: 45b8b963b (image change 83ce61be plus the appended perf-changelog.yaml entry; prior changelog bytes preserved, validate_perf_changelog.py and process_changelog.py pass locally).

  • Matrix from the changelog entry: 8 agentic points, TP8/EP1 DRAM HiCache MTP at concurrency 2 / 4 / 8 / 10 / 12 / 16 / 20 / 24 on cluster:h200-dgxc, plus the default agentic gsm8k eval at concurrency 24. No scenario, eval-selection or append-only modifiers.

  • Label: full-sweep-enabled applied after a passing capacity check; PR stays draft.

  • Run: run-sweep.yml 34486082107 (labeled event, 13:59 UTC) on head 45b8b963b; the push-triggered 34486032361 (13:58 UTC) is the same head's synchronize event.

  • Result: passed. run-sweep.yml 34486082107 completed success on head 45b8b963b (16:38 UTC): all 8 agentic points green (concurrency 2 / 8 / 16 / 20 at 15:17-15:21, 10 / 24 at 15:47-15:49, 4 / 12 at 16:34-16:36 UTC), the concurrency-24 gsm8k eval green at 14:29 UTC, klaud-sweep-manifest and changelog-metadata uploaded, zero request errors on every point, power valid on every point. Every point ran lmsysorg/sglang:v0.5.19-cu130 with TP8/EP1, MTP and the HiCache backend. gsm8k: em_strict 0.9765 / em_flexible 0.9757 versus published 0.9773 / 0.9765 (within one standard error).

    Per-point comparison against the published 2026-09-08 curve (same TP8/EP1 DRAM HiCache MTP recipe, same 256k with-subagents corpus, baseline → this sweep):

Conc Output tok/s/GPU Input tok/s/GPU Median TTFT (s) Median TPOT (ms) Median E2E (s)
2 24.08 → 24.03 (-0.2%) 2158.7 → 2140.7 (-0.8%) 0.541 → 0.548 (+1.3%) 3.26 → 3.42 (+4.9%) 1.849 → 1.870 (+1.1%)
4 31.85 → 32.43 (+1.8%) 2945.2 → 2955.3 (+0.3%) 0.517 → 0.533 (+3.1%) 3.56 → 3.74 (+5.1%) 2.036 → 2.191 (+7.6%)
8 48.46 → 46.99 (-3.0%) 4891.9 → 4874.9 (-0.3%) 0.604 → 0.601 (-0.5%) 4.37 → 4.39 (+0.5%) 2.252 → 2.267 (+0.7%)
10 59.54 → 58.31 (-2.1%) 5922.7 → 5849.6 (-1.2%) 0.694 → 0.664 (-4.3%) 5.15 → 5.26 (+2.1%) 2.805 → 2.831 (+0.9%)
12 71.28 → 70.95 (-0.5%) 7638.9 → 7658.3 (+0.3%) 0.628 → 0.672 (+6.9%) 5.92 → 6.10 (+3.0%) 2.612 → 2.703 (+3.5%)
16 86.74 → 86.46 (-0.3%) 9771.9 → 9718.1 (-0.6%) 0.772 → 0.792 (+2.6%) 7.29 → 7.32 (+0.4%) 3.652 → 3.786 (+3.7%)
20 99.68 → 99.15 (-0.5%) 11578.5 → 11494.9 (-0.7%) 0.847 → 0.877 (+3.5%) 9.51 → 9.63 (+1.3%) 4.129 → 4.173 (+1.1%)
24 110.16 → 109.52 (-0.6%) 13411.3 → 13362.6 (-0.4%) 0.926 → 0.924 (-0.3%) 11.55 → 11.64 (+0.8%) 5.061 → 5.129 (+1.3%)

Throughput is flat to slightly lower (worst -3.0% output tok/s/GPU at concurrency 8, +1.8% at concurrency 4). Latency regresses mildly on the release: median TPOT is 0.4-5.1% slower at every point, median TTFT is slower at 6 of 8 points (up to +6.9% at concurrency 12), median E2E is 0.7-7.6% slower everywhere. These are reported as regressions without a rejection threshold; the release trades the nightly-only Qwen3.5 GDN prefill and EAGLE optimizations listed in the Initial attempt for a tagged, reproducible engine. Server-reported GPU cache hit rate is not compared: this sweep's values exceed 1.0 at concurrency 12-24, so the counter is not directly comparable to the baseline's 0.94-0.99.

  • Status: the sweep itself passed, but finish rejected this head's changelog entry for carrying a scenario-type modifier (coverage-selection rule, not a benchmark problem). The entry is corrected in Repair 1/5 and the complete sweep is re-run on the corrected head; see that comment for the authoritative final validation.

最终完整 sweep

  • Head: 45b8b963b(镜像改动 83ce61be 加上追加的 perf-changelog.yaml 条目;历史 changelog 字节完整保留,本地 validate_perf_changelog.pyprocess_changelog.py 均通过)。

  • 由 changelog 条目生成的矩阵: 8 个 agentic 点,cluster:h200-dgxc 上 TP8/EP1 DRAM HiCache MTP 并发 2 / 4 / 8 / 10 / 12 / 16 / 20 / 24,另加并发 24 的默认 agentic gsm8k 评测。无 scenario、评测选择或 append-only 修饰符。

  • 标签: 容量检查通过后应用 full-sweep-enabled;PR 保持草稿。

  • 运行: head 45b8b963b 上的 run-sweep.yml 34486082107(labeled 事件,13:59 UTC);由推送触发的 34486032361(13:58 UTC)是同一 head 的 synchronize 事件。

  • 结果: 通过。run-sweep.yml 34486082107 在 head 45b8b963b 上以 success 完成(16:38 UTC):全部 8 个 agentic 点通过(并发 2 / 8 / 16 / 20 于 15:17-15:21,10 / 24 于 15:47-15:49,4 / 12 于 16:34-16:36 UTC),并发 24 的 gsm8k 评测于 14:29 UTC 通过,klaud-sweep-manifestchangelog-metadata 已上传,每个点请求错误为零,每个点功耗有效。所有点均运行 lmsysorg/sglang:v0.5.19-cu130,TP8/EP1、MTP 与 HiCache 后端。gsm8k:em_strict 0.9765 / em_flexible 0.9757,已发布值为 0.9773 / 0.9765(在一个标准误内)。

    与 2026-09-08 已发布曲线的逐点对比(相同 TP8/EP1 DRAM HiCache MTP 配方、相同 256k 含子代理语料,基线 → 本次 sweep):

并发 输出 tok/s/GPU 输入 tok/s/GPU TTFT 中位数 (s) TPOT 中位数 (ms) 端到端中位数 (s)
2 24.08 → 24.03 (-0.2%) 2158.7 → 2140.7 (-0.8%) 0.541 → 0.548 (+1.3%) 3.26 → 3.42 (+4.9%) 1.849 → 1.870 (+1.1%)
4 31.85 → 32.43 (+1.8%) 2945.2 → 2955.3 (+0.3%) 0.517 → 0.533 (+3.1%) 3.56 → 3.74 (+5.1%) 2.036 → 2.191 (+7.6%)
8 48.46 → 46.99 (-3.0%) 4891.9 → 4874.9 (-0.3%) 0.604 → 0.601 (-0.5%) 4.37 → 4.39 (+0.5%) 2.252 → 2.267 (+0.7%)
10 59.54 → 58.31 (-2.1%) 5922.7 → 5849.6 (-1.2%) 0.694 → 0.664 (-4.3%) 5.15 → 5.26 (+2.1%) 2.805 → 2.831 (+0.9%)
12 71.28 → 70.95 (-0.5%) 7638.9 → 7658.3 (+0.3%) 0.628 → 0.672 (+6.9%) 5.92 → 6.10 (+3.0%) 2.612 → 2.703 (+3.5%)
16 86.74 → 86.46 (-0.3%) 9771.9 → 9718.1 (-0.6%) 0.772 → 0.792 (+2.6%) 7.29 → 7.32 (+0.4%) 3.652 → 3.786 (+3.7%)
20 99.68 → 99.15 (-0.5%) 11578.5 → 11494.9 (-0.7%) 0.847 → 0.877 (+3.5%) 9.51 → 9.63 (+1.3%) 4.129 → 4.173 (+1.1%)
24 110.16 → 109.52 (-0.6%) 13411.3 → 13362.6 (-0.4%) 0.926 → 0.924 (-0.3%) 11.55 → 11.64 (+0.8%) 5.061 → 5.129 (+1.3%)

吞吐持平或略降(最差为并发 8 的输出 tok/s/GPU -3.0%,并发 4 为 +1.8%)。发布版的延迟轻微回退:TPOT 中位数在每个点慢 0.4-5.1%,TTFT 中位数在 8 个点中的 6 个变慢(并发 12 最多 +6.9%),端到端中位数各点慢 0.7-7.6%。以上作为回退如实报告,不设拒绝阈值;发布版以初始尝试中列出的 nightly 独有 Qwen3.5 GDN prefill 与 EAGLE 优化,换取带标签、可复现的引擎。服务器上报的 GPU 缓存命中率不作比较:本次 sweep 在并发 12-24 的值超过 1.0,该计数器与基线的 0.94-0.99 不可直接比较。

  • 状态: sweep 本身通过,但 finish 因该 head 的 changelog 条目带有 scenario-type 修饰符而拒绝(覆盖选择规则,并非基准问题)。条目已在修复 1/5 中更正,并在更正后的 head 上重跑完整 sweep;最终验证以该评论为准。

@github-actions

Copy link
Copy Markdown
Contributor

@Klaud-Cold

Klaud-Cold commented Sep 10, 2026

Copy link
Copy Markdown
Collaborator Author

Repair 1/5

| Conc | Output tok/s/GPU | Input tok/s/GPU | Median TTFT (s) | Median TPOT (ms) | Median E2E (s) | profiled/errors | power |
| 2 | 24.08 → 24.11 (+0.1%) | 2158.7 → 2167.1 (+0.4%) | 0.541 → 0.591 (+9.3%) | 3.26 → 3.44 (+5.5%) | 1.849 → 1.927 (+4.2%) | 617/0 | 1 |
| 4 | 31.85 → 31.13 (-2.3%) | 2945.2 → 2861.6 (-2.8%) | 0.517 → 0.512 (-1.0%) | 3.56 → 3.64 (+2.2%) | 2.036 → 2.012 (-1.2%) | 946/0 | 1 |
| 8 | 48.46 → 46.32 (-4.4%) | 4891.9 → 4835.3 (-1.2%) | 0.604 → 0.607 (+0.5%) | 4.37 → 4.39 (+0.5%) | 2.252 → 2.296 (+2.0%) | 1446/0 | 1 |
| 10 | 59.54 → 58.71 (-1.4%) | 5922.7 → 5850.6 (-1.2%) | 0.694 → 0.693 (-0.2%) | 5.15 → 5.24 (+1.7%) | 2.805 → 2.831 (+0.9%) | 1680/0 | 1 |
| 12 | 71.28 → 71.01 (-0.4%) | 7638.9 → 7638.5 (-0.0%) | 0.628 → 0.658 (+4.7%) | 5.92 → 6.17 (+4.2%) | 2.612 → 2.783 (+6.5%) | 2234/0 | 1 |
| 16 | 86.74 → 86.24 (-0.6%) | 9771.9 → 9696.9 (-0.8%) | 0.772 → 0.814 (+5.4%) | 7.29 → 7.49 (+2.7%) | 3.652 → 3.750 (+2.7%) | 2678/0 | 1 |
| 20 | 99.68 → 99.16 (-0.5%) | 11578.5 → 11518.0 (-0.5%) | 0.847 → 0.877 (+3.5%) | 9.51 → 9.56 (+0.5%) | 4.129 → 4.238 (+2.6%) | 3162/0 | 1 |
| 24 | 110.16 → 109.42 (-0.7%) | 13411.3 → 13330.5 (-0.6%) | 0.926 → 0.937 (+1.2%) | 11.55 → 11.60 (+0.4%) | 5.061 → 5.132 (+1.4%) | 3600/2 | 1 |

Throughput is flat to slightly lower (worst -4.4% output tok/s/GPU at concurrency 8; -2.3% at 4). Latency regresses mildly on the release: median TPOT is 0.4-5.5% slower at every point, median TTFT slower at 6 of 8 points (up to +9.3% at concurrency 2), median E2E slower at 7 of 8 (up to +6.5% at concurrency 12). The concurrency-24 point dropped 2 of 3600 profiled records as errors (ClientOSError: 2); the earlier green sweep on 45b8b963b had zero at that point. These regressions are reported without a rejection threshold; the picture matches the first full sweep on 45b8b963b (same image, same results within run-to-run noise), so it reflects the release engine rather than a one-off. Server-reported GPU cache hit rate is not compared because this sweep's counter exceeds 1.0 at several points.

  • Status: targeted smoke passed, exact-head final validation passed on 8604d93ed; 1 of 5 repairs used (changelog modifier and base drift, no image or recipe change). Handing to finish to mark the PR ready for automatic reviews; global approval remains with CODEOWNERS.

修复 1/5

| Conc | Output tok/s/GPU | Input tok/s/GPU | Median TTFT (s) | Median TPOT (ms) | Median E2E (s) | profiled/errors | power |
| 2 | 24.08 → 24.11 (+0.1%) | 2158.7 → 2167.1 (+0.4%) | 0.541 → 0.591 (+9.3%) | 3.26 → 3.44 (+5.5%) | 1.849 → 1.927 (+4.2%) | 617/0 | 1 |
| 4 | 31.85 → 31.13 (-2.3%) | 2945.2 → 2861.6 (-2.8%) | 0.517 → 0.512 (-1.0%) | 3.56 → 3.64 (+2.2%) | 2.036 → 2.012 (-1.2%) | 946/0 | 1 |
| 8 | 48.46 → 46.32 (-4.4%) | 4891.9 → 4835.3 (-1.2%) | 0.604 → 0.607 (+0.5%) | 4.37 → 4.39 (+0.5%) | 2.252 → 2.296 (+2.0%) | 1446/0 | 1 |
| 10 | 59.54 → 58.71 (-1.4%) | 5922.7 → 5850.6 (-1.2%) | 0.694 → 0.693 (-0.2%) | 5.15 → 5.24 (+1.7%) | 2.805 → 2.831 (+0.9%) | 1680/0 | 1 |
| 12 | 71.28 → 71.01 (-0.4%) | 7638.9 → 7638.5 (-0.0%) | 0.628 → 0.658 (+4.7%) | 5.92 → 6.17 (+4.2%) | 2.612 → 2.783 (+6.5%) | 2234/0 | 1 |
| 16 | 86.74 → 86.24 (-0.6%) | 9771.9 → 9696.9 (-0.8%) | 0.772 → 0.814 (+5.4%) | 7.29 → 7.49 (+2.7%) | 3.652 → 3.750 (+2.7%) | 2678/0 | 1 |
| 20 | 99.68 → 99.16 (-0.5%) | 11578.5 → 11518.0 (-0.5%) | 0.847 → 0.877 (+3.5%) | 9.51 → 9.56 (+0.5%) | 4.129 → 4.238 (+2.6%) | 3162/0 | 1 |
| 24 | 110.16 → 109.42 (-0.7%) | 13411.3 → 13330.5 (-0.6%) | 0.926 → 0.937 (+1.2%) | 11.55 → 11.60 (+0.4%) | 5.061 → 5.132 (+1.4%) | 3600/2 | 1 |

吞吐持平或略降(最差为并发 8 的输出 tok/s/GPU -4.4%;并发 4 为 -2.3%)。发布版延迟轻微回退:TPOT 中位数在每个点慢 0.4-5.5%,TTFT 中位数在 8 个点中的 6 个变慢(并发 2 最多 +9.3%),端到端中位数在 8 个点中的 7 个变慢(并发 12 最多 +6.5%)。并发 24 点在 3600 条计入记录中有 2 条因错误被丢弃(ClientOSError: 2);此前 45b8b963b 上的绿色 sweep 在该点为零。以上回退如实报告,不设拒绝阈值;结果与 45b8b963b 上的首次完整 sweep 一致(相同镜像,差异在运行间噪声内),因此反映的是发布版引擎本身而非偶发。服务器上报的 GPU 缓存命中率因本次 sweep 多个点的计数器超过 1.0 而不作比较。

  • 状态: 定向 smoke 通过,8604d93ed 上的精确 head 最终验证通过;已使用 5 次修复中的 1 次(changelog 修饰符与基线漂移,未改动镜像或配方)。交由 finish 将 PR 标记为 ready 以启动自动评审;全局批准仍由 CODEOWNERS 决定。

Klaud-Cold and others added 2 commits September 10, 2026 16:51
…ge to v0.5.19-cu130

Move the H200 Qwen3.5 FP8 SGLang AgentX HiCache MTP recipe from the
lmsysorg/sglang:nightly-dev-cu13-20260907-30705c00 dev nightly to the
v0.5.19 release image lmsysorg/sglang:v0.5.19-cu130 (digest
sha256:d6e7288627be8b02be88e4bba38e73f6d50e2826869f753c13a4c4385ab3eda9,
build commit sgl-project/sglang@0bcd822).
Model, TP8/EP1 topology, EAGLE MTP settings, DRAM HiCache offload,
concurrency list and the launch script are unchanged.

将 H200 Qwen3.5 FP8 SGLang AgentX HiCache MTP 配方的镜像从开发 nightly
lmsysorg/sglang:nightly-dev-cu13-20260907-30705c00 切换到 v0.5.19 发布版
lmsysorg/sglang:v0.5.19-cu130(摘要
sha256:d6e7288627be8b02be88e4bba38e73f6d50e2826869f753c13a4c4385ab3eda9,
构建提交 sgl-project/sglang@0bcd822)。
模型、TP8/EP1 拓扑、EAGLE MTP 设置、DRAM HiCache 卸载、并发列表和启动脚本均保持不变。

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…pdate to v0.5.19-cu130

Append the perf-changelog entry for moving the H200 Qwen3.5 FP8 SGLang AgentX HiCache MTP recipe to lmsysorg/sglang:v0.5.19-cu130 (PR #2966), rebased onto current main after other changelog entries landed. The entry selects the whole family without scenario, eval-selection or append-only modifiers.

为将 H200 Qwen3.5 FP8 SGLang AgentX HiCache MTP 配方切换到 lmsysorg/sglang:v0.5.19-cu130(PR #2966)追加 perf-changelog 条目;在其他 changelog 条目合入后已重新基于当前 main。该条目选择整个配方族,不带 scenario、评测选择或 append-only 修饰符。

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@Klaud-Cold
Klaud-Cold force-pushed the klaud/auto-8e051edc9c729447-0a3fae3324b49107 branch from 4531276 to 8604d93 Compare September 10, 2026 16:51
@github-actions

Copy link
Copy Markdown
Contributor

@Klaud-Cold
Klaud-Cold marked this pull request as ready for review September 10, 2026 18:16
@Klaud-Cold
Klaud-Cold requested a review from a team September 10, 2026 18:16
@Klaud-Cold

Copy link
Copy Markdown
Collaborator Author

Klaud Cold: validated. All owned runs are terminal. Repairs: 1. Runs: 34477898041, 34486032361, 34486082107, 34504699566, 34504764956.

The full sweep is verified; this PR remains ready for review.


Klaud Cold:validated。所有自有运行均已结束。修复次数:1。运行:34477898041, 34486032361, 34486082107, 34504699566, 34504764956

完整 sweep 已通过验证;PR 保持就绪,等待审查。

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I reviewed this PR and didn't find any bugs. Because configs/nvidia-master.yaml has named CODEOWNERs (ankur-singh, kedarpotdar-nv) beyond the default wildcard, a human look would still be worthwhile before merging.

What was reviewed: the single-line image bump in configs/nvidia-master.yaml for the qwen3.5-fp8-h200-sglang-agentic-hicache-mtp recipe, and the corresponding append-only perf-changelog.yaml entry (confirmed it only adds 7 lines at EOF, no historical bytes touched). Checked whether the new image: value carries a @ sha256:... digest matching the one cited in the changelog text (it does not — the config uses the bare tag lmsysorg/sglang:v0.5.19-cu130), and whether losing the old nightly tag's baked-in date/commit specificity matters for reproducibility; both were already flagged as investigated/ruled-out candidates rather than new findings.

Extended reasoning...

This is a two-file, mechanical-looking change: a single image: tag bump in configs/nvidia-master.yaml plus a corresponding append to perf-changelog.yaml. The diff is minimal and follows the repository's documented image-bump and append-only changelog conventions (verified the changelog addition is strictly appended at EOF with no modification to prior entries, and that no model.container field exists in this single-node recipe to keep in sync).

No security risks are present — this is a config data change with no code execution paths, credentials, or auth logic involved.

The bug-hunting system reported zero findings, and two candidate issues (digest mismatch between the changelog's claimed sha256 pin and the actual undigested image tag; loss of the old nightly tag's baked-in immutability) were investigated and explicitly ruled out rather than raised as findings. Despite the low complexity, the CODEOWNERS file designates specific named owners (ankur-singh, kedarpotdar-nv) for configs/nvidia-master.yaml beyond the default wildcard team, which per the review guidelines means a human should still weigh in rather than approving outright.

No outstanding CHANGES_REQUESTED or unaddressed third-party objections were visible in the available timeline metadata, and I am not restating any inline findings since none exist. Given the CODEOWNERS designation, deferring rather than approving is the appropriate call here.

@adibarra

Copy link
Copy Markdown
Collaborator

/reuse-sweep-run 34504764956

合并 main 到 PR #2966,保留已验证的配置并复用完整 sweep 结果。
@adibarra
adibarra merged commit fac862f into main Sep 10, 2026
31 checks passed
@adibarra
adibarra deleted the klaud/auto-8e051edc9c729447-0a3fae3324b49107 branch September 10, 2026 21:23
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

Development

Successfully merging this pull request may close these issues.

2 participants