[Klaud Cold] Update dsr1-fp8-mi325x-sglang-mtp SGLang ROCm image to v0.5.19-rocm700-mi30x / 将 dsr1-fp8-mi325x-sglang-mtp 的 SGLang ROCm 镜像更新至 v0.5.19-rocm700-mi30x - #2844
Conversation
…0-mi30x Bump the dsr1-fp8-mi325x-sglang-mtp master family image from lmsysorg/sglang:v0.5.12-rocm700-mi30x to lmsysorg/sglang:v0.5.19-rocm700-mi30x. No other fields or families change. 将 dsr1-fp8-mi325x-sglang-mtp 家族的镜像从 lmsysorg/sglang:v0.5.12-rocm700-mi30x 更新为 lmsysorg/sglang:v0.5.19-rocm700-mi30x。不改动其他字段或家族。 Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
|
Thanks for the contribution! Please reach out to respective companies' CODEOWNER to fill in the latest PR_REVIEW_CHECKLIST.md before pinging core maintainer on Slack for review. In order for the signoff PR check bot to trigger, you must follow the PR_REVIEW_CHECKLIST.md template correctly, including the phrase For PR verification, add the PR authors are responsible for ensuring that after merging, all GitHub Action jobs fully pass. A lot of the time, failures are just flakes and simply re-running the failed jobs will fix it. See GitHub's docs on re-running failed jobs 感谢你的贡献!请联系相应公司的 CODEOWNER 填写最新的 PR_REVIEW_CHECKLIST.md,然后再在 Slack 上联系核心维护者进行审阅。为了触发 signoff PR 检查机器人,你必须正确遵循 PR_REVIEW_CHECKLIST.md 模板,包括保留英文语句 如需进行 PR 验证,请为此 PR 添加 PR 作者有责任确保合并后所有 GitHub Action 任务完全通过。 很多时候失败只是偶发抖动(flake),重新运行失败的任务即可解决。参见 GitHub 关于重新运行失败任务的文档 |
Append the perf-changelog.yaml entry for the dsr1-fp8-mi325x-sglang-mtp image update to lmsysorg/sglang:v0.5.19-rocm700-mi30x (PR #2844). 为 dsr1-fp8-mi325x-sglang-mtp 更新至 lmsysorg/sglang:v0.5.19-rocm700-mi30x 的镜像变更追加 perf-changelog.yaml 条目(PR #2844)。 Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Summary / 摘要
Klaud Cold image refresh for the master family
dsr1-fp8-mi325x-sglang-mtp(configs/amd-master.yaml): updateimagefromlmsysorg/sglang:v0.5.12-rocm700-mi30xtolmsysorg/sglang:v0.5.19-rocm700-mi30x, plus the requiredperf-changelog.yamlentry. Model, precision, topology (TP8/EP1), MTP speculative decoding, 8k/1k workload, benchmark script, and eval selection are unchanged. Only this family is touched; the non-MTP siblingdsr1-fp8-mi325x-sglangand all other families stay on their current images.Klaud Cold 为主配置家族
dsr1-fp8-mi325x-sglang-mtp(configs/amd-master.yaml)刷新镜像:将image从lmsysorg/sglang:v0.5.12-rocm700-mi30x更新为lmsysorg/sglang:v0.5.19-rocm700-mi30x,并追加必需的perf-changelog.yaml条目。模型、精度、拓扑(TP8/EP1)、MTP 投机解码、8k/1k 负载、基准脚本与评测选择均保持不变。仅改动该家族;非 MTP 兄弟家族dsr1-fp8-mi325x-sglang及其他家族保持原镜像。Candidate evidence / 候选证据
aac067bd5766d07d-d7346b01615deec4; base SHAe52162e5c4805ccf6933199034dbf2f7ba9f5b55; branchklaude/auto-aac067bd5766d07d-d7346b01615deec4./api/v1/latest-images): dsr1 / mi325x / sglang / fp8 / mtp / single-node / 8192→1024 onlmsysorg/sglang:v0.5.12-rocm700-mi30x, published 2026-05-18. Release feed (/api/v1/framework-releases): sglangv0.5.19.lmsysorg/sglang:v0.5.19-rocm700-mi30xexists on Docker Hub (linux/amd64, pushed 2026-09-04T21:28Z, digestsha256:590a815c128d7d5c83ef771ad768c9f8be82f64f7d0dbd98dc7b9f085b268a58). It keeps the same ROCm 7.0 / MI30x (gfx942) base line as the current pin. Old image digest:sha256:4548e936fd76e707184c449ed4720739220ca53734944e1d330c8c4c1c5215ed. The image runs as shipped; no engine patches, wheel overlays, or script changes.runner: mi325xexpands inconfigs/runners.yamlonly to themi325x-amds_*fleet; the single telemetry clustermi325xpassedcheck-capacitybefore edits, before branch/PR creation and before dispatch. Generated matrix (test-config --config-files configs/amd-master.yaml --config-keys dsr1-fp8-mi325x-sglang-mtp): 5 single-node points (conc 4/8/16/32/64, TP8 EP1, MTP), evals at conc 32 and 64; identical to base exceptimage.benchmarks/single_node/fixed_seq_len/dsr1_fp8_mi325x{,_mtp}.sh.Published baseline (frozen) / 已发布基线(冻结)
Source:
GET /api/v1/benchmarks?model=DeepSeek-R1-0528&date=2026-05-18&exact=truefiltered to hardware=mi325x, framework=sglang, model=dsr1, precision=fp8, spec_method=mtp, disagg=false, isl=8192, osl=1024, image=lmsysorg/sglang:v0.5.12-rocm700-mi30x;GET /api/v1/workflow-info?date=2026-05-18;GET /api/v1/evaluations. Producer run for every point: https://github.com/SemiAnalysisAI/InferenceX/actions/runs/26048239142/attempts/2 (PR #1500; curve snapshot id 1899 is not a producer id). Published date: 2026-05-18.Baseline evals (gsm8k, 2026-05-18, same producer run): conc 32 em_strict 0.9575 / em_flexible 0.9583; conc 64 em_strict 0.9575 / em_flexible 0.9568.
Baseline comparability caveat / 基线可比性说明: the retained baseline server log (
GET /api/v1/server-log-search?id=412356&q=EAGLE, alsoq=speculative_num_steps) showsspeculative_algorithm=None,speculative_num_steps=None,mem_fraction_static=0.68andmax_running_requests=None→4096, i.e. the published "mtp" curve was served without speculative decoding and with the non-MTP launch flags (its metrics also match the non-MTPdsr1-fp8-mi325x-sglangcurve of the same day within noise). The updated-image run below really ran EAGLE/MTP (accept length ≈2.0 on random benchmark traffic, ≈2.3–2.7 on gsm8k). The per-point deltas are therefore reported as informational only and are not claimed as image improvements; they mix the engine update with the MTP-on vs MTP-off difference. / 基线保留的服务端日志显示speculative_algorithm=None(未启用投机解码,且为非 MTP 启动参数),而本次新镜像运行确实启用了 EAGLE/MTP。因此下列逐点变化仅供参考,不作为镜像本身的性能提升声明。Attempts / 尝试记录
lmsysorg/sglang:v0.5.12-rocm700-mi30xklaud-run=true,fail-fast=true)lmsysorg/sglang:v0.5.19-rocm700-mi30x@97045041f5c0dbe62b2d1baef9b56e05b84a5958results_bmk/agg_bmk.json(allpower_valid=1); gsm8k (1319 samples) conc32 em_strict 0.9553 / flexible 0.9553 (baseline 0.9575 / 0.9583, Δ −0.2 pt); conc64 em_strict 0.9507 / flexible 0.9515 (baseline 0.9575 / 0.9568, Δ −0.7 / −0.5 pt, ≈1.2 SE)SGLANG_ENABLE_SPEC_V2is now a removed no-op (spec always runs the V2 worker); the engine now resetsmax_running_requeststo 48 under speculative decoding, so the conc-64 point queues 16 requests server-side and its median TTFT rises to 13.85 s (regression, reported not fixed; recipe flags are out of Klaud scope) / 镜像可直接运行;新引擎在投机解码下将 max_running_requests 重置为 48,conc-64 点排队导致 TTFT 中位数升至 13.85 s(回归,仅报告)full-sweep-enabled) / 最终全量 sweeplmsysorg/sglang:v0.5.19-rocm700-mi30x@1fceee6e753e4cf7959fee2e52bd46fcc276feec(image + changelog)run-sweep.ymlruns concludedskippedwith every job skipped.run-sweep.ymlgatescheck-changelog,reuse-sweep-gateandsetupon!github.event.pull_request.draft, and this PR must stay a draft per Klaud Cold policy, so no GPU sweep ran on the PR head / 不适用 — 未执行:两次 run-sweep 运行均为 skipped(所有作业跳过),因为该工作流要求 PR 非草稿,而 Klaud Cold 须保持草稿状态ready_for_reviewevent is arun-sweep.ymltrigger, so a maintainer converting this draft to ready will start the labeled full sweep on the PR head; nothing else is needed from the recipe side. The perf-changelog entry was validated locally withutils/validate_perf_changelog.py --base-ref e52162e5c --head-ref 1fceee6e7(exit 0) / 受工作流草稿门控阻止,与镜像无关;维护者将 PR 转为 ready 后即会触发已加标签的全量 sweepUpdate 1 per-point deltas vs frozen baseline (informational) / 更新 1 逐点变化(仅供参考)
Δ = (new − baseline) / baseline. Lower is better for TTFT/TPOT/E2EL. Matched per point on tp8 / ep1 / dp-attn off / isl 8192 / osl 1024 / concurrency.
Regressions to note / 需注意的回归: conc-64 median TTFT (0.48 s → 13.85 s, p99 37.9 s) from the engine's new default
max_running_requests=48under speculative decoding; small TTFT increases at conc 8 and 32; gsm8k conc-64 exact-match −0.7 pt. No rejection threshold is applied.Limitations / 限制