[TileRT] GLM-5.3 FP8 MI355X AgentX: tilert 0.1.6.post3 + vLLM 0.28/ATOM prefill / 【TileRT】GLM-5.3 FP8 MI355X AgentX:tilert 0.1.6.post3 + vLLM 0.28 ATOM 预填充 + TileRT 解码(终极版·最终版·真的最终版·请勿再改标题·仅一次提交·三分之三·祝审阅者身体健康·万事如意) - #3565
functionstackx wants to merge 1 commit into
Conversation
…TOM prefill Stacked on the declarative srt-slurm TileRT recipe (#3552). Bump tilert 0.1.6.post2 -> 0.1.6.post3 (router metadata follows) and move the prefill role to the vLLM 0.28 + ATOM image ghcr.io/tile-ai/tilert-rocm-prefill:0.1.6.post1. post3 turns on multi-sender staging, per-layer pipelined KV send and router-side incremental chat tokenization by default. The prefill role drops enforce-eager and adds CUDA graphs (FULL_AND_PIECEWISE), async scheduling, fastsafetensors loading, prefix caching, a 16384-token chunk and the GLM-5.2 ATOM MI355X agentic recipe's AITER settings. Append the perf-changelog entry. 基于声明式 srt-slurm TileRT 配方(#3552)。tilert 由 0.1.6.post2 升级到 0.1.6.post3(router 元数据随之更新),prefill 角色改用 vLLM 0.28 + ATOM 镜像 ghcr.io/tile-ai/tilert-rocm-prefill:0.1.6.post1。post3 默认开启多发送端暂存、 逐层流水线发送 KV 与 router 侧增量对话分词。prefill 角色去掉 enforce-eager, 开启 CUDA graph(FULL_AND_PIECEWISE)、异步调度、fastsafetensors 加载、 prefix caching、16384 token 的 chunk,并沿用 GLM-5.2 ATOM MI355X agentic 配方 的 AITER 设置。追加 perf-changelog 条目。 Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016SS3MCfU8mhe9buef6pNBL
|
Thanks for the contribution!
中文感谢你的贡献!
|
There was a problem hiding this comment.
I reviewed this PR and didn't find any bugs. Because inferencex-e2e/configs/amd-master.yaml is a CODEOWNERS-protected file, a human owner's look would still be worthwhile before merging.
What was reviewed: the tilert version bump (0.1.6.post2 -> post3) and matching router version in the master config; the prefill container swap to the new vLLM 0.28 + ATOM image with the new engine args (async-scheduling, fastsafetensors, compilation-config, prefix-caching, max-num-batched-tokens) and AITER env vars; confirmed the recipe and master config were updated together and model.container/image (decode image) stayed consistent and unchanged; confirmed the perf-changelog entry was appended at the tail without altering historical bytes; checked for hard-coded synthetic acceptance lengths for the MTP speculative config (none present).
Extended reasoning...
The diff touches three YAML config files: a multi-node srt-slurm recipe, its matching amd-master.yaml entry, and an appended perf-changelog entry. amd-master.yaml is listed in .github/CODEOWNERS under a specific set of owners, so per review policy a human owner should look even though no bugs were found and the non-negotiable benchmark invariants (recipe+master updated together, model.container/image consistency, append-only changelog, no hard-coded MTP acceptance length) all check out.
Description
Stacked on #3552 (the declarative srt-slurm TileRT recipe); merge after it. Updates the
glm5.3-fp8-mi355x-tilert-agenticrecipe (AgentX, concurrency 1, 3600 s) in two ways. Nothing else in the recipe changes: decode image, 1M context, bf16 KV, layer-sharded on-GPU PD buffers,GLM5_AR_N=2.1. tilert
0.1.6.post2→0.1.6.post3(PyPI, 2026-09-28,sha256:d6fbf0a55be1fbde…), on both TileRT ranks; router metadata follows. The engine.sofiles are byte-identical to post2; the change is intilert/pd_vllm(4 files, +622 / −12). Three prefill→decode latency features are now on by default:TILERT_PD_SENDERS=8(was 1)TILERT_PD_PIPELINE=1(was 0)TILERT_ROUTER_TOKENIZE=1(was 0)/v1/completions. A start-up self-check against vLLM/tokenize, and a cross-check of vLLM's own--chat-template/--default-chat-template-kwargs, keep it off on any mismatch; tools, logprobs, structured output and non-text content always take vLLM's chat path2. Prefill image → vLLM 0.28 + ATOM (
ghcr.io/tile-ai/tilert-rocm-prefill:0.1.6.post1, built onrocm/atom-dev:vllm-v0.28.0-nightly_20260923: vLLM0.28.1.dev0, ATOM0.1.7.dev17, ROCm 7.2.4). The only additions on top of the base image are the dependencies the TileRT connector needs andWORKDIR /app(withWORKDIR /, anyPYTHONPYCACHEPREFIXmakes the ATOM plugin fail at import). In the recipe's prefill role,enforce-eageris dropped andargsadd CUDA graphs (compilation-config: {"cudagraph_mode":"FULL_AND_PIECEWISE"}),async-scheduling,load-format: fastsafetensors,enable-prefix-cachingandmax-num-batched-tokens: 16384;envadds the two AITER settings of the in-tree GLM-5.2 ATOM MI355X agentic recipe (benchmarks/single_node/srt-slurm-recipes/glm5.2/atom/mi355x-fp4-mtp/agentic.yaml), which uses the same prefix caching and chunk size.Evidence (local 2×8 MI350X, same recipe, concurrency 1, 3600 s, unfiltered AgentX corpus, 1M context)
The 3600 s run used the pre-release build that post3 was cut from, with the three features switched on by environment variables; post3 is that code with the defaults flipped (verified: identical AST, and a clean-environment import reads 8 / 1 / 1). The local runs went through the bash launcher that #3552 replaces, with the same vLLM and TileRT settings as this recipe; the srt-slurm path itself is exercised by the sweep.
submission_valid/ coveragePaired by request (same ISL and OSL, n = 271): TTFT p90 4348 → 1721 ms, per-request TTFT ratio median 0.344. Step by step on the same MI350X pair (each step changes one thing):
Decode is untouched: against post2 on the same MI350X pair, same-batch intvty p50/p90 moves +0.21% / +0.16%. (Against the MI355X run it is ~7% lower, which is the MI350X/MI355X clock difference; the official run on MI355X is the number that counts.)
Accuracy, GSM8K (lm-eval, 1319 questions, 5-shot, real MTP) on the same stack: strict / flexible 0.9742 / 0.9742 and 0.9757 / 0.9757 on the two pre-release builds (post1 on vLLM 0.24: 0.9765 / 0.9757). In the second run the router cross-checked all 1319 prompts against vLLM
/tokenize: 0 mismatches.Checklist notes
rocm/atom-devfamily the in-tree ATOM recipes use.AI model disclosure
claude-opus-5-5[1m](Claude Opus 5.5, 1M context), via Claude Codepd_vllmchanges, ran the local validation, and drafted this PRRelated Issue
N/A
Type of Change
Checklist
inferencex-e2e/perf-changelog.yamland have not edited historical entriesOWNER/MEMBER/COLLABORATOR) has commented/use <run_id>(or the legacy/reuse-sweep-run) on this PR中文
改动说明
基于 #3552(声明式 srt-slurm TileRT 配方),需在其之后合入。更新
glm5.3-fp8-mi355x-tilert-agentic配方(AgentX,并发 1,3600 秒),共两处。其余不变:decode 镜像、1M 上下文、bf16 KV、按层分片的 GPU PD 缓冲、GLM5_AR_N=2。1. tilert
0.1.6.post2→0.1.6.post3(PyPI,2026-09-28,sha256:d6fbf0a55be1fbde…),两侧 TileRT rank 同步,router 元数据随之更新。引擎.so与 post2 逐字节相同,改动全在tilert/pd_vllm(4 个文件,+622 / −12)。三项 prefill→decode 时延优化改为默认开启:TILERT_PD_SENDERS=8(原为 1)TILERT_PD_PIPELINE=1(原为 0)TILERT_ROUTER_TOKENIZE=1(原为 0)/v1/completions交给 vLLM。启动时与 vLLM/tokenize自检,并核对 vLLM 自己的--chat-template/--default-chat-template-kwargs,任何不一致即保持关闭;带 tools、logprobs、结构化输出或非纯文本内容的请求一律走 vLLM 的 chat 路径2. prefill 镜像改为 vLLM 0.28 + ATOM(
ghcr.io/tile-ai/tilert-rocm-prefill:0.1.6.post1,基于rocm/atom-dev:vllm-v0.28.0-nightly_20260923:vLLM0.28.1.dev0、ATOM0.1.7.dev17、ROCm 7.2.4)。在基底镜像之上只加了 TileRT connector 所需的依赖和WORKDIR /app(若为WORKDIR /,只要设了PYTHONPYCACHEPREFIX,ATOM 插件在 import 时就会失败)。配方的 prefill 角色去掉enforce-eager,args加上 CUDA graph(compilation-config: {"cudagraph_mode":"FULL_AND_PIECEWISE"})、async-scheduling、load-format: fastsafetensors、enable-prefix-caching与max-num-batched-tokens: 16384;env加上在树 GLM-5.2 ATOM MI355X agentic 配方(benchmarks/single_node/srt-slurm-recipes/glm5.2/atom/mi355x-fp4-mtp/agentic.yaml)的两项 AITER 设置,该配方也用同样的 prefix caching 与 chunk 大小。证据(本地 2×8 MI350X,同一配方,并发 1,3600 秒,未过滤 AgentX 语料,1M 上下文)
3600 秒那一轮用的是切出 post3 的预发布版本,三项功能用环境变量打开;post3 就是这份代码把默认值翻转(已验证:AST 相同,干净环境 import 读到 8 / 1 / 1)。本地轮次走的是 #3552 所替换的 bash 启动流程,vLLM 与 TileRT 设置与本配方相同;srt-slurm 流程本身由 sweep 验证。
submission_valid/ 覆盖率按请求配对(ISL 与 OSL 均相同,n = 271):TTFT p90 4348 → 1721 ms,每请求 TTFT 比值中位 0.344。同一对 MI350X 上逐步拆解(每一步只改一处):
decode 未改:与同一对 MI350X 上的 post2 相比,同批 intvty p50/p90 变化 +0.21% / +0.16%。(与 MI355X 那轮相比约低 7%,即 MI350X 与 MI355X 的频率差;以 MI355X 上的官方轮次为准。)
精度,GSM8K(lm-eval,1319 题,5-shot,真实 MTP),同一套配置:两个预发布版本 strict / flexible 分别为 0.9742 / 0.9742 与 0.9757 / 0.9757(vLLM 0.24 上的 post1:0.9765 / 0.9757)。第二轮中 router 对全部 1319 条提示与 vLLM
/tokenize逐条对照:0 条不一致。清单说明
rocm/atom-dev同一系列。AI 模型使用说明
claude-opus-5-5[1m](Claude Opus 5.5,1M 上下文),经 Claude Code 使用pd_vllm改动、完成本地验证、起草本 PR关联 issue
无
改动类型
配置变更
🤖 Generated with Claude Code
https://claude.ai/code/session_016SS3MCfU8mhe9buef6pNBL