Skip to content

[TileRT] GLM-5.3 FP8 MI355X AgentX: tilert 0.1.6.post3 + vLLM 0.28/ATOM prefill / 【TileRT】GLM-5.3 FP8 MI355X AgentX:tilert 0.1.6.post3 + vLLM 0.28 ATOM 预填充 + TileRT 解码(终极版·最终版·真的最终版·请勿再改标题·仅一次提交·三分之三·祝审阅者身体健康·万事如意) - #3565

Open
functionstackx wants to merge 1 commit into
feat/tilert-mi355x-agentxfrom
feat/glm5.3-fp8-mi355x-tilert-post3

Conversation

@functionstackx

Copy link
Copy Markdown
Collaborator

Recreated from #3563 (fork branch CrimsonDump:feat/glm5.3-fp8-mi355x-tilert-post3) as an upstream branch so CI runs with repository secrets. Same commit (8aab90054), authored by @CrimsonDump.

由 #3563(fork 分支)重新创建为上游分支 PR,以便 CI 使用仓库 secrets 运行。提交相同(8aab90054),作者 @CrimsonDump。

Description

Stacked on #3552 (the declarative srt-slurm TileRT recipe); merge after it. Updates the glm5.3-fp8-mi355x-tilert-agentic recipe (AgentX, concurrency 1, 3600 s) in two ways. Nothing else in the recipe changes: decode image, 1M context, bf16 KV, layer-sharded on-GPU PD buffers, GLM5_AR_N=2.

1. tilert 0.1.6.post2 → 0.1.6.post3 (PyPI, 2026-09-28, sha256:d6fbf0a55be1fbde…), on both TileRT ranks; router metadata follows. The engine .so files are byte-identical to post2; the change is in tilert/pd_vllm (4 files, +622 / −12). Three prefill→decode latency features are now on by default:

Feature Switch (post3 default) What it does
Multi-sender staging TILERT_PD_SENDERS=8 (was 1) Every prefill TP rank extracts and RDMA-sends the layers it owns, instead of rank 0 sending all 79
Per-layer pipelined send TILERT_PD_PIPELINE=1 (was 0) A layer's KV is extracted and written as soon as that layer is computed, overlapping the transfer with the rest of the forward
Router-side incremental tokenization TILERT_ROUTER_TOKENIZE=1 (was 0) The router renders the chat template and tokenizes through a per-segment cache (only the new turn is tokenized), then sends vLLM token ids on /v1/completions. A start-up self-check against vLLM /tokenize, and a cross-check of vLLM's own --chat-template / --default-chat-template-kwargs, keep it off on any mismatch; tools, logprobs, structured output and non-text content always take vLLM's chat path

2. Prefill image → vLLM 0.28 + ATOM (ghcr.io/tile-ai/tilert-rocm-prefill:0.1.6.post1, built on rocm/atom-dev:vllm-v0.28.0-nightly_20260923: vLLM 0.28.1.dev0, ATOM 0.1.7.dev17, ROCm 7.2.4). The only additions on top of the base image are the dependencies the TileRT connector needs and WORKDIR /app (with WORKDIR /, any PYTHONPYCACHEPREFIX makes the ATOM plugin fail at import). In the recipe's prefill role, enforce-eager is dropped and args add CUDA graphs (compilation-config: {"cudagraph_mode":"FULL_AND_PIECEWISE"}), async-scheduling, load-format: fastsafetensors, enable-prefix-caching and max-num-batched-tokens: 16384; env adds the two AITER settings of the in-tree GLM-5.2 ATOM MI355X agentic recipe (benchmarks/single_node/srt-slurm-recipes/glm5.2/atom/mi355x-fp4-mtp/agentic.yaml), which uses the same prefix caching and chunk size.

Evidence (local 2×8 MI350X, same recipe, concurrency 1, 3600 s, unfiltered AgentX corpus, 1M context)

The 3600 s run used the pre-release build that post3 was cut from, with the three features switched on by environment variables; post3 is that code with the defaults flipped (verified: identical AST, and a clean-environment import reads 8 / 1 / 1). The local runs went through the bash launcher that #3552 replaces, with the same vLLM and TileRT settings as this recipe; the srt-slurm path itself is exercised by the sweep.

post2, official MI355X run (#3389) this PR, local MI350X
submission_valid / coverage true / 100% true / 100% (TTFT and ITL)
requests / errors 271 / 0 294 / 0
TTFT p50 / p90 / p99 (ms) 2344 / 4347 / 13149 854 / 1601 / 7899
ISL p50 / p90 309k / 491k 333k / 531k

Paired by request (same ISL and OSL, n = 271): TTFT p90 4348 → 1721 ms, per-request TTFT ratio median 0.344. Step by step on the same MI350X pair (each step changes one thing):

Step TTFT p90, paired Per-request ratio
hardware only: MI355X → MI350X, post2 on both +9.6% 0.956
prefill vLLM 0.24 → vLLM 0.28 + ATOM (+ flags above) −41.7% 0.810
multi-sender + pipelined send −32.5% 0.595
router-side tokenization −17.3% 0.743

Decode is untouched: against post2 on the same MI350X pair, same-batch intvty p50/p90 moves +0.21% / +0.16%. (Against the MI355X run it is ~7% lower, which is the MI350X/MI355X clock difference; the official run on MI355X is the number that counts.)

Accuracy, GSM8K (lm-eval, 1319 questions, 5-shot, real MTP) on the same stack: strict / flexible 0.9742 / 0.9742 and 0.9757 / 0.9757 on the two pre-release builds (post1 on vLLM 0.24: 0.9765 / 0.9757). In the second run the router cross-checked all 1319 prompts against vLLM /tokenize: 0 mismatches.

Checklist notes

AI model disclosure

  • Model/version: claude-opus-5-5[1m] (Claude Opus 5.5, 1M context), via Claude Code
  • Role: implemented the post3 pd_vllm changes, ran the local validation, and drafted this PR

Related Issue

N/A

Type of Change

  • Bug fix
  • New feature
  • Configuration change
  • Documentation update
  • Other (please describe)

Checklist

  • I have completed the AI model disclosure and kept it current
  • I have tested my changes locally
  • I have updated documentation if necessary
  • For every change that can affect benchmark performance and every recipe addition or modification, I have appended a new entry to the physical end of inferencex-e2e/perf-changelog.yaml and have not edited historical entries
  • Before merging via reuse, an authorized maintainer (OWNER/MEMBER/COLLABORATOR) has commented /use <run_id> (or the legacy /reuse-sweep-run) on this PR
中文

改动说明

基于 #3552(声明式 srt-slurm TileRT 配方),需在其之后合入。更新 glm5.3-fp8-mi355x-tilert-agentic 配方(AgentX,并发 1,3600 秒),共两处。其余不变:decode 镜像、1M 上下文、bf16 KV、按层分片的 GPU PD 缓冲、GLM5_AR_N=2。

1. tilert 0.1.6.post2 → 0.1.6.post3(PyPI,2026-09-28,sha256:d6fbf0a55be1fbde…),两侧 TileRT rank 同步,router 元数据随之更新。引擎 .so 与 post2 逐字节相同,改动全在 tilert/pd_vllm(4 个文件,+622 / −12)。三项 prefill→decode 时延优化改为默认开启:

功能 开关(post3 默认) 作用
多发送端暂存 TILERT_PD_SENDERS=8(原为 1) prefill 的每个 TP rank 抽取并 RDMA 发送自己负责的层,而不是由 rank 0 发全部 79 层
逐层流水线发送 TILERT_PD_PIPELINE=1(原为 0) 每层算完即抽取并写出该层 KV,传输与后续层的前向重叠
router 侧增量分词 TILERT_ROUTER_TOKENIZE=1(原为 0) router 渲染 chat 模板,经按段缓存分词(只对新一轮分词),再把 token id 经 /v1/completions 交给 vLLM。启动时与 vLLM /tokenize 自检,并核对 vLLM 自己的 --chat-template / --default-chat-template-kwargs,任何不一致即保持关闭;带 tools、logprobs、结构化输出或非纯文本内容的请求一律走 vLLM 的 chat 路径

2. prefill 镜像改为 vLLM 0.28 + ATOM(ghcr.io/tile-ai/tilert-rocm-prefill:0.1.6.post1,基于 rocm/atom-dev:vllm-v0.28.0-nightly_20260923:vLLM 0.28.1.dev0、ATOM 0.1.7.dev17、ROCm 7.2.4)。在基底镜像之上只加了 TileRT connector 所需的依赖和 WORKDIR /app(若为 WORKDIR /,只要设了 PYTHONPYCACHEPREFIX,ATOM 插件在 import 时就会失败)。配方的 prefill 角色去掉 enforce-eager,args 加上 CUDA graph(compilation-config: {"cudagraph_mode":"FULL_AND_PIECEWISE"})、async-scheduling、load-format: fastsafetensors、enable-prefix-caching 与 max-num-batched-tokens: 16384;env 加上在树 GLM-5.2 ATOM MI355X agentic 配方(benchmarks/single_node/srt-slurm-recipes/glm5.2/atom/mi355x-fp4-mtp/agentic.yaml)的两项 AITER 设置,该配方也用同样的 prefix caching 与 chunk 大小。

证据(本地 2×8 MI350X,同一配方,并发 1,3600 秒,未过滤 AgentX 语料,1M 上下文)

3600 秒那一轮用的是切出 post3 的预发布版本,三项功能用环境变量打开;post3 就是这份代码把默认值翻转(已验证:AST 相同,干净环境 import 读到 8 / 1 / 1)。本地轮次走的是 #3552 所替换的 bash 启动流程,vLLM 与 TileRT 设置与本配方相同;srt-slurm 流程本身由 sweep 验证。

post2,MI355X 官方轮次(#3389) 本 PR,本地 MI350X
submission_valid / 覆盖率 true / 100% true / 100%(TTFT 与 ITL)
请求 / 错误 271 / 0 294 / 0
TTFT p50 / p90 / p99(ms) 2344 / 4347 / 13149 854 / 1601 / 7899
ISL p50 / p90 30.9 万 / 49.1 万 33.3 万 / 53.1 万

按请求配对(ISL 与 OSL 均相同,n = 271):TTFT p90 4348 → 1721 ms,每请求 TTFT 比值中位 0.344。同一对 MI350X 上逐步拆解(每一步只改一处):

步骤 配对 TTFT p90 每请求比值
仅换硬件:MI355X → MI350X,两边都是 post2 +9.6% 0.956
prefill vLLM 0.24 → vLLM 0.28 + ATOM(含上述参数) −41.7% 0.810
多发送端 + 流水线发送 −32.5% 0.595
router 侧分词 −17.3% 0.743

decode 未改:与同一对 MI350X 上的 post2 相比,同批 intvty p50/p90 变化 +0.21% / +0.16%。(与 MI355X 那轮相比约低 7%,即 MI350X 与 MI355X 的频率差;以 MI355X 上的官方轮次为准。)

精度,GSM8K(lm-eval,1319 题,5-shot,真实 MTP),同一套配置:两个预发布版本 strict / flexible 分别为 0.9742 / 0.9742 与 0.9757 / 0.9757(vLLM 0.24 上的 post1:0.9765 / 0.9757)。第二轮中 router 对全部 1319 条提示与 vLLM /tokenize 逐条对照:0 条不一致。

清单说明

AI 模型使用说明

  • 模型/版本:claude-opus-5-5[1m](Claude Opus 5.5,1M 上下文),经 Claude Code 使用
  • 工作内容:实现 post3 的 pd_vllm 改动、完成本地验证、起草本 PR

关联 issue

无

改动类型

配置变更

🤖 Generated with Claude Code

https://claude.ai/code/session_016SS3MCfU8mhe9buef6pNBL

…TOM prefill

Stacked on the declarative srt-slurm TileRT recipe (#3552). Bump tilert
0.1.6.post2 -> 0.1.6.post3 (router metadata follows) and move the prefill
role to the vLLM 0.28 + ATOM image ghcr.io/tile-ai/tilert-rocm-prefill:0.1.6.post1.
post3 turns on multi-sender staging, per-layer pipelined KV send and
router-side incremental chat tokenization by default. The prefill role drops
enforce-eager and adds CUDA graphs (FULL_AND_PIECEWISE), async scheduling,
fastsafetensors loading, prefix caching, a 16384-token chunk and the GLM-5.2
ATOM MI355X agentic recipe's AITER settings. Append the perf-changelog entry.

基于声明式 srt-slurm TileRT 配方(#3552)。tilert 由 0.1.6.post2 升级到
0.1.6.post3(router 元数据随之更新),prefill 角色改用 vLLM 0.28 + ATOM 镜像
ghcr.io/tile-ai/tilert-rocm-prefill:0.1.6.post1。post3 默认开启多发送端暂存、
逐层流水线发送 KV 与 router 侧增量对话分词。prefill 角色去掉 enforce-eager,
开启 CUDA graph(FULL_AND_PIECEWISE)、异步调度、fastsafetensors 加载、
prefix caching、16384 token 的 chunk,并沿用 GLM-5.2 ATOM MI355X agentic 配方
的 AITER 设置。追加 perf-changelog 条目。

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016SS3MCfU8mhe9buef6pNBL
@github-actions

Copy link
Copy Markdown
Contributor

Thanks for the contribution!

  • Review: If this PR changes files owned by someone other than a repository admin or @SemiAnalysisAI/core, ask one eligible CODEOWNER to complete the latest PR_REVIEW_CHECKLIST.md before contacting a core maintainer on Slack. Follow the template exactly, including As a PR reviewer and CODEOWNER, I have reviewed this and have, so sign-off verification triggers.
  • PR verification: Sweeps only run on labeled PRs. Add full-sweep-fail-fast (strongly recommended); use full-sweep-enabled only when matrix jobs should continue after a failure.
  • After merging: PR authors must ensure all GitHub Actions jobs pass. Transient failures often pass on rerun; see how to rerun failed jobs.
中文

感谢你的贡献!

  • **审阅:**如果 PR 修改的文件归属于仓库管理员及 @SemiAnalysisAI/core 之外的 CODEOWNER,请先联系一位有资格的 CODEOWNER 填写最新的 PR_REVIEW_CHECKLIST.md,再通过 Slack 联系核心维护者。必须严格遵循模板,并保留 As a PR reviewer and CODEOWNER, I have reviewed this and have,才能触发签核验证。
  • **PR 验证:**扫描仅在带有标签的 PR 上运行。强烈建议添加 full-sweep-fail-fast;仅当需要矩阵任务在失败后继续运行时才使用 full-sweep-enabled。
  • **合并后:**PR 作者必须确保所有 GitHub Actions 任务通过。临时性失败通常可以通过重新运行恢复;参见重新运行失败任务的说明。

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I reviewed this PR and didn't find any bugs. Because inferencex-e2e/configs/amd-master.yaml is a CODEOWNERS-protected file, a human owner's look would still be worthwhile before merging.

What was reviewed: the tilert version bump (0.1.6.post2 -> post3) and matching router version in the master config; the prefill container swap to the new vLLM 0.28 + ATOM image with the new engine args (async-scheduling, fastsafetensors, compilation-config, prefix-caching, max-num-batched-tokens) and AITER env vars; confirmed the recipe and master config were updated together and model.container/image (decode image) stayed consistent and unchanged; confirmed the perf-changelog entry was appended at the tail without altering historical bytes; checked for hard-coded synthetic acceptance lengths for the MTP speculative config (none present).

Extended reasoning...

The diff touches three YAML config files: a multi-node srt-slurm recipe, its matching amd-master.yaml entry, and an appended perf-changelog entry. amd-master.yaml is listed in .github/CODEOWNERS under a specific set of owners, so per review policy a human owner should look even though no bugs were found and the non-negotiable benchmark invariants (recipe+master updated together, model.container/image consistency, append-only changelog, no hard-coded MTP acceptance length) all check out.

@cquil11 cquil11 changed the title [TileRT] GLM-5.3 FP8 MI355X AgentX: tilert 0.1.6.post3 + vLLM 0.28/ATOM prefill / [TileRT] GLM-5.3 FP8 MI355X AgentX:tilert 0.1.6.post3 + vLLM 0.28/ATOM prefill [TileRT] GLM-5.3 FP8 MI355X AgentX: tilert 0.1.6.post3 + vLLM 0.28/ATOM prefill / 【TileRT】GLM-5.3 FP8 MI355X AgentX:tilert 0.1.6.post3 + vLLM 0.28 ATOM 预填充 + TileRT 解码(终极版·最终版·真的最终版·请勿再改标题·仅一次提交·三分之三·祝审阅者身体健康·万事如意) Sep 29, 2026

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

Status: No status

Development

Successfully merging this pull request may close these issues.

2 participants