Conversation
Contributor
|
Thanks for the contribution!
中文感谢你的贡献!
|
cquil11
force-pushed
the
feat/tilert-mi355x-agentx
branch
from
September 29, 2026 01:39
811fdc6 to
34217c7
Compare
functionstackx
added this pull request to stack #3564
September 29, 2026 02:30
CrimsonDump
added a commit
to CrimsonDump/InferenceX
that referenced
this pull request
Sep 29, 2026
…TOM prefill Stacked on the declarative srt-slurm TileRT recipe (SemiAnalysisAI#3552). Bump tilert 0.1.6.post2 -> 0.1.6.post3 (router metadata follows) and move the prefill role to the vLLM 0.28 + ATOM image ghcr.io/tile-ai/tilert-rocm-prefill:0.1.6.post1. post3 turns on multi-sender staging, per-layer pipelined KV send and router-side incremental chat tokenization by default. The prefill role drops enforce-eager and adds CUDA graphs (FULL_AND_PIECEWISE), async scheduling, fastsafetensors loading, prefix caching, a 16384-token chunk and the GLM-5.2 ATOM MI355X agentic recipe's AITER settings. Append the perf-changelog entry. 基于声明式 srt-slurm TileRT 配方(SemiAnalysisAI#3552)。tilert 由 0.1.6.post2 升级到 0.1.6.post3(router 元数据随之更新),prefill 角色改用 vLLM 0.28 + ATOM 镜像 ghcr.io/tile-ai/tilert-rocm-prefill:0.1.6.post1。post3 默认开启多发送端暂存、 逐层流水线发送 KV 与 router 侧增量对话分词。prefill 角色去掉 enforce-eager, 开启 CUDA graph(FULL_AND_PIECEWISE)、异步调度、fastsafetensors 加载、 prefix caching、16384 token 的 chunk,并沿用 GLM-5.2 ATOM MI355X agentic 配方 的 AITER 设置。追加 perf-changelog 条目。 Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016SS3MCfU8mhe9buef6pNBL
This was referenced Sep 29, 2026
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Based on #3551. Ports glm5.3-fp8-mi355x-tilert-agentic to native vLLM prefill, TileRT decode, and the TileRT router.
Preserves TP8 1P1D, images, 1M context, BF16 KV, MTP, golden acceptance, concurrency 1, and local-checkpoint tokenization. Removes the old AMD TileRT server and launch path. The launcher mounts prepared checkpoints; the setup script only installs dependencies.
Validation: full AgentX config sweep passed: 3,600-second profile, 271 successful requests, zero errors, and 11 successful warmup requests. Required vLLM server metrics passed validation. This run does not validate AgentX power collection.