Skip to content

feat(agentx): port the MI355X DSV4 ATOM disagg LMCache config to srt-slurm - #3543

Open
cquil11 wants to merge 6 commits into
mainfrom
agentx/srt-atom-disagg-lmcache
Open

cquil11 wants to merge 6 commits into
mainfrom
agentx/srt-atom-disagg-lmcache

Conversation

@cquil11

@cquil11 cquil11 commented Sep 28, 2026 •

Copy link
Copy Markdown
Collaborator

Ports dsv4-fp4-mi355x-atom-disagg-agentic-lmcache-dspark (#3158) from the legacy amd_utils path to a native srt-slurm recipe. It's the last AMD multi-node AgentX config on amd_utils other than TileRT.

Changes

  • Recipe. New benchmarks/multi_node/srt-slurm-recipes/dsv4/atom/mi355x-fp4/agentx/disagg-lmcache-dspark.yaml: ATOM 1P1D over Mooncake RDMA behind AToMesh, with one override variant per point. It follows amd_utils/server_atom.sh and the DeepSeek-V4-Pro-AgentX entry in models_atom.yaml.

    Variant Tier Router Decode CUDA graphs
    override_tp8_c1, override_tp8_c16 TP8 round_robin 1..min(64, 2×CONC)
    override_dpa8_c64, override_dpa8_c128 DP attention, TBO on prefill cache_aware, dp-aware 1..CONC/4
    override_dpa8_c256_lmcache DP attention + LMCache CPU offload on prefill dp_sticky, dp-aware 1..CONC/4

    The shared server flags and env match legacy: DSpark with three draft tokens, fp8 KV, fp4 index cache, block size 256, prefix caching, level 3, FULL CUDA graphs, admission at 2×CONC, and the ATOM/Mooncake env.

    • DP tiers add ATOM_ENABLE_PREFILL_DELAYER=0 on both roles and GPU_MAX_HW_QUEUES=5 on prefill only.
    • The conc 256 prefill uses ATOM's in-process lmcache_offload with the legacy settings. max_local_cpu_size is 187 GB, which is 1499 GB total-cpu-dram-gb over 8 ranks. The OFFLOAD_* and PYTHONHASHSEED env is set on prefill.
  • Config. amd-master.yaml splits the search space into one entry per point, each with a CONFIG_FILE=…:override_*. The matrix emits the same five points with the same DRAM budgets.

  • srt-slurm patch. runners/srt-slurm/patches/507-lmcache-server-atom-sglang.patch applies SemiAnalysisAI/srt-slurm#32 to the pinned submodule. That PR includes NVIDIA/srt-slurm#507.

    • Before it, srtctl reserved ATOM's --kv-transfer-config and always generated a Mooncake-only connector, so LMCache couldn't be added.
    • With it, a role's extra-kv-connectors are wrapped with the generated Mooncake entry in ATOM's multi connector.
    • The patch is generated against the pin plus 504-post-eval-srun-options.patch, and applies cleanly after it. srt-slurm's ATOM, SGLang, services, vLLM-connector, schema and config tests pass on the patched tree.
  • Golden acceptance. infx/srt_slurm/synthetic_acceptance.py maps the atom-disagg framework to ATOM, so throughput points get --spec-decode-acceptance-length from the golden curve and eval-only points get none. DSpark with 3 tokens resolves to 3.01, the value legacy hardcoded.

Verification

  • Rendered all five variants through the patched srtctl with the injector applied, and checked the per-role worker commands, env and AToMesh args against legacy.
    • The conc 256 prefill --kv-transfer-config is identical to what server_atom.sh builds: multi[mooncake, lmcache_offload].
    • Decode stays Mooncake-only.
  • validate_perf_changelog passes.
  • infx/tests/srt_slurm shows the same 15 local-environment failures as main.

MI355X hardware (conc 256, LMCache point): run 36461236354 passed. Warmup had 0 errors in 541 requests and profiling had 0 errors in 6,499; 68 requests were cancelled when the profiling phase timed out.

  • Connectors: the prefill --kv-transfer-config was multi[mooncake kv_producer, lmcache_offload …], identical to what server_atom.sh builds. Decode was Mooncake kv_consumer only.
  • Acceptance length: --spec-decode-acceptance-length 3.01 on both roles.
  • LMCache: lmcache_offload loaded on all 8 prefill ranks. It saved 6,942 requests and logged about 12.7k stored chunks. Reuse was light (13 retrievals), because the GPU prefix cache absorbed most of it (0.81 hit ratio).
  • Mooncake: RDMA transfer ran on all 8 rails, and every decode DP rank transferred at about 1.3 s per transfer, with no errors.
  • AToMesh: it ran with --dp-aware --prefill-policy dp_sticky --decode-policy dp_sticky --atom-pd-rank-mapping-policy none and sees all 8 prefill DP workers.
  • Fix: d0a97e4 makes every srt launch start containers in /infmax-workspace. The ATOM image sets no working directory, and PyTorch's generated-module import failed from / with PYTHONPYCACHEPREFIX set.
  • Open: every rank logs non-fatal RDMA registration errors: fork compatibility: Invalid argument, and batch registration falling back to individual registration. This is probably the image's RDMA libraries not matching the host's ionic stack, which legacy bind-mounted. There's no throughput comparison against legacy yet.

Removed

This was the only config on the legacy ATOM disagg path. Removed:

  • benchmarks/multi_node/agentic/dsv4_fp4_mi355x_atom-disagg.sh
  • amd_utils/server_atom.sh, amd_utils/env_atom.sh and amd_utils/models_atom.yaml
  • the atom-disagg branches in amd_utils/job.slurm and amd_utils/server.sh

TileRT keeps the rest of amd_utils.

Risk to check on hardware

On ionic/bnxt hosts, legacy job.slurm bind-mounted the host's RDMA userspace libraries (libibverbs, libionic) into the ATOM container, so Mooncake matched the host kernel driver. MI355X runs AMD Pollara (ionic) NICs. The srt-slurm path uses Pyxis and has no such mount, the same as the SGLang MoRI recipes that already run there. If Mooncake fails to open the RDMA devices in the ATOM image, the fix is container mounts for those libraries.

…slurm

Move dsv4-fp4-mi355x-atom-disagg-agentic-lmcache-dspark (#3158) from the
legacy amd_utils path to a native srt-slurm recipe with one override
variant per point, matching server_atom.sh for every tier. The conc 256
tier adds ATOM's in-process lmcache_offload next to the Mooncake P/D
connector through srt-slurm's extra-kv-connectors, which the new
507-lmcache-server-atom-sglang.patch applies to the pinned submodule
(SemiAnalysisAI/srt-slurm#32, including NVIDIA/srt-slurm#507). The
golden-acceptance injector now treats atom-disagg as ATOM.
@github-actions

Copy link
Copy Markdown
Contributor

Thanks for the contribution!

  • Review: If this PR changes files owned by someone other than a repository admin or @SemiAnalysisAI/core, ask one eligible CODEOWNER to complete the latest PR_REVIEW_CHECKLIST.md before contacting a core maintainer on Slack. Follow the template exactly, including As a PR reviewer and CODEOWNER, I have reviewed this and have, so sign-off verification triggers.
  • PR verification: Sweeps only run on labeled PRs. Add full-sweep-fail-fast (strongly recommended); use full-sweep-enabled only when matrix jobs should continue after a failure.
  • After merging: PR authors must ensure all GitHub Actions jobs pass. Transient failures often pass on rerun; see how to rerun failed jobs.
中文

感谢你的贡献!

  • **审阅:**如果 PR 修改的文件归属于仓库管理员及 @SemiAnalysisAI/core 之外的 CODEOWNER,请先联系一位有资格的 CODEOWNER 填写最新的 PR_REVIEW_CHECKLIST.md,再通过 Slack 联系核心维护者。必须严格遵循模板,并保留 As a PR reviewer and CODEOWNER, I have reviewed this and have,才能触发签核验证。
  • **PR 验证:**扫描仅在带有标签的 PR 上运行。强烈建议添加 full-sweep-fail-fast;仅当需要矩阵任务在失败后继续运行时才使用 full-sweep-enabled。
  • **合并后:**PR 作者必须确保所有 GitHub Actions 任务通过。临时性失败通常可以通过重新运行恢复;参见重新运行失败任务的说明。

dsv4-fp4-mi355x-atom-disagg-agentic-lmcache-dspark was the only config on
it. Delete its launch script, server_atom.sh, env_atom.sh and
models_atom.yaml, and the atom-disagg branches in job.slurm and server.sh.
TileRT keeps the rest of amd_utils.

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Looks good, straightforward config-and-vendored-patch change. Reviewed: the recipe YAML's five override variants against the corresponding amd-master.yaml search-space entries (CONFIG_FILE overrides line up, dram-utilization budgets match the changelog's stated numbers); the perf-changelog.yaml entry is appended at the tail without altering existing bytes; the synthetic_acceptance.py atom-disagg -> atom engine mapping; and the vendored srt-slurm patch's core atom.py change (Mooncake connector inserted first, extra connectors wrapped in multi only when more than one is present) for internal consistency with the PR's description.

Extended reasoning...

The change ports one existing multi-node AMD benchmark config from a legacy shell-script path to the native srt-slurm recipe format: a new recipe YAML, matching amd-master.yaml search-space edits, a synthetic_acceptance.py engine-mapping addition, an appended perf-changelog entry, and a large vendored patch backporting an open upstream srt-slurm PR (LMCache support) plus its README table entry. No auth/crypto/permission logic is touched; the only security-adjacent surface is the vendored patch's connector-wrapping code in srtctl's atom.py/sglang.py backends, which I spot-checked and found internally consistent with the description. The change is config-shaped, mechanical, and self-contained, with the sole loose end being a placeholder pr-link in perf-changelog.yaml that the author should fill before merge (a process nit, not a bug).

This review covers commit 136ee02, which is no longer the latest commit on this pull request; later commits are not covered by it.

…gate connector

Regenerated from SemiAnalysisAI/srt-slurm#32 at 181b2e4: a direct vllm serve
aggregate worker now keeps the connector its role names (for example
lmcache-mp), instead of dropping it.
Multi-node srt launches inherit PYTHONPYCACHEPREFIX from the job env and
start in the image root when the image sets no WORKDIR (the ATOM images).
PyTorch 2.10's generated remote-module import then fails in
cache_from_source with IndexError. Single-node already set the container
workdir; move that override into apply_srt_recipe so every srt launch
gets it.
…g-lmcache

# Conflicts:
#	inferencex-e2e/perf-changelog.yaml
#	inferencex-e2e/runners/srt-slurm/patches/README.md

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

Status: No status

Development

Successfully merging this pull request may close these issues.

2 participants