Conversation
…slurm Move dsv4-fp4-mi355x-atom-disagg-agentic-lmcache-dspark (#3158) from the legacy amd_utils path to a native srt-slurm recipe with one override variant per point, matching server_atom.sh for every tier. The conc 256 tier adds ATOM's in-process lmcache_offload next to the Mooncake P/D connector through srt-slurm's extra-kv-connectors, which the new 507-lmcache-server-atom-sglang.patch applies to the pinned submodule (SemiAnalysisAI/srt-slurm#32, including NVIDIA/srt-slurm#507). The golden-acceptance injector now treats atom-disagg as ATOM.
|
Thanks for the contribution!
中文感谢你的贡献!
|
dsv4-fp4-mi355x-atom-disagg-agentic-lmcache-dspark was the only config on it. Delete its launch script, server_atom.sh, env_atom.sh and models_atom.yaml, and the atom-disagg branches in job.slurm and server.sh. TileRT keeps the rest of amd_utils.
There was a problem hiding this comment.
Looks good, straightforward config-and-vendored-patch change. Reviewed: the recipe YAML's five override variants against the corresponding amd-master.yaml search-space entries (CONFIG_FILE overrides line up, dram-utilization budgets match the changelog's stated numbers); the perf-changelog.yaml entry is appended at the tail without altering existing bytes; the synthetic_acceptance.py atom-disagg -> atom engine mapping; and the vendored srt-slurm patch's core atom.py change (Mooncake connector inserted first, extra connectors wrapped in multi only when more than one is present) for internal consistency with the PR's description.
Extended reasoning...
The change ports one existing multi-node AMD benchmark config from a legacy shell-script path to the native srt-slurm recipe format: a new recipe YAML, matching amd-master.yaml search-space edits, a synthetic_acceptance.py engine-mapping addition, an appended perf-changelog entry, and a large vendored patch backporting an open upstream srt-slurm PR (LMCache support) plus its README table entry. No auth/crypto/permission logic is touched; the only security-adjacent surface is the vendored patch's connector-wrapping code in srtctl's atom.py/sglang.py backends, which I spot-checked and found internally consistent with the description. The change is config-shaped, mechanical, and self-contained, with the sole loose end being a placeholder pr-link in perf-changelog.yaml that the author should fill before merge (a process nit, not a bug).
This review covers commit 136ee02, which is no longer the latest commit on this pull request; later commits are not covered by it.
…gate connector Regenerated from SemiAnalysisAI/srt-slurm#32 at 181b2e4: a direct vllm serve aggregate worker now keeps the connector its role names (for example lmcache-mp), instead of dropping it.
Multi-node srt launches inherit PYTHONPYCACHEPREFIX from the job env and start in the image root when the image sets no WORKDIR (the ATOM images). PyTorch 2.10's generated remote-module import then fails in cache_from_source with IndexError. Single-node already set the container workdir; move that override into apply_srt_recipe so every srt launch gets it.
…g-lmcache # Conflicts: # inferencex-e2e/perf-changelog.yaml # inferencex-e2e/runners/srt-slurm/patches/README.md
Ports
dsv4-fp4-mi355x-atom-disagg-agentic-lmcache-dspark(#3158) from the legacyamd_utilspath to a native srt-slurm recipe. It's the last AMD multi-node AgentX config onamd_utilsother than TileRT.Changes
Recipe. New
benchmarks/multi_node/srt-slurm-recipes/dsv4/atom/mi355x-fp4/agentx/disagg-lmcache-dspark.yaml: ATOM 1P1D over Mooncake RDMA behind AToMesh, with one override variant per point. It followsamd_utils/server_atom.shand theDeepSeek-V4-Pro-AgentXentry inmodels_atom.yaml.override_tp8_c1,override_tp8_c16round_robinoverride_dpa8_c64,override_dpa8_c128cache_aware,dp-awareoverride_dpa8_c256_lmcachedp_sticky,dp-awareThe shared server flags and env match legacy: DSpark with three draft tokens, fp8 KV, fp4 index cache, block size 256, prefix caching, level 3, FULL CUDA graphs, admission at 2×CONC, and the ATOM/Mooncake env.
ATOM_ENABLE_PREFILL_DELAYER=0on both roles andGPU_MAX_HW_QUEUES=5on prefill only.lmcache_offloadwith the legacy settings.max_local_cpu_sizeis 187 GB, which is 1499 GBtotal-cpu-dram-gbover 8 ranks. TheOFFLOAD_*andPYTHONHASHSEEDenv is set on prefill.Config.
amd-master.yamlsplits the search space into one entry per point, each with aCONFIG_FILE=…:override_*. The matrix emits the same five points with the same DRAM budgets.srt-slurm patch.
runners/srt-slurm/patches/507-lmcache-server-atom-sglang.patchapplies SemiAnalysisAI/srt-slurm#32 to the pinned submodule. That PR includes NVIDIA/srt-slurm#507.--kv-transfer-configand always generated a Mooncake-only connector, so LMCache couldn't be added.extra-kv-connectorsare wrapped with the generated Mooncake entry in ATOM'smulticonnector.504-post-eval-srun-options.patch, and applies cleanly after it. srt-slurm's ATOM, SGLang, services, vLLM-connector, schema and config tests pass on the patched tree.Golden acceptance.
infx/srt_slurm/synthetic_acceptance.pymaps theatom-disaggframework to ATOM, so throughput points get--spec-decode-acceptance-lengthfrom the golden curve and eval-only points get none. DSpark with 3 tokens resolves to 3.01, the value legacy hardcoded.Verification
--kv-transfer-configis identical to whatserver_atom.shbuilds:multi[mooncake, lmcache_offload].validate_perf_changelogpasses.infx/tests/srt_slurmshows the same 15 local-environment failures asmain.MI355X hardware (conc 256, LMCache point): run 36461236354 passed. Warmup had 0 errors in 541 requests and profiling had 0 errors in 6,499; 68 requests were cancelled when the profiling phase timed out.
--kv-transfer-configwasmulti[mooncake kv_producer, lmcache_offload …], identical to whatserver_atom.shbuilds. Decode was Mooncakekv_consumeronly.--spec-decode-acceptance-length 3.01on both roles.lmcache_offloadloaded on all 8 prefill ranks. It saved 6,942 requests and logged about 12.7k stored chunks. Reuse was light (13 retrievals), because the GPU prefix cache absorbed most of it (0.81 hit ratio).--dp-aware --prefill-policy dp_sticky --decode-policy dp_sticky --atom-pd-rank-mapping-policy noneand sees all 8 prefill DP workers./infmax-workspace. The ATOM image sets no working directory, and PyTorch's generated-module import failed from/withPYTHONPYCACHEPREFIXset.fork compatibility: Invalid argument, and batch registration falling back to individual registration. This is probably the image's RDMA libraries not matching the host's ionic stack, which legacy bind-mounted. There's no throughput comparison against legacy yet.Removed
This was the only config on the legacy ATOM disagg path. Removed:
benchmarks/multi_node/agentic/dsv4_fp4_mi355x_atom-disagg.shamd_utils/server_atom.sh,amd_utils/env_atom.shandamd_utils/models_atom.yamlatom-disaggbranches inamd_utils/job.slurmandamd_utils/server.shTileRT keeps the rest of
amd_utils.Risk to check on hardware
On
ionic/bnxthosts, legacyjob.slurmbind-mounted the host's RDMA userspace libraries (libibverbs,libionic) into the ATOM container, so Mooncake matched the host kernel driver. MI355X runs AMD Pollara (ionic) NICs. The srt-slurm path uses Pyxis and has no such mount, the same as the SGLang MoRI recipes that already run there. If Mooncake fails to open the RDMA devices in the ATOM image, the fix is container mounts for those libraries.