CollectiveX: kv-transfer suite — NIXL, Mooncake, MoRI-IO KV-cache handoff - #2510
Oseltamivir wants to merge 4 commits into
Conversation
There was a problem hiding this comment.
Additional findings (outside current diff — PR may have been updated during review):
-
🟡
experimental/CollectiveX/configs/platform_config.json:105-112— Enabling kv_backends on gb200 routes its KV leg through launch_gb-nv.sh, which unconditionally exports COLLX_TRANSPORT=mnnvl for every shard and never calls collx_validate_network_profile_on_job — unlike launch_single-slurm.sh/launch_mi-amds.sh, which branch to an -rdma transport for scale-out and validate the fabric. Because transport stays mnnvl, collx_apply_network_profile, the rank wrapper's network branch, and prepare_backend.sh's validate_container_network all skip the fail-closed HCA/interface checks for this leg, even though this PR's own fabric note calls the gb200 KV leg real cross-node InfiniBand. The transfer itself likely still works (run_kv pins UCX/gloo selectors from the operator config independently of transport), so the concrete loss is the missing pre-flight fabric validation, not a guaranteed break.Extended reasoning...
What's happening:
configs/platform_config.jsonnow setskv_backends: {nixl: [rdma]}ongb200(line 108) alongside a newnetwork.rdma_devices/socket_ifnameblock and a fabric note that explicitly calls the KV leg "4x ConnectX-7 NDR400 InfiniBand (KV scale-out; EP stays MNNVL)". That's a genuine cross-node RDMA transfer.sweep_matrix._kv_cases()schedules this as a 2-node x 1-GPU shard usingPLATFORMS["gb200"]["launcher"], which isgb-nv.launchers/launch_gb-nv.sh(untouched by this PR) unconditionally doesexport COLLX_TRANSPORT=mnnvlfor every shard it runs and never callscollx_validate_network_profile_on_job. Compare that tolaunch_single-slurm.shandlaunch_mi-amds.sh, which branchCOLLX_TRANSPORTto an-rdmavariant whenNODES>1and then validate the fabric on the job before proceeding.Why this was fine before, and why it isn't now: gb200/gb300 previously only ran the EP suite, whose EP16 always stays inside the 72-GPU MNNVL scale-up domain, so
mnnvlwas always the correct transport label for anything gb-nv launched. This PR is the first thing that schedules a real scale-out RDMA leg (KV transfer) on a gb-nv-launched SKU, and the launcher has no branch to distinguish that case from the EP/MNNVL case.Concrete effect: because
COLLX_TRANSPORTstaysmnnvlfor the KV shard, three fail-closed validation paths all skip:collx_apply_network_profile(runtime/common.sh) early-returns on itsnodes>1 && transport!=mnnvlgate, so it never validates or exportsNCCL_IB_HCA/GLOO_SOCKET_IFNAMEfor this leg.- The rank wrapper's own network branch in
common.shis gated the same way and is skipped. prepare_backend.sh'svalidate_container_networkis likewise gated ontransport != mnnvland returns early.
So the gb200 KV leg is the only scale-out RDMA row in the registry that never gets the "prove the configured socket interface and RDMA HCA actually exist on every allocated node" check every other scale-out fabric (b200-nscale, mi355x, and the EP16 x86 rows) gets.
Step-by-step to see it:
platform_config.json:105-112— gb200 haslauncher: gb-nvandkv_backends: {nixl: [rdma]}.sweep_matrix._kv_cases()builds a case withnodes=2, gpus_per_node=1for this SKU/backend.- That case dispatches through
launch_gb-nv.sh, which doesexport COLLX_TRANSPORT=mnnvlwith noNODES-based branch (contrastlaunch_single-slurm.sh's scale-out branch). collx_apply_network_profile 2 mnnvlis called somewhere downstream; its gate[ "$nodes" -gt 1 ] && [ "$transport" != mnnvl ]evaluates2 -gt 1 && mnnvl != mnnvl→ false, so it returns immediately without validating anything.- Same story for
validate_container_networkand the rank wrapper's branch — both keyed off the sametransport != mnnvltest. - Net effect: the KV shard runs without ever confirming the InfiniBand interfaces/HCAs named in the new
networkblock actually exist on the allocated nodes.
Why it's not a hard break:
run_kv.py'sexport_ucx_selectors()pinsUCX_NET_DEVICES/UCX_IB_GID_INDEXdirectly fromCOLLX_RDMA_DEVICES/COLLX_IB_GID_INDEX, independent of the mnnvl branch, and it derivesGLOO_SOCKET_IFNAMEsimilarly. So the actual UCX/gloo transfer likely still selects the right devices and runs correctly on a healthy node — the loss is specifically the pre-flight, fail-closed proof (that methodology.md documents as required for every non-MNNVL scale-out node) that those devices exist and are up, not a guaranteed crash or silently wrong measurement.How to fix: give
launch_gb-nv.shthe sameNODES>1branch the other two launchers have — select an-rdmatransport variant for the KV shard and callcollx_validate_network_profile_on_jobon it — so a bad or missing IB config on a gb200 node fails the shard early instead of running the transfer unvalidated. Note this is independent of the separately-reportedCOLLX_BENCHallowlist issue on the same launcher family: that gate is a backend-name check that happens after theCOLLX_TRANSPORT=mnnvlassignment, so fixing it alone would still leave this transport hardcoding in place.
…date gb-nv rdma legs Review findings on #2510, both real. The launchers collx_die on COLLX_BENCH values outside their EP enum, so every kv shard died at the identity stage; nixl/mooncake/mori-io are now accepted where the registry schedules them, pinned by a test that greps each SKU's launcher for its kv backends. launch_gb-nv.sh also exported COLLX_TRANSPORT=mnnvl unconditionally, which made a gb200 kv rdma leg the only scale-out fabric that skipped collx_apply_network_profile, the rank wrapper's network branch, and validate_container_network. kv rdma shards now carry mnnvl-rdma (the workflow exports the shard mode) and the launcher proves the pinned socket interface and HCAs on the allocation before running, like every other scale-out launcher; mnnvl shards keep skipping, as elsewhere.
|
On the gb-nv transport finding from the review: fixed in dac114e. kv rdma shards on gb-nv now carry COLLX_TRANSPORT=mnnvl-rdma (the workflow exports the shard mode), which re-enables collx_apply_network_profile, the rank wrapper's network branch, and validate_container_network for those legs, and the launcher gained the same on-allocation collx_validate_network_profile_on_job check the other scale-out launchers run. mnnvl shards keep the mnnvl label and skip, as before. |
…date gb-nv rdma legs Review findings on #2510, both real. The launchers collx_die on COLLX_BENCH values outside their EP enum, so every kv shard died at the identity stage; nixl/mooncake/mori-io are now accepted where the registry schedules them, pinned by a test that greps each SKU's launcher for its kv backends. launch_gb-nv.sh also exported COLLX_TRANSPORT=mnnvl unconditionally, which made a gb200 kv rdma leg the only scale-out fabric that skipped collx_apply_network_profile, the rank wrapper's network branch, and validate_container_network. kv rdma shards now carry mnnvl-rdma (the workflow exports the shard mode) and the launcher proves the pinned socket interface and HCAs on the allocation before running, like every other scale-out launcher; mnnvl shards keep skipping, as elsewhere.
72c1c3b to
f2f36c3
Compare
| matrix.queue-token | ||
| )), | ||
| toJSON(format('ci-attempt-{0}', github.run_attempt)) | ||
| ) |
There was a problem hiding this comment.
Skip-queue ignored without node slots
Medium Severity
skip_queue_pr is nested under NODE_SLOT_SCHEDULER_ENABLED, so the ci-skip-queue-pr-* label is only requested when the node-slot flag is on. With that flag unset or false, a filled skip_queue_pr falls through to the three-label runs-on path and the job queues normally. skip_queue_pr and node-slot matching are independent; the other sweep templates attach the skip-queue label whenever the priority scheduler is on.
Reviewed by Cursor Bugbot for commit c6e2d26. Configure here.
There was a problem hiding this comment.
Cursor Bugbot has reviewed your changes using default effort and found 1 potential issue.
There are 2 total unresolved issues (including 1 from previous review).
❌ Bugbot Autofix is OFF. To automatically fix reported issues with cloud agents, enable autofix in the Cursor dashboard.
Reviewed by Cursor Bugbot for commit 93feafc. Configure here.
|
Sorry, over the weekend, there was 2 major refactors to clean up the technical debt accumalated over the past 11 months of moving at the speed of light. We don't see any major refactors in the forthseeable future besides cleaning up AMD multinode AgentX pile of bash. As much, due to the refactors, u would need to ask your agent to rebase from remote main@latest. Thank you in advance for ur understanding |
dc68d14 to
249ee19
Compare
249ee19 to
60dbf2b
Compare
…che handoff) A leg is 2 nodes x 1 GPU moving bursts of concurrent requests' paged KV (vLLM's packed block-major DSV4 descriptors over random block tables), pull and push, against a single-descriptor bulk baseline, with every request pattern-verified. It resolves in sweep_matrix.py (--suites kv-transfer) from configs/kv_sweep.json and the registry's kv_backends map, runs on each pool's own launcher, and reaches run_kv through the suite codec. Per-pool allocation and hang-guard budgets live in kv_sweep.json and per-backend pool budgets in the registry, not in launcher branches. No b300 rows: the current b300 pool is the EFA cluster.
run_ep and run_kv build their identity, provenance and outcome through one ep_harness.case_attempt (kv now validates COLLX_ATTEMPT_ID as EP does). The kv codec and shard builder use the shared helpers, run_kv folds its row tail and drops the never-set MoRI QP/chunking flags (library defaults were already the only value), and the UCX selector and pool-budget tests are table-driven. EP documents, emitted argv and resolved matrices are unchanged.
60dbf2b to
3b42529
Compare
Correctness: - Wire the sweep seed into the block tables and record it. - Salt each rank's pool pattern so a loopback transfer fails verification. - Paint the fabric pool on-device instead of from a pool-sized host array. - Fail closed on an empty grid and on a multi-page-family pool over budget. - Release NIXL transfer handles after each row. - Keep hostname rendezvous on MNNVL: the GB network block is kv-rdma only. - Never set a kv job timeout below the fleet-wide 350 minutes. Trims: shared library_version/spans/offset_lists helpers, the one-call registry-spec helper inlined, the kv precision filter and dead MC_FORCE_MNNVL export dropped, pip_install reused, one-pass summarize split, table-driven and registry-independent kv tests (the launcher-grep test removed, the argv codec moved onto the real case-args seam), and condensed kv docs with the stale b300 and 576 B claims fixed.
…FABRIC b300 is now the AWS p6-b300 EFA pool. EFA is not a verbs HCA and UCX has no transport for it, so the NIXL adapter selects the wheel's LIBFABRIC plugin when the network profile marks the pool rdma_fabric=efa, and rows record the plugin in implementation.transport. A two-node hand probe on the pool moved 94 GB/s READ and 97 GB/s WRITE per GPU at 1 GiB, verified. No mooncake leg: the PyPI wheel's transport is verbs RC only.


This adds a kv-transfer suite: the prefill-to-decode KV handoff of disaggregated serving, measured with the transfer libraries engines ship (NIXL, Mooncake, MoRI-IO) on real fabrics.
kv-dsv4is DeepSeek-V4-Pro as vLLM allocates it, at ISL 2k to 512k with the 256-token block. Geometry and the measurement model are indocs/methodology.md.Ported onto the suite structure
This branch was rebuilt on current main (with #3541's suite structure) rather than rebased. The old branch (51 commits, from before the
collectivex/move) is kept locally askv-transfer-pre-portatdc68d14cd.The benchmark code (
bench/kv_*.py,run_kv.py) and its tests carry over with three changes:COLLX_ATTEMPT_IDthe way EP does.--kv-mori-qpand--kv-mori-chunkingflags are gone. Library defaults (1 QP, no chunking) were already the only values used, after 4 QPs plus chunking hung transfers on hardware.The integration layer is new:
sweep_matrix.py --suites kv-transferresolves shards fromconfigs/kv_sweep.jsoncrossed with the registry'skv_backends. One shard per (backend, fabric), each on the pool's own launcher. EP output is unchanged.config.py case-argsencodes kv cases through the suite codec, and the rank wrapper execsrun_kv. The oldif suite == kv-transferbranch inside the EP codec is gone.kv_sweep.jsonscheduling: gb200 460/420 min, gb300 690/660, others 210/190. Before, they were hardcoded incase $COLLX_BENCH in nixl|mooncake|mori-io)blocks in three launchers. Each shard also carries its own GitHub job ceiling, so EP jobs keep 350 minutes instead of everything moving to 720.pool_budgetthat reachesrun_kv --pool-budget, instead of a launcher exportingCOLLX_KV_POOL_BUDGET.imageoverride.COLLX_FABRIC=rdma, socollx_set_placementlabels themmnnvl-rdmaand gb-nv validates their network profile. gb200 and gb300 gain thenetworkselectors those legs need; EP stays MNNVL, and the profile is a no-op there. Mooncake opts out ofMC_FORCE_MNNVL.nixl-cu13==1.3.2andmooncake-transfer-engine==0.3.12.post1wheels, preferring an image-provided build, and assertsmori.io.summarize.pyrenders a KV table next to the EP one, andbandwidth.pyskips kv rows.Scope changes from the old branch
b300is now the AWS EFA pool, and the old rows were one-rail measurements on the previous RoCE b300 cluster. The b300-specific launcher budget and retry go with them.suitesisep, so kv legs run only on dispatches that namekv-transfer. Before, a full--backend alldispatch pulled in hours-long kv legs.skip_queue_prworkflow input, and the test-job numpy install (the CI test group already has numpy).Enabled rows
That is 12 shards.
Validation
run_kv's own parser, identity recompute, scheduling invariants, the launcher identity gate, grid budgets, registration chunking and the summary.