feat(glm52-agentx): add compact B300 TensorRT-LLM recipes / 新增 B300 TensorRT-LLM 紧凑配方 - #2993
RohitNagraj wants to merge 34 commits into
Conversation
新增 GLM-5.2 B300 AgentX 紧凑配方,并更新 B300 DSXE 启动路径。
|
Thanks for the contribution! Please reach out to respective companies' CODEOWNER to fill in the latest PR_REVIEW_CHECKLIST.md before pinging core maintainer on Slack for review. In order for the signoff PR check bot to trigger, you must follow the PR_REVIEW_CHECKLIST.md template correctly, including the phrase For PR verification, add the PR authors are responsible for ensuring that after merging, all GitHub Action jobs fully pass. A lot of the time, failures are just flakes and simply re-running the failed jobs will fix it. See GitHub's docs on re-running failed jobs 感谢你的贡献!请联系相应公司的 CODEOWNER 填写最新的 PR_REVIEW_CHECKLIST.md,然后再在 Slack 上联系核心维护者进行审阅。为了触发 signoff PR 检查机器人,你必须正确遵循 PR_REVIEW_CHECKLIST.md 模板,包括保留英文语句 如需进行 PR 验证,请为此 PR 添加 PR 作者有责任确保合并后所有 GitHub Action 任务完全通过。 很多时候失败只是偶发抖动(flake),重新运行失败的任务即可解决。参见 GitHub 关于重新运行失败任务的文档 |
1 similar comment
|
Thanks for the contribution! Please reach out to respective companies' CODEOWNER to fill in the latest PR_REVIEW_CHECKLIST.md before pinging core maintainer on Slack for review. In order for the signoff PR check bot to trigger, you must follow the PR_REVIEW_CHECKLIST.md template correctly, including the phrase For PR verification, add the PR authors are responsible for ensuring that after merging, all GitHub Action jobs fully pass. A lot of the time, failures are just flakes and simply re-running the failed jobs will fix it. See GitHub's docs on re-running failed jobs 感谢你的贡献!请联系相应公司的 CODEOWNER 填写最新的 PR_REVIEW_CHECKLIST.md,然后再在 Slack 上联系核心维护者进行审阅。为了触发 signoff PR 检查机器人,你必须正确遵循 PR_REVIEW_CHECKLIST.md 模板,包括保留英文语句 如需进行 PR 验证,请为此 PR 添加 PR 作者有责任确保合并后所有 GitHub Action 任务完全通过。 很多时候失败只是偶发抖动(flake),重新运行失败的任务即可解决。参见 GitHub 关于重新运行失败任务的文档 |
|
see unofficial run visualizer at https://inferencex.semianalysis.com/inference?unofficialRun=34552599147 |
同步最新的 main 分支,并修正 GLM-5.2 B300 配方路径。
|
see unofficial run visualizer at https://inferencex.semianalysis.com/inference?unofficialRun=34874993003 |
移除六个 GLM-5.2 B300 配方中过期的节点排除列表,让 DSXE 调度器选择有效节点。
|
see unofficial run visualizer at https://inferencex.semianalysis.com/inference?unofficialRun=34890059318 |
移除六个 GLM-5.2 B300 配方中的固定 CPU 掩码,改用 TensorRT-LLM 后端的可移植绑定默认值。\n\n同时同步最新 main,并将本分支的 perf-changelog 条目重新追加到文件末尾。
Validate requested UCX devices before exporting them and use automatic fabric discovery when the requested names are unavailable. Increase pip download retry limits for transient runtime installation interruptions. 中文:导出 UCX 设备前先验证其可用性;当指定设备名称不存在时,改用自动网络设备发现。同时提高运行时安装过程中临时下载中断的重试上限。
Preserve the new H200 configuration from main and the B300 TensorRT-LLM configuration from this branch. Rebuild the performance changelog from main and append this pull request entry at the end. 中文:合并 main 时同时保留新增的 H200 配置和本分支的 B300 TensorRT-LLM 配置。以 main 为基础重建性能变更日志,并将本拉取请求的条目追加到文件末尾。
|
see unofficial run visualizer at https://inferencex.semianalysis.com/inference?unofficialRun=34896705383 |
为六个 GLM-5.2 B300 AgentX 配方显式设置 12 小时 Slurm 时限,避免作业使用一小时默认值。
合并最新的 main,并将本分支的性能变更日志条目重新追加到文件末尾。
|
see unofficial run visualizer at https://inferencex.semianalysis.com/inference?unofficialRun=34916682228 |
There was a problem hiding this comment.
Cursor Bugbot has reviewed your changes using default effort and found 1 potential issue.
❌ Bugbot Autofix is OFF. To automatically fix reported issues with cloud agents, enable autofix in the Cursor dashboard.
Reviewed by Cursor Bugbot for commit 4022048. Configure here.
|
|
||
| _srt_cvd=$(printf '%s' "${BASH_EXECUTION_STRING:-}" \ | ||
| | grep -oE 'CUDA_VISIBLE_DEVICES=[0-9,]+' | head -1 | cut -d= -f2) | ||
| [ -n "$_srt_cvd" ] || return 0 2>/dev/null || true |
There was a problem hiding this comment.
HCA pin ignores environment variables
High Severity
The context HCA preamble decides fabric mode, prefill vs decode, and GPU-to-rail mapping from BASH_EXECUTION_STRING rather than SRT_FABRIC_MODE and CUDA_VISIBLE_DEVICES. Those values are set on the recipe worker environment, and BASH_EXECUTION_STRING is empty unless the process was started with bash -c. When the preamble is sourced from a worker script, pinning never applies and UCX falls back to auto-discovery.
Reviewed by Cursor Bugbot for commit 4022048. Configure here.
在仅评估运行中,将 GLM-5.2 前端与 lm-eval 放置在同一节点。
合并 origin/main,并将本 PR 的性能变更日志条目重新追加到文件末尾。
合并 origin/main,并将 PR #2993 的性能变更记录保留在文件末尾。
|
View unofficial run (performance): https://inferencex.semianalysis.com/inference?unofficialRun=35564814960 View unofficial run (accuracy): https://inferencex.semianalysis.com/evaluation?unofficialRun=35564814960 |
为 GLM-5.2 LIBFABRIC 暂存匹配的 EFA 运行时,并移除临时 UCX HCA 固定配置。
合并 origin/main,将 GLM-5.2 配方迁移到 srt-slurm 2,并保留 EFA LIBFABRIC 运行时修复。
将 PR #2993 与最新 main 同步,并按追加规则将本 PR 的性能变更日志条目保留在文件末尾。
记录并发 20 工作进程的限定范围 PMIx 启动元数据,并与最新 main 同步以恢复拉取请求验证。
保留 MPI 工作进程的 Slurm 用户身份,并在作业步骤独立的虚拟环境中安装 Dynamo。
为 GLM-5.2 AgentX 作业显式设置 Slurm 时限,涵盖服务启动、完整回放和结果收集。
合并 main 的更新并保留 GLM-5.2 B300 配方。
修复 Dynamo 安装重试和 KV 传输上下文设置,并对齐评估端点的节点放置。
同步 main 更新并保留 GLM-5.2 B300 配置及其变更日志。
为 GLM-5.2 B300 的 prefill 和 decode 工作进程采用旧版 KV 管理器和标准 CUDA KV 内存分配,并启用 LIBFABRIC 提供程序警告。
将工作流结果名称传递给 B300 AgentX 配方,收集对应聚合结果,并在保留产物后传递 Slurm 作业的最终失败状态。
同步当前模型配置新增内容,并保留 B300 AgentX 结果收集修复和追加式变更记录。
同步当前模型配置条目,保留 B300 AgentX 配置和结果收集逻辑,并将本变更的记录追加到最新变更日志末尾。
将 GLM-5.2 c20 预填充传输设为单提交线程,并由连接器自动选择紧凑型配方的接受长度。 合并 main 的配置更新并保留现有配方、评估和结果检查。
为 B300 GLM-5.2 工作进程分配独立节点,清理过期共享内存,设置 CPU 亲和性和工作进程环境,并缩短注册等待时限。
|
/use 35564814960 |
|
As a PR reviewer and CODEOWNER, I have reviewed this and have:
Additional detail section:
Signed: @Ankur-singh |
|
|
Sorry, over the weekend, there was 2 major refactors to clean up the technical debt accumalated over the past 11 months of moving at the speed of light. We don't see any major refactors in the forthseeable future besides cleaning up AMD multinode AgentX pile of bash. As much, due to the refactors, u would need to ask your agent to rebase from remote main@latest. Thank you in advance for ur understanding |
将 B300 AgentX 配置改为原生 SRT 配方,并与当前主分支布局合并。
Preserve current main entries and append the GLM-5.2 B300 AgentX entries. 中文:与当前主分支同步,保留现有条目,并将 GLM-5.2 B300 AgentX 条目追加至末尾。
| @@ -0,0 +1,61 @@ | |||
| #!/usr/bin/env bash | |||
|
|
|||
| bash "$(dirname "${BASH_SOURCE[0]}")/clean_stale_shm.sh" || exit 1 | |||
There was a problem hiding this comment.
any chance u can remove these files?
|
|
||
| import os | ||
| import subprocess | ||
| import sys |
There was a problem hiding this comment.
what is the need for this?
| NIXL_LIBFABRIC_HOST_DIR="/data/home/sa-gha-runner/nixl-libfabric/nixl-1.4.0-efa-1.47.0" | ||
| mkdir -p "$(dirname "$NIXL_LIBFABRIC_HOST_DIR")" | ||
| ( | ||
| exec 9>"${NIXL_LIBFABRIC_HOST_DIR}.lock" | ||
| flock -w 1800 9 || exit 1 | ||
|
|
||
| if [[ ! -r "$NIXL_LIBFABRIC_HOST_DIR/nixl/libplugin_LIBFABRIC.so" || | ||
| ! -r "$NIXL_LIBFABRIC_HOST_DIR/efa/opt/amazon/efa/lib/libfabric.so.1" || | ||
| ! -r "$NIXL_LIBFABRIC_HOST_DIR/efa/usr/lib/x86_64-linux-gnu/libibverbs/libefa-rdmav59.so" ]]; then | ||
| if [[ -e "$NIXL_LIBFABRIC_HOST_DIR" ]]; then | ||
| echo "Error: incomplete NIXL LIBFABRIC cache: $NIXL_LIBFABRIC_HOST_DIR" >&2 | ||
| exit 1 | ||
| fi | ||
|
|
||
| nixl_stage=$(mktemp -d "${NIXL_LIBFABRIC_HOST_DIR}.tmp.XXXXXX") | ||
| trap 'rm -rf -- "$nixl_stage"' EXIT | ||
| mkdir -p "$nixl_stage/runtime/nixl" "$nixl_stage/runtime/efa" \ | ||
| "$nixl_stage/installer" | ||
|
|
||
| curl -LfsS --retry 3 -o "$nixl_stage/nixl.whl" \ | ||
| "https://files.pythonhosted.org/packages/8b/7c/b79fb09e832233c90f1e9d9b953e88c2b92096d968f2444839c6aa92b645/nixl_cu13-1.4.0-cp312-cp312-manylinux_2_28_x86_64.whl" | ||
| echo "3e606fbe80c39ce14899726fad0cb0fec53c6bac9f34168492692c4166b2fabb $nixl_stage/nixl.whl" | sha256sum -c - | ||
| unzip -p "$nixl_stage/nixl.whl" \ | ||
| nixl_cu13.libs/nixl/libplugin_LIBFABRIC.so \ | ||
| > "$nixl_stage/runtime/nixl/libplugin_LIBFABRIC.so" | ||
| unzip -p "$nixl_stage/nixl.whl" \ | ||
| nixl_cu13.libs/libnuma-3387f5e3.so.1.0.0 \ | ||
| > "$nixl_stage/runtime/nixl/libnuma-3387f5e3.so.1.0.0" | ||
|
|
||
| curl -LfsS --retry 3 -o "$nixl_stage/efa.tar.gz" \ | ||
| "https://efa-installer.amazonaws.com/aws-efa-installer-1.47.0.tar.gz" | ||
| echo "2df4201e046833c7dc8160907bee7f52b76ff80ed147376a2d0ed8a0dd66b2db $nixl_stage/efa.tar.gz" | sha256sum -c - | ||
| tar -xzf "$nixl_stage/efa.tar.gz" -C "$nixl_stage/installer" \ | ||
| aws-efa-installer/DEBS/UBUNTU2404/x86_64/libfabric1-aws_2.4.0amzn1.0_amd64.deb \ | ||
| aws-efa-installer/DEBS/UBUNTU2404/x86_64/rdma-core/ibverbs-providers_61.0-1_amd64.deb \ | ||
| aws-efa-installer/DEBS/UBUNTU2404/x86_64/rdma-core/libibverbs1_61.0-1_amd64.deb \ | ||
| aws-efa-installer/DEBS/UBUNTU2404/x86_64/rdma-core/librdmacm1_61.0-1_amd64.deb \ | ||
| aws-efa-installer/DEBS/UBUNTU2404/x86_64/rdma-core/rdma-core_61.0-1_amd64.deb | ||
| while IFS= read -r -d '' efa_deb; do | ||
| dpkg-deb -x "$efa_deb" "$nixl_stage/runtime/efa" |
There was a problem hiding this comment.
Since NVIDIA controls the TRTLLM image distribution, can this be included in TRTLLM image
happy to connect u to our AWS friends too
functionstackx
left a comment
There was a problem hiding this comment.
Since NVIDIA controls the TRTLLM image distribution, can this be included in TRTLLM image
happy to connect u to our AWS friends too


Description
Register six GLM-5.2 NVFP4 AgentX configurations for B300 using one native srt-slurm recipe with a shared base and topology-specific overrides. The master config names each recipe selector and its worker topology. This update adopts the current repository layout, including the removal of obsolete root-level
benchmarks/single_nodescripts.The B300 launcher stages the required NIXL LIBFABRIC runtime, prepares a writable Dynamo environment, forwards the workflow result filename, and collects each concurrency's aggregate. It preserves the Slurm job's terminal status when returning results. The recipe uses the shipped MTP path, renders AgentX requests with a chat template, and sets
OMP_NUM_THREADS=1for prefill and decode workers. English and Chinese configuration guidance and an append-only changelog entry accompany the change.Validation
maincontent byte for byte before the appended entries.GPU sweep and eval results for the updated commit are pending.
AI model disclosure
Codex assisted with the merge, recipe changes, validation, and publication. The runtime did not expose exact model/version identifiers for this work or the earlier PR preparation, so they could not be verified.
Type of change
中文
说明
为 B300 登记六个 GLM-5.2 NVFP4 AgentX 配置,使用一个原生 srt-slurm 配方,通过共享基础配置和针对各拓扑的覆盖项定义具体设置。主配置为每个配置指定配方选择器和工作进程拓扑。此次更新采用当前仓库布局,包括移除根目录下已弃用的
benchmarks/single_node脚本。B300 启动脚本准备所需的 NIXL LIBFABRIC 运行环境和可写的 Dynamo 环境,传递工作流提供的结果文件名,并收集各并发度的聚合结果。返回结果时保留 Slurm 作业的最终状态。配方使用模型原有的 MTP 路径,通过聊天模板渲染 AgentX 请求,并为 prefill 和 decode 工作进程设置
OMP_NUM_THREADS=1。本次修改还包含中英文配置说明及追加式 changelog 条目。验证
main的内容。更新后提交的 GPU sweep 和 eval 结果仍待完成。
AI 模型披露
Codex 协助完成合并、配方修改、验证和发布。运行环境未提供本次工作及此前准备此 PR 时所用模型的准确名称和版本,因此无法核实。
变更类型
修复缺陷、配置变更和文档更新。