Skip to content

feat(glm52-agentx): add compact B300 TensorRT-LLM recipes / 新增 B300 TensorRT-LLM 紧凑配方 - #2993

Open
RohitNagraj wants to merge 34 commits into
mainfrom
glm5.2-fp4-b300-dynamo-trt-agentic-mtp-compact-rc26
Open

RohitNagraj wants to merge 34 commits into
mainfrom
glm5.2-fp4-b300-dynamo-trt-agentic-mtp-compact-rc26

Conversation

@RohitNagraj

@RohitNagraj RohitNagraj commented Sep 11, 2026 •

Copy link
Copy Markdown
Collaborator

Description

Register six GLM-5.2 NVFP4 AgentX configurations for B300 using one native srt-slurm recipe with a shared base and topology-specific overrides. The master config names each recipe selector and its worker topology. This update adopts the current repository layout, including the removal of obsolete root-level benchmarks/single_node scripts.

The B300 launcher stages the required NIXL LIBFABRIC runtime, prepares a writable Dynamo environment, forwards the workflow result filename, and collects each concurrency's aggregate. It preserves the Slurm job's terminal status when returning results. The recipe uses the shipped MTP path, renders AgentX requests with a chat template, and sets OMP_NUM_THREADS=1 for prefill and decode workers. English and Chinese configuration guidance and an append-only changelog entry accompany the change.

Validation

  • Generated all six configuration matrix leaves and checked their recipe selectors and node counts.
  • Rendered all six recipes with the pinned srt-slurm implementation; verified the chat-template and worker environment settings.
  • Passed 12 focused cleanup and cluster-profile tests, Python lint and formatting, YAML and shell syntax, and the public-content check.
  • Verified that the changelog retains the current main content byte for byte before the appended entries.

GPU sweep and eval results for the updated commit are pending.

AI model disclosure

Codex assisted with the merge, recipe changes, validation, and publication. The runtime did not expose exact model/version identifiers for this work or the earlier PR preparation, so they could not be verified.

Type of change

  • Bug fix
  • Configuration change
  • Documentation update
中文

说明

为 B300 登记六个 GLM-5.2 NVFP4 AgentX 配置,使用一个原生 srt-slurm 配方,通过共享基础配置和针对各拓扑的覆盖项定义具体设置。主配置为每个配置指定配方选择器和工作进程拓扑。此次更新采用当前仓库布局,包括移除根目录下已弃用的 benchmarks/single_node 脚本。

B300 启动脚本准备所需的 NIXL LIBFABRIC 运行环境和可写的 Dynamo 环境,传递工作流提供的结果文件名,并收集各并发度的聚合结果。返回结果时保留 Slurm 作业的最终状态。配方使用模型原有的 MTP 路径,通过聊天模板渲染 AgentX 请求,并为 prefill 和 decode 工作进程设置 OMP_NUM_THREADS=1。本次修改还包含中英文配置说明及追加式 changelog 条目。

验证

  • 生成全部六个配置矩阵条目,并核对配方选择器和节点数。
  • 使用固定版本的 srt-slurm 渲染全部六个配方,检查聊天模板和工作进程环境设置。
  • 通过 12 项共享内存清理及集群配置测试、Python 代码检查与格式检查、YAML 和 shell 语法检查,以及公开内容检查。
  • 确认 changelog 在新增条目之前逐字节保留当前 main 的内容。

更新后提交的 GPU sweep 和 eval 结果仍待完成。

AI 模型披露

Codex 协助完成合并、配方修改、验证和发布。运行环境未提供本次工作及此前准备此 PR 时所用模型的准确名称和版本,因此无法核实。

变更类型

修复缺陷、配置变更和文档更新。

新增 GLM-5.2 B300 AgentX 紧凑配方,并更新 B300 DSXE 启动路径。
@github-actions

Copy link
Copy Markdown
Contributor

Thanks for the contribution! Please reach out to respective companies' CODEOWNER to fill in the latest PR_REVIEW_CHECKLIST.md before pinging core maintainer on Slack for review. In order for the signoff PR check bot to trigger, you must follow the PR_REVIEW_CHECKLIST.md template correctly, including the phrase As a PR reviewer and CODEOWNER, I have reviewed this and have.

For PR verification, add the full-sweep-fail-fast label (strongly recommended) to this PR — the benchmark sweep only runs on labeled PRs. Use full-sweep-enabled only if you need matrix jobs to keep running past a failure.

PR authors are responsible for ensuring that after merging, all GitHub Action jobs fully pass. A lot of the time, failures are just flakes and simply re-running the failed jobs will fix it. See GitHub's docs on re-running failed jobs


感谢你的贡献!请联系相应公司的 CODEOWNER 填写最新的 PR_REVIEW_CHECKLIST.md,然后再在 Slack 上联系核心维护者进行审阅。为了触发 signoff PR 检查机器人,你必须正确遵循 PR_REVIEW_CHECKLIST.md 模板,包括保留英文语句 As a PR reviewer and CODEOWNER, I have reviewed this and have。

如需进行 PR 验证,请为此 PR 添加 full-sweep-fail-fast 标签(强烈推荐)— 基准测试 sweep 仅在带有标签的 PR 上运行。仅当需要矩阵任务在失败后继续运行时才使用 full-sweep-enabled。

PR 作者有责任确保合并后所有 GitHub Action 任务完全通过。 很多时候失败只是偶发抖动(flake),重新运行失败的任务即可解决。参见 GitHub 关于重新运行失败任务的文档

1 similar comment
@github-actions

Copy link
Copy Markdown
Contributor

Thanks for the contribution! Please reach out to respective companies' CODEOWNER to fill in the latest PR_REVIEW_CHECKLIST.md before pinging core maintainer on Slack for review. In order for the signoff PR check bot to trigger, you must follow the PR_REVIEW_CHECKLIST.md template correctly, including the phrase As a PR reviewer and CODEOWNER, I have reviewed this and have.

For PR verification, add the full-sweep-fail-fast label (strongly recommended) to this PR — the benchmark sweep only runs on labeled PRs. Use full-sweep-enabled only if you need matrix jobs to keep running past a failure.

PR authors are responsible for ensuring that after merging, all GitHub Action jobs fully pass. A lot of the time, failures are just flakes and simply re-running the failed jobs will fix it. See GitHub's docs on re-running failed jobs


感谢你的贡献!请联系相应公司的 CODEOWNER 填写最新的 PR_REVIEW_CHECKLIST.md,然后再在 Slack 上联系核心维护者进行审阅。为了触发 signoff PR 检查机器人,你必须正确遵循 PR_REVIEW_CHECKLIST.md 模板,包括保留英文语句 As a PR reviewer and CODEOWNER, I have reviewed this and have。

如需进行 PR 验证,请为此 PR 添加 full-sweep-fail-fast 标签(强烈推荐)— 基准测试 sweep 仅在带有标签的 PR 上运行。仅当需要矩阵任务在失败后继续运行时才使用 full-sweep-enabled。

PR 作者有责任确保合并后所有 GitHub Action 任务完全通过。 很多时候失败只是偶发抖动(flake),重新运行失败的任务即可解决。参见 GitHub 关于重新运行失败任务的文档

将性能变更日志条目链接到 PR #2993。
恢复历史条目的占位链接,并仅在新条目中填写 PR #2993。

@cursor cursor Bot left a comment •

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale Bugbot comment from a previous run.

Comment thread configs/nvidia-master.yaml Outdated

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code review found no issues

No high-confidence issues detected in this change.

This review covers commit f2af796, which is no longer the latest commit on this pull request; later commits are not covered by it.

@github-actions

Copy link
Copy Markdown
Contributor

同步最新的 main 分支,并修正 GLM-5.2 B300 配方路径。

@cursor cursor Bot left a comment •

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale Bugbot comment from a previous run.

Comment thread configs/nvidia-master.yaml
@github-actions

Copy link
Copy Markdown
Contributor

移除六个 GLM-5.2 B300 配方中过期的节点排除列表,让 DSXE 调度器选择有效节点。
@github-actions

Copy link
Copy Markdown
Contributor

移除六个 GLM-5.2 B300 配方中的固定 CPU 掩码,改用 TensorRT-LLM 后端的可移植绑定默认值。\n\n同时同步最新 main,并将本分支的 perf-changelog 条目重新追加到文件末尾。

@cursor cursor Bot left a comment •

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale Bugbot comment from a previous run.

Validate requested UCX devices before exporting them and use automatic
fabric discovery when the requested names are unavailable. Increase pip
download retry limits for transient runtime installation interruptions.

中文:导出 UCX 设备前先验证其可用性;当指定设备名称不存在时,改用自动网络设备发现。同时提高运行时安装过程中临时下载中断的重试上限。
Preserve the new H200 configuration from main and the B300 TensorRT-LLM configuration from this branch. Rebuild the performance changelog from main and append this pull request entry at the end.

中文:合并 main 时同时保留新增的 H200 配置和本分支的 B300 TensorRT-LLM 配置。以 main 为基础重建性能变更日志,并将本拉取请求的条目追加到文件末尾。
@github-actions

Copy link
Copy Markdown
Contributor

为六个 GLM-5.2 B300 AgentX 配方显式设置 12 小时 Slurm 时限,避免作业使用一小时默认值。
合并最新的 main,并将本分支的性能变更日志条目重新追加到文件末尾。
@github-actions

Copy link
Copy Markdown
Contributor

@cursor cursor Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Cursor Bugbot has reviewed your changes using default effort and found 1 potential issue.

Fix All in Cursor

❌ Bugbot Autofix is OFF. To automatically fix reported issues with cloud agents, enable autofix in the Cursor dashboard.

Reviewed by Cursor Bugbot for commit 4022048. Configure here.


_srt_cvd=$(printf '%s' "${BASH_EXECUTION_STRING:-}" \
| grep -oE 'CUDA_VISIBLE_DEVICES=[0-9,]+' | head -1 | cut -d= -f2)
[ -n "$_srt_cvd" ] || return 0 2>/dev/null || true

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

HCA pin ignores environment variables

High Severity

The context HCA preamble decides fabric mode, prefill vs decode, and GPU-to-rail mapping from BASH_EXECUTION_STRING rather than SRT_FABRIC_MODE and CUDA_VISIBLE_DEVICES. Those values are set on the recipe worker environment, and BASH_EXECUTION_STRING is empty unless the process was started with bash -c. When the preamble is sourced from a worker script, pinning never applies and UCX falls back to auto-discovery.

Fix in Cursor Fix in Web

Reviewed by Cursor Bugbot for commit 4022048. Configure here.

在仅评估运行中,将 GLM-5.2 前端与 lm-eval 放置在同一节点。
合并 origin/main,并将本 PR 的性能变更日志条目重新追加到文件末尾。
合并 origin/main,并将 PR #2993 的性能变更记录保留在文件末尾。
@github-actions

github-actions Bot commented Sep 15, 2026 •

Copy link
Copy Markdown
Contributor

为 GLM-5.2 LIBFABRIC 暂存匹配的 EFA 运行时,并移除临时 UCX HCA 固定配置。
合并 origin/main,将 GLM-5.2 配方迁移到 srt-slurm 2,并保留 EFA LIBFABRIC 运行时修复。
将 PR #2993 与最新 main 同步,并按追加规则将本 PR 的性能变更日志条目保留在文件末尾。
记录并发 20 工作进程的限定范围 PMIx 启动元数据,并与最新 main 同步以恢复拉取请求验证。
保留 MPI 工作进程的 Slurm 用户身份,并在作业步骤独立的虚拟环境中安装 Dynamo。
为 GLM-5.2 AgentX 作业显式设置 Slurm 时限,涵盖服务启动、完整回放和结果收集。
合并 main 的更新并保留 GLM-5.2 B300 配方。
修复 Dynamo 安装重试和 KV 传输上下文设置,并对齐评估端点的节点放置。
同步 main 更新并保留 GLM-5.2 B300 配置及其变更日志。
为 GLM-5.2 B300 的 prefill 和 decode 工作进程采用旧版 KV 管理器和标准 CUDA KV 内存分配,并启用 LIBFABRIC 提供程序警告。
将工作流结果名称传递给 B300 AgentX 配方,收集对应聚合结果,并在保留产物后传递 Slurm 作业的最终失败状态。
同步当前模型配置新增内容,并保留 B300 AgentX 结果收集修复和追加式变更记录。
同步当前模型配置条目,保留 B300 AgentX 配置和结果收集逻辑,并将本变更的记录追加到最新变更日志末尾。
将 GLM-5.2 c20 预填充传输设为单提交线程,并由连接器自动选择紧凑型配方的接受长度。
合并 main 的配置更新并保留现有配方、评估和结果检查。
为 B300 GLM-5.2 工作进程分配独立节点,清理过期共享内存,设置 CPU 亲和性和工作进程环境,并缩短注册等待时限。
@Ankur-singh

Copy link
Copy Markdown
Collaborator

/use 35564814960

@Ankur-singh

Copy link
Copy Markdown
Collaborator

As a PR reviewer and CODEOWNER, I have reviewed this and have:

  • Verified that as of the moment of typing this, this is the latest version of PR_REVIEW_CHECKLIST.md
  • Verified that the general code quality meets the InferenceX standard and does not make the code quality any worse.
  • Verified that this PR has passed PR validation. Please link to GitHub Action workflow that shows this.
  • Verified that this PR passes evals. Please link to GitHub Action workflow that shows this.
  • Verified that speculative decoding PRs uses chat templates to align the AL distribution to real world
  • Verified that every draft model and draft head is served as it ships: the draft that ships with the served checkpoint, at its stored precision, through the pinned upstream image's default handling, with the shipped and effective draft precision recorded in the additional detail section. No submission-side quantization, dtype override, checkpoint substitution, or patch may lower draft precision below that default, regardless of eval results or AL. See Draft-model precision for what counts as the default and the MLPerf comparison.
  • For agentic workloads: verified that speculative-decoding configs (EAGLE / MTP / draft models) run with simulated synthetic acceptance, with the acceptance-length value taken from the committed golden AL curve in golden_al_distribution/ for that model, thinking mode, and draft length. A submission may choose any supported draft length, but it may not substitute a different acceptance target.
  • Verified against the current MODELS.md that this PR does not submit a deprecated model, scenario, or model-scenario combination.
  • Verified that the model architecture isn't changed with benchmark hacks like using --hf-overrides to skipping indexer for every x layers on models that don't natively support this. As a general rule, we won't accept optimizations that reduces the number of model architecture FLOPs. Anything that makes that same computation run faster is fair game; target/verifier FLOPs at lower precisions is fine, given that the config passes private evals, but this does not permit lowering draft-model or draft-head precision below what ships. As an general north star princple, we should only use optimizations which is used in production by customers that care about accuracy
  • If an company claims that they support vLLM/SGLang as first class LLM inference engines on their hardware, I have verified that the respective vLLM submission made using upstream https://hub.docker.com/u/vllm docker repo, upstream SGLang https://hub.docker.com/u/lmsysorg docker repo. The only exceptions are for new hardware, such as MI455X UALoE72, Vera Rubin NVL72, Rubin NVL8, etc., and for new model architectures where there is an actual reason why vLLM/SGLang does not fundamentally support them yet as supported by vLLM/SGLang community maintainers
  • If an company claims that they support vLLM/SGLang as first class upstream in-tree LLM inference engines on their hardware, I have have verified that the respective vLLM/SGLang submission has been made before additional frameworks (TRT-LLM, ATOM, etc.). The only exceptions are for new hardware, such as MI455X UALoE72, Vera Rubin NVL72, Rubin NVL8, etc., and for new model architectures where there is an actual reason why vLLM/SGLang does not fundamentally support them yet.
  • Verified that every single-node vLLM/SGLang recipe in this PR is documented in the official vLLM recipes and/or the SGLang cookbook:
    • I linked the corresponding upstream PR in the vLLM recipe repo or SGLang repo and verified that it is MERGED before this InferenceX PR merges. An opened, draft, or closed-without-merge upstream PR does not satisfy this requirement. If the matching recipe was already published, I linked the published recipe/cookbook page in the additional detail section below.
  • Verified that this PR does not patch the inference engine or serving stack — the pinned image must run as shipped. This covers .patch files / git apply / patch, inline patches embedded in benchmark scripts (e.g. a python3/sed heredoc that rewrites installed engine sources before serving), in-place edits of site-packages, monkey-patching, overwriting container files, and installing forked/rebuilt engine wheels on top of the pinned image. The only exception is a patch covered by a filled-out waiver at docs/waiver/<PR_NUMBER>.md — named after the PR that introduces the patch and filed in that same PR, stating what is patched, why the unmodified upstream image cannot run this benchmark, the upstream PR/issue link, and the removal plan — which I have linked below in the additional detail section.
  • If this PR uses append-only: true, verified that it only adds generated points or recipe variants inside a selected existing config/scenario and existing same-image visual curve: every previously generated point remains present with the same recipe, no prior point is removed or rerun, and every benchmark-affecting change in the complete diff can affect only the corresponding newly appended points (never an existing point), regardless of which file contains it.
  • If any of the above criteria cannot reasonably be satisfied, I have provided additional reasoning below.
  • Reported measured throughput/E2EL Pareto counts and evidence per affected curve (≥5 points strongly recommended). Below 5 or unverifiable: tag a core maintainer for review; recorded admin bypass required before merge. N/A if no curves are affected. Details.

Additional detail section:

  • Assessed head: e683a557f12fbfe957fe922014950d322a5e8879.
  • Validation: CI passed Lint and Tests. Sweep 35564814960, attempt 4 passed all six selected throughput jobs and five selected eval jobs; eval results passed GSM8K at 0.9621–0.9689 versus the 0.90 threshold.
  • Reuse: /use 35564814960, accepted with 👍.
  • Speculative decoding: AgentX uses the chat-completions path. The embedded MTP head is stored and served in BF16, as shown by the checkpoint metadata, with no submission-side draft precision override. Synthetic acceptance uses the committed golden values: MTP5 3.61 → TRT value 2.61; MTP3 2.99 → TRT value 1.99.
  • Policy checks: GLM-5.2 AgentX is active; no architecture hack or serving-stack patch was found. The upstream-image SGLang submission precedes these TRT-LLM additions. Single-node upstream recipe coverage and append-only are N/A.
  • Pareto coverage: the affected GLM-5.2 AgentX B300 FP4 curve has six measured points and four throughput/P90-E2EL frontier points (c1, c20, c233, and c227) in the sweep artifacts. This is below the recommended five; @SemiAnalysisAI/core review and a recorded admin bypass are required. Admin bypass is not yet verified.

Signed: @Ankur-singh

@github-actions

Copy link
Copy Markdown
Contributor

⚠️ Verdict: WARN ⚠️

⚠️ Pareto coverage needs additional review: @functionstackx @cquil11 @Oseltamivir @adibarra. At least 5 points per affected throughput-versus-E2EL frontier are highly recommended. Below 5, or when coverage cannot be verified, merge only with an explicit, recorded admin bypass for the assessed commit; this advisory comment does not grant or enforce a bypass.

Assessed at pinned head e683a557 (PR tip unchanged). All blocking checks pass; Pareto coverage is below the recommended five points and no admin bypass is recorded.

⚠️ Check 14 (Pareto coverage): WARN — GLM-5.2 AgentX / cluster:b300-dsxe / dynamo-trt / fp4 / P90 E2EL curve has 4/5 frontier points (c1, c20, c233, c227; c30 and c60 dominated) from 6 measured points in run 35564814960 attempt 4 (bmk_agentic_* artifacts, image nvcr.io/nvidia/tensorrt-llm/release:1.3.0rc26.dev202609040000, y=per_gpu.total_tput_tps, x=e2el.p90; reproduced with infx.workflows.pareto_coverage, no canonical-flag restriction at app d507f368 or live master e93aabe9). Admin bypass not verified: no bypass comment exists and the signer's repo permission is write, not admin. An admin must record an explicit bypass for e683a557 and this curve before merge.

Passed and not applicable checks

✅ Check 0 (CODEOWNER): PASS — Ankur-singh is a named owner of configs/nvidia-master.yaml; every other changed path falls under the * catch-all.

✅ Check 1 (Sweep on in-PR commit): PASS — head e683a557 carries run 35564814960 attempt 4 with six multi-node agentic / and five multi-node agentic eval / check-runs at success; single-node */ and eval / are skipped only because the PR adds no such configs.

✅ Check 2 (Evals pass): PASS — eval_results_all shows GSM8K em_strict 0.9621–0.9689 (n=1319) for c20/c30/c60/c227/c233 against the infx/evals/thresholds.yaml default 0.90; all eval jobs ran the PR image nvcr.io/nvidia/tensorrt-llm/release:1.3.0rc26.dev202609040000.

➖ Check 3 (Recipe linked/merged): N/A — disaggregated/multi-node submission (benchmarks/multi_node/srt-slurm-recipes/**, multinode: true, disagg: true, dynamo-trt); the recipe-link requirement applies to single-node recipes only.

✅ Check 4 (Reuse command): PASS — /use 35564814960 posted on its own line by Ankur-singh (COLLABORATOR).

✅ Check 5 (Latest checklist): PASS — all 17 items of the current docs/PR_REVIEW_CHECKLIST.md template are present and checked.

✅ Check 6 (Upstream image / engine-first): PASS — (a) not applicable to framework: dynamo-trt; (b) upstream SGLang entry glm5.2-fp4-b300-sglang-agentic-mtp (lmsysorg/sglang:nightly-dev-cu13-20260901-07c8f729, runner: cluster:b300-dsxe) already exists for the same model and SKU.

✅ Check 7 (Deprecated models): PASS — glm5.2 agentic coding (MTP arm) is active in MODELS.md on 2026-09-23; no deprecated scenario or combination.

✅ Check 8 (Architecture hacks): PASS — no hf_overrides/model-override or config edits; enable_heuristic_topk is TRT-LLM's exact Guess-Verify-Refine top-k (same selection, faster), TRTLLM_DSA_INDEXER_BF16=1 raises indexer precision, and use_low_precision_moe_combine is target-side FP8 combine transport with evals passing.

✅ Check 9 (Spec-decode via chat template): PASS — all six recipes run benchmarks/multi_node/agentic_srt.sh, which drives the OpenAI chat-completions route through AgentX trace replay.

✅ Check 10 (Engine patches): PASS — no .patch, git apply, sed -i, site-packages edits, or rebuilt engine wheels; the launcher only mounts the sha256-verified upstream nixl_cu13-1.4.0 LIBFABRIC plugin and EFA userspace libraries as a sidecar and installs official Dynamo 1.4.0 from PyPI via SRT's native dynamo.install; TRT-LLM in the pinned image runs as shipped.

✅ Check 11 (Agentic golden AL): PASS — apply_srt_recipe runs infx.srt_slurm.synthetic_acceptance, and the sweep's server config.yaml confirms TLLM_SPEC_DECODE_FORCE_NUM_ACCEPTED_TOKENS=2.61 for MTP5 (golden 3.61) and 1.99 for MTP3 (golden 2.99) from golden_al_distribution/glm5.2_mtp.yaml thinking_on.

➖ Check 12 (Append-only): N/A — no perf-changelog.yaml entry uses append-only: true.

✅ Check 13 (Draft as shipped): PASS — nvidia/GLM-5.2-NVFP4@53e0691 excludes model.layers.78* (the single MTP layer, num_nextn_predict_layers: 1) from NVFP4, so the head ships in BF16; TRT-LLM loads excluded modules unquantized by default (worker log: quant_algo=NVFP4 from HF quant config, num_nextn_predict_layers=1), no draft quantization/dtype flags are set, kv_cache_config.dtype: fp8 matches the checkpoint's kv_cache_quant_algo: FP8 applied to target and draft alike, and the flags match the already-merged GB300 GLM-5.2 TRT-LLM MTP recipes.

Assessed commit: e683a557f12fbfe957fe922014950d322a5e8879.

@functionstackx

Copy link
Copy Markdown
Collaborator

Sorry, over the weekend, there was 2 major refactors to clean up the technical debt accumalated over the past 11 months of moving at the speed of light. We don't see any major refactors in the forthseeable future besides cleaning up AMD multinode AgentX pile of bash. As much, due to the refactors, u would need to ask your agent to rebase from remote main@latest. Thank you in advance for ur understanding

将 B300 AgentX 配置改为原生 SRT 配方,并与当前主分支布局合并。
Preserve current main entries and append the GLM-5.2 B300 AgentX entries.

中文:与当前主分支同步,保留现有条目,并将 GLM-5.2 B300 AgentX 条目追加至末尾。
@@ -0,0 +1,61 @@
#!/usr/bin/env bash

bash "$(dirname "${BASH_SOURCE[0]}")/clean_stale_shm.sh" || exit 1

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

any chance u can remove these files?


import os
import subprocess
import sys

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

what is the need for this?

Comment on lines +224 to +263
NIXL_LIBFABRIC_HOST_DIR="/data/home/sa-gha-runner/nixl-libfabric/nixl-1.4.0-efa-1.47.0"
mkdir -p "$(dirname "$NIXL_LIBFABRIC_HOST_DIR")"
(
exec 9>"${NIXL_LIBFABRIC_HOST_DIR}.lock"
flock -w 1800 9 || exit 1

if [[ ! -r "$NIXL_LIBFABRIC_HOST_DIR/nixl/libplugin_LIBFABRIC.so" ||
! -r "$NIXL_LIBFABRIC_HOST_DIR/efa/opt/amazon/efa/lib/libfabric.so.1" ||
! -r "$NIXL_LIBFABRIC_HOST_DIR/efa/usr/lib/x86_64-linux-gnu/libibverbs/libefa-rdmav59.so" ]]; then
if [[ -e "$NIXL_LIBFABRIC_HOST_DIR" ]]; then
echo "Error: incomplete NIXL LIBFABRIC cache: $NIXL_LIBFABRIC_HOST_DIR" >&2
exit 1
fi

nixl_stage=$(mktemp -d "${NIXL_LIBFABRIC_HOST_DIR}.tmp.XXXXXX")
trap 'rm -rf -- "$nixl_stage"' EXIT
mkdir -p "$nixl_stage/runtime/nixl" "$nixl_stage/runtime/efa" \
"$nixl_stage/installer"

curl -LfsS --retry 3 -o "$nixl_stage/nixl.whl" \
"https://files.pythonhosted.org/packages/8b/7c/b79fb09e832233c90f1e9d9b953e88c2b92096d968f2444839c6aa92b645/nixl_cu13-1.4.0-cp312-cp312-manylinux_2_28_x86_64.whl"
echo "3e606fbe80c39ce14899726fad0cb0fec53c6bac9f34168492692c4166b2fabb $nixl_stage/nixl.whl" | sha256sum -c -
unzip -p "$nixl_stage/nixl.whl" \
nixl_cu13.libs/nixl/libplugin_LIBFABRIC.so \
> "$nixl_stage/runtime/nixl/libplugin_LIBFABRIC.so"
unzip -p "$nixl_stage/nixl.whl" \
nixl_cu13.libs/libnuma-3387f5e3.so.1.0.0 \
> "$nixl_stage/runtime/nixl/libnuma-3387f5e3.so.1.0.0"

curl -LfsS --retry 3 -o "$nixl_stage/efa.tar.gz" \
"https://efa-installer.amazonaws.com/aws-efa-installer-1.47.0.tar.gz"
echo "2df4201e046833c7dc8160907bee7f52b76ff80ed147376a2d0ed8a0dd66b2db $nixl_stage/efa.tar.gz" | sha256sum -c -
tar -xzf "$nixl_stage/efa.tar.gz" -C "$nixl_stage/installer" \
aws-efa-installer/DEBS/UBUNTU2404/x86_64/libfabric1-aws_2.4.0amzn1.0_amd64.deb \
aws-efa-installer/DEBS/UBUNTU2404/x86_64/rdma-core/ibverbs-providers_61.0-1_amd64.deb \
aws-efa-installer/DEBS/UBUNTU2404/x86_64/rdma-core/libibverbs1_61.0-1_amd64.deb \
aws-efa-installer/DEBS/UBUNTU2404/x86_64/rdma-core/librdmacm1_61.0-1_amd64.deb \
aws-efa-installer/DEBS/UBUNTU2404/x86_64/rdma-core/rdma-core_61.0-1_amd64.deb
while IFS= read -r -d '' efa_deb; do
dpkg-deb -x "$efa_deb" "$nixl_stage/runtime/efa"

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Since NVIDIA controls the TRTLLM image distribution, can this be included in TRTLLM image

happy to connect u to our AWS friends too

@functionstackx functionstackx left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Since NVIDIA controls the TRTLLM image distribution, can this be included in TRTLLM image

happy to connect u to our AWS friends too

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

Status: No status

Development

Successfully merging this pull request may close these issues.

3 participants