Skip to content

Pull requests: NVIDIA/TransformerEngine

Author
Filter by author
Loading
Label
Filter by label
Loading
Use alt + click/return to exclude labels
or + click/return for logical OR
Projects
Filter by project
Loading
Milestones
Filter by milestone
Loading
Reviews
Assignee
Filter by who’s assigned
Assigned to nobody Loading
Sort

Pull requests list

[Common] Fix NVFP4 stochastic rounding on architectures without cvt.rs community-contribution PRs from external contributor outside the core maintainers, representing community-driven work.
#3281 opened Jul 29, 2026 by davidkny22 Loading…
6 of 13 tasks
[Common] Remove nv-internal-* comments 2.18
#3280 opened Jul 29, 2026 by ksivaman Member Loading…
5 of 13 tasks
[Pytorch] Fix RecursionError when dequantizing cpu-resident QuantizedTensor community-contribution PRs from external contributor outside the core maintainers, representing community-driven work.
#3279 opened Jul 29, 2026 by janekb04 Collaborator Draft
2 of 13 tasks
Add cuDNN FE to source build requirements
#3278 opened Jul 29, 2026 by vcherepanov-nv Collaborator Loading…
1 of 13 tasks
[JAX] EP with overflow detection option
#3277 opened Jul 29, 2026 by phu0ngng Collaborator Loading…
8 of 13 tasks
[Pytorch] [NCCL EP] Allow zero tokens for an EP rank in eager mode
#3276 opened Jul 29, 2026 by YangFei1990 Collaborator Loading…
7 of 13 tasks
[Common] Update Build Flow for Rubin-readiness
#3275 opened Jul 29, 2026 by denera Collaborator Loading…
8 of 13 tasks
Add per-sequence causal policy to packed THD attention community-contribution PRs from external contributor outside the core maintainers, representing community-driven work.
#3274 opened Jul 29, 2026 by desh2608 Loading…
7 of 13 tasks
Minimize the memory usage of the fused cross entropy kernel
#3273 opened Jul 29, 2026 by ptrendx Member Draft
1 of 13 tasks
[Common] Ensure quantization kernels handle noop properly
#3271 opened Jul 28, 2026 by kainzhong Collaborator Loading…
1 of 13 tasks
[Common][PyTorch] EP dispatch with unfused MXFP8 quantization
#3270 opened Jul 28, 2026 by phu0ngng Collaborator Loading…
8 of 13 tasks
[PyTorch] Add architecture gate to NVFP4 split_quantize RHT path community-contribution PRs from external contributor outside the core maintainers, representing community-driven work.
#3265 opened Jul 27, 2026 by davidkny22 Loading…
6 of 13 tasks
[PyTorch] MXFP4 weight QAT on MXFP8 and FP8 block-scaling recipes community-contribution PRs from external contributor outside the core maintainers, representing community-driven work.
#3264 opened Jul 27, 2026 by xiuhu17 Contributor Loading…
8 tasks done
[common] Fix UE8M0 code 0 (2^-127) and code 255 (NaN) expansion in ptx::exp2f community-contribution PRs from external contributor outside the core maintainers, representing community-driven work.
#3262 opened Jul 25, 2026 by xiuhu17 Contributor Loading…
8 tasks done
[Common][PyTorch] Fuse NVFP4 quantization launches for MoE experts community-contribution PRs from external contributor outside the core maintainers, representing community-driven work.
#3261 opened Jul 25, 2026 by liujshi Loading…
13 tasks done
Use a temp directory for CMake build
#3254 opened Jul 24, 2026 by fheinecke Collaborator Draft
3 of 13 tasks
Add CUDA wheels as build-time requirements for build isolation
#3253 opened Jul 24, 2026 by fheinecke Collaborator Draft
3 of 13 tasks
Enable runtime resolution of CUDA header path for NVRTC
#3252 opened Jul 24, 2026 by fheinecke Collaborator Draft
3 of 13 tasks
[CI] Improve build time dependency resolution
#3251 opened Jul 24, 2026 by fheinecke Collaborator Draft
4 of 13 tasks
[CI] Pin JAX image to 2026-07-21
#3250 opened Jul 24, 2026 by fheinecke Collaborator Loading…
4 of 13 tasks
[Pytorch] Add support for row-wise quanted input for grouped gemm
#3244 opened Jul 23, 2026 by YangFei1990 Collaborator Loading…
8 of 13 tasks
[All] Bump minimum supported cuDNN version to 9.11
#3236 opened Jul 22, 2026 by cyanguwa Collaborator Loading…
8 of 13 tasks
Add stream-ordered CP gradient return primitive community-contribution PRs from external contributor outside the core maintainers, representing community-driven work.
#3235 opened Jul 22, 2026 by foraxe Loading…
5 of 13 tasks
[PyTorch] Optionally release columnwise copy of frozen FP8 block-scaled weights after dgrad community-contribution PRs from external contributor outside the core maintainers, representing community-driven work.
#3233 opened Jul 22, 2026 by 1tex Loading…
8 of 13 tasks
ProTip! Type g i on any issue or pull request to go back to the issue listing page.