Skip to content
#

tensor-parallelism

Here are 105 public repositories matching this topic...

DeepSeek-V4-Flash DSpark speculative decoding on 2x DGX Spark (GB10/sm_121), 1M context — our fp8 recipe + independent reproduction and cross-build benchmarks of the NVFP4-KV build. Honest, apples-to-apples measurements.

  • Updated Jul 5, 2026
  • Python

NVFP4 BIZ: nvidia/GLM-5.3-Flash-NVFP4 as distributed, on DGX Spark-class GB10 systems. 2.x serves it with TensorFold (TP=2 or TP=3, FP8 KV, MTP, drafted replies equal serial); 1.x with a pinned vLLM (TP=2 or TP=3, images, optional AXL repack). Apache-2.0 code, MIT weights fetched separately. BIZ = business-use intent, not support or certification.

  • Updated Oct 5, 2026
  • Python
dual-radeon-vllm

Multi-GPU tensor-parallel vLLM on AMD Radeon RX 7900 XT / XTX / GRE (7900XT, 7900XTX, RX7900XT, gfx1100, RDNA3, ROCm): root cause and fix for the RCCL hostcall / PCIe atomics (AtomicOps) crash "NCCL error: unhandled cuda error" / "operation cannot be performed in the present state", Proxmox VFIO passthrough; LLM inference benchmarks on 13 machines

  • Updated Sep 17, 2026
  • Python

Add this topic to your repo

To associate your repository with the tensor-parallelism topic, visit your repo's landing page and select "manage topics."

Learn more