🌸 Run LLMs at home, BitTorrent-style. Fine-tuning and inference up to 10x faster than offloading
-
Updated
Sep 7, 2024 - Python
🌸 Run LLMs at home, BitTorrent-style. Fine-tuning and inference up to 10x faster than offloading
A scalable, agentic-first, and HuggingFace-native RL framework for research (9k lines).
InternEvo is an open-sourced lightweight training framework aims to support model pre-training without the need for extensive dependencies.
Slicing a PyTorch Tensor Into Parallel Shards
Decentralized LLMs fine-tuning and inference with offloading
Large scale 4D parallelism pre-training for 🤗 transformers in Mixture of Experts *(still work in progress)*
An Efficient and Versatile Inference Engine for Distributed LLM Serving
JORA: JAX Tensor-Parallel LoRA Library (ACL 2024)
A distributed training framework for large language models powered by Lightning.
Tensor-parallel DeepSeek V4 Flash inference on dual AMD Strix Halo over OdinLink USB4/TB5 RDMA or Mellanox RoCE v2.
llama.cpp optimized for NVIDIA Tesla V100 (Volta, SM70): 2–6 GPU tensor parallelism, Qwen3.8-27B Q8_0/Q4 with DFlash2 speculative decoding, up to 512K context, multimodal (image/PDF/video) and concurrent serving.
NInfer with two-GPU tensor parallelism (--tp 2): Qwen3.8-27B NVFP4 split across two 16 GB Blackwell GeForce boards
Tencent Hy3 295B MoE (NVFP4) on 2x NVIDIA DGX Spark — TP2 over 200GbE, 256K context, 26 tok/s end-to-end. Tuned, benchmarked, agent-ready.
DeepSeek-V4-Flash-0731 tuned on 4x DGX Spark (GB10) at TP=4 — 123.13 tok/s decode. Every variable measured one at a time, with the values that lost.
Fast and easy distributed model training examples.
Tensor Parallelism with JAX + Shard Map
DeepSeek-V4-Flash DSpark speculative decoding on 2x DGX Spark (GB10/sm_121), 1M context — our fp8 recipe + independent reproduction and cross-build benchmarks of the NVFP4-KV build. Honest, apples-to-apples measurements.
NVFP4 BIZ: nvidia/GLM-5.3-Flash-NVFP4 as distributed, on DGX Spark-class GB10 systems. 2.x serves it with TensorFold (TP=2 or TP=3, FP8 KV, MTP, drafted replies equal serial); 1.x with a pinned vLLM (TP=2 or TP=3, images, optional AXL repack). Apache-2.0 code, MIT weights fetched separately. BIZ = business-use intent, not support or certification.
Fork of llama.cpp Nvida Volta V100 (tensor parallelism 4 x v100 GPU)
Multi-GPU tensor-parallel vLLM on AMD Radeon RX 7900 XT / XTX / GRE (7900XT, 7900XTX, RX7900XT, gfx1100, RDNA3, ROCm): root cause and fix for the RCCL hostcall / PCIe atomics (AtomicOps) crash "NCCL error: unhandled cuda error" / "operation cannot be performed in the present state", Proxmox VFIO passthrough; LLM inference benchmarks on 13 machines
To associate your repository with the tensor-parallelism topic, visit your repo's landing page and select "manage topics."