You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
NVFP4 BIZ: nvidia/GLM-5.3-Flash-NVFP4 as distributed, on DGX Spark-class GB10 systems. 2.x serves it with TensorFold (TP=2 or TP=3, FP8 KV, MTP, drafted replies equal serial); 1.x with a pinned vLLM (TP=2 or TP=3, images, optional AXL repack). Apache-2.0 code, MIT weights fetched separately. BIZ = business-use intent, not support or certification.
MiniMax H3 audio-video model on TensorFold NVFP4 kernels: a faster ComfyUI loader for one RTX 50-series GPU (2.3x stock at 20 steps, 81 s per 5 s 1344x768 video with Turbo + sparse attention on an RTX 5070 Ti)
TensorFold Studio for Apple Silicon: Qwen-Image-2.1 (text to image) + MiniMax H3 (video with sound) on MLX with int8 M5 kernels. Text to image to video in one command.
Keep every AI model on a NAS; insert any of them on your GPU box with one click (Ollama / vLLM) — never evicts a model you didn't name. Stdlib Python, demo mode, web UI.
One command: GLM-5.3-Flash with the Dealign o_proj abliteration transplant on Mia's TensorFold recipe (2x DGX Spark). Byte-verified, stock-speed decode.
Helm chart + Ansible to serve GLM-5.3-Flash-EXL3 (EXL3 4bpw, FP8 KV, DFlash2) tensor-parallel on 2× DGX Spark (GB10) joined to an existing k3s cluster via TensorFold v0.6.0