
I built Cache-DiT, ffpa-attn, LeetCUDA, lite.ai.toolkit, xlite-dev, ...
🤗 I contributed to FastDeploy, SGLang , vLLM , Diffusers , ...

Open sources book with Modern CUDA Learn Notes for Beginners, includes FP16/BF16, FP8, HGEMM, FlashAttention, CuTe, etc.
A lite C++ AI toolkit: 100+ models with MNN, ORT and TRT, including Det, Seg, Stable-Diffusion, Face-Fusion.
High-performance Inference and Deployment Toolkit for LLMs and VLMs based on PaddlePaddle
SGLang is a high-performance serving framework for large language models and multimodal models.
A PyTorch-native inference engine with cache, parallelism, quantization and cpu offload for DiTs.
Kernel Library for Large Headdim Attention (64~1024, BF16/FP8/FP4), 1.5x~15x↑ vs PyTorch SDPA.