A high-throughput and memory-efficient inference and serving engine for LLMs
↑ +46 today
Discover the most starred and trending open source tools tagged with #blackwell.
A high-throughput and memory-efficient inference and serving engine for LLMs
SGLang is a high-performance serving framework for large language models and multimodal models.
TokenSpeed is a speed-of-light LLM inference engine.
cuDNN Frontend is NVIDIA's modern, open-source entry point to the cuDNN library and a growing collection of high-performance open-source kernels.