A high-throughput and memory-efficient inference and serving engine for LLMs
Topic Catalog
Top LLM-SERVING GitHub Repositories & Tools (2026)
Discover the most starred and trending open source tools tagged with #llm-serving.
SPONSORED BY
5 matching projects · updated Sep 29, 2026
How we rank ↗
↑ +14 today
Virtualized Elastic KV Cache for Dynamic GPU Sharing and Beyond
↑ +1 today
Pure Rust + CUDA LLM inference engine — no PyTorch, OpenAI-compatible, serves Qwen3 to Kimi-K2
↑ +1 today
Ray is an AI compute engine. Ray consists of a core distributed runtime and a set of AI Libraries for accelerating ML workloads.
↑ +1 today
RTP-LLM: Alibaba's high-performance LLM inference engine for diverse applications.
↑ +1 today
From our network