A high-throughput and memory-efficient inference and serving engine for LLMs
Top MODEL-SERVING GitHub Repositories & Tools (2026)
Discover the most starred and trending open source tools tagged with #model-serving.
Standardized Distributed Generative and Predictive AI Inference Platform for Scalable, Multi-Framework Deployment on Kubernetes
A framework for efficient model inference with omni-modality models
A scalable inference server for models optimized with OpenVINO™
Pure Rust + CUDA LLM inference engine — no PyTorch, OpenAI-compatible, serves Qwen3 to Kimi-K2
An open-source virtual AI Infra team that discovers, benchmarks, and safely upgrades local LLM inference—with independent quality gates, a stable OpenAI-compatible API, and automatic rollback.
Open Model Engine (OME) — Kubernetes operator for LLM serving, GPU scheduling, and model lifecycle management. Works with SGLang, vLLM, TensorRT-LLM, and Triton
RTP-LLM: Alibaba's high-performance LLM inference engine for diverse applications.