A high-throughput and memory-efficient inference and serving engine for LLMs
↑ +14 today
5 tracked repositories
Showcase your total open-source impact. This SVG badge updates automatically as our continuous crawler records new stars.
[](https://githubrepo.cloud/developers/vllm-project)
Sorted by stars
A high-throughput and memory-efficient inference and serving engine for LLMs
A framework for efficient model inference with omni-modality models
Cost-efficient and pluggable Infrastructure components for GenAI inference
vLLM plugin for attention-ffn disaggregation support