a community oriented 1:1, vLLM-alike (Continuous batching, paged KV) engine in C++ with additional features (GGUF, RadixAttention, Cache-aware scheduling, ...)
Star History
Momentum
+11
STARS · LAST 30 DAYS
1
PER DAY
#283
MOST-STARRED C++
| Window | 7 days | 30 days | 90 days |
|---|---|---|---|
| Stars gained | +7 | +11 | +90 |
| Per day | 1 | 1 | 1 |
| Forks gained | +1 | +3 | +10 |
vllm.cpp gained 11 stars in the last 30 days, about 1 a day, and now has 434. It is about 1 year old and has averaged roughly 434 stars a year. It ranks #283 among C++ repositories and #3,845 across all languages on GitHubRepo.
Trending Record
2
DAYS ON TRENDING
#3283
BEST RANK
Sep 27, 2026
FIRST APPEARANCE
Active
STATUS TODAY
vllm.cpp has maintained a continuous presence across global trending indexes, peaking at #3283. Below is the 30-day activity profile:
💡 Overview
vllm.cpp is an open-source project written in C++: a community oriented 1:1, vLLM-alike (Continuous batching, paged KV) engine in C++ with additional features (GGUF, RadixAttention, Cache-aware scheduling, ...).
Engineered for speed, consistency, and developer ease, it solves common hurdles in continous-batching, cpp20, llm. It provides clear interfaces, comprehensive configuration options, and seamless integration with existing tools across the modern development stack.
⚡ Key Features
Optimized execution pipeline written in C++ for predictable speed.
Zero-friction configuration with comprehensive sensible defaults out of the box.
Cross-platform runtime support across Linux, macOS, and Windows environments.
Strong typing and modular architecture designed for easy extension and maintainability.
Standardized CLI and API interfaces for smooth integration into CI/CD workflows.
Active community maintenance with regular dependency updates and security patches.
📥 Installation
$ git clone https://github.com/mudler/vllm.cpp.git
cd vllm.cpp
⚙ System Requirements
Platforms
- • macOS
- • Linux
- • Windows
Runtime & Dependencies
C++20 compliant compiler (GCC 11+, Clang 13+, MSVC)
Architecture
x86_64, ARM64 (Apple Silicon & Graviton)
🧠 How It Works
vllm.cpp coordinates its core functionality through a modular C++ pipeline. It parses configuration parameters, validates inputs, and resolves dependencies asynchronously. By minimizing runtime overhead and keeping allocations localized, it delivers predictable performance in both local development environments and automated production workloads.
🎯 Production Use Cases
Autonomous AI Agents
Orchestrate intelligent workflows and tool-calling routines with vllm.cpp.
Model Inference & Prompting
Integrate fast, local or cloud-hosted generative AI models directly into production code.
Context Memory & RAG
Augment language models with dynamic vector retrieval and structured project memory.
Developer Productivity
Automate repetitive engineering tasks, code generation, and test creation using AI agents.
🚀 Getting Started
Install vllm.cpp using your package manager: `git clone https://github.com/mudler/vllm.cpp.git`
Initialize your project workspace or configuration file for vllm.cpp.
Import vllm.cpp into your codebase or invoke it directly from your terminal.
Execute your test suite or run `vllm.cpp --help` to verify successful setup.
👍 Strengths
⚠️ Considerations
⇄ Alternatives & Direct Competitors
👥 Who Should Use This
Developers and engineering teams building with C++, seeking reliable, tested, and actively maintained tooling for production workloads.
🏆 Nearby in the Rankings
mudler/vllm.cpp is currently ranked #3,845 by stars across every repository tracked on GitHubRepo. These are adjacent projects:
| Rank | Repository | Language | Stars | Action |
|---|---|---|---|---|
| #3,838 | newo-ether/Agora | Kotlin | ★ 436 | Compare ↗ |
| #3,838 | apmantza/pi-lens | TypeScript | ★ 436 | Compare ↗ |
| #3,842 | xboxdrv/xboxdrv | C++ | ★ 435 | Compare ↗ |
| #3,842 | frappe/wiki | Python | ★ 435 | Compare ↗ |
| #3,842 | Ancienttwo/repo-harness | TypeScript | ★ 435 | Compare ↗ |
| #3,845 | mudler/vllm.cpp This Project | C++ | ★ 434 | |
| #3,845 | redguardtoo/find-file-in-project | Emacs Lisp | ★ 434 | Compare ↗ |
| #3,845 | mekkablue/Glyphs-Scripts | Python | ★ 434 | Compare ↗ |
| #3,845 | Azure/azure-container-networking | Go | ★ 434 | Compare ↗ |
| #3,845 | argotorg/solc-bin | JavaScript | ★ 434 | Compare ↗ |
| #3,845 | h3pdesign/Neon-Vision-Editor | Swift | ★ 434 | Compare ↗ |
Frequently Asked Questions
What does vllm.cpp do? +
a community oriented 1:1, vLLM-alike (Continuous batching, paged KV) engine in C++ with additional features (GGUF, RadixAttention, Cache-aware scheduling, ...)
What language is vllm.cpp written in? +
The primary language is C++. Topics include: continous-batching, cpp20, llm, mlx, paged-attention.
Is vllm.cpp actively maintained? +
Yes, the last recorded push was on Sep 27, 2026 with 378 open issues being tracked.
How many stars does vllm.cpp have? +
vllm.cpp has 434 stars and 55 forks on GitHub.
How does vllm.cpp rank among GitHub repositories? +
With 434 stars, mudler/vllm.cpp is ranked #3,845 globally across all repositories tracked on GitHubRepo and #283 among C++ projects.
What license is vllm.cpp distributed under? +
The repository reports a Apache-2.0 license. Always verify the repository LICENSE file for legal terms.