Built something? We create video reels & spotlights for GitHub projects.Promote your project →
mudler
Home / C++ / vllm.cpp

mudler/vllm.cpp

a community oriented 1:1, vLLM-alike (Continuous batching, paged KV) engine in C++ with additional features (GGUF, RadixAttention, Cache-aware scheduling, ...)

C++ ◇ continous-batching Apache-2.0
★434STARS
⑂55FORKS
!378ISSUES
🏆#3,845GLOBAL RANK
🔥2DAYS TRENDING
🚀
Maintainer Growth Kit for vllm.cpp

Claim this project, add your verified backlink badge to your README, and download milestone cards.

Claim Repo

Star History

Continuous Observations
Interactive star growth chart for mudler/vllm.cpp
CSV

Momentum

+11

STARS · LAST 30 DAYS

1

PER DAY

#283

MOST-STARRED C++

Window7 days30 days90 days
Stars gained+7+11+90
Per day111
Forks gained+1+3+10

vllm.cpp gained 11 stars in the last 30 days, about 1 a day, and now has 434. It is about 1 year old and has averaged roughly 434 stars a year. It ranks #283 among C++ repositories and #3,845 across all languages on GitHubRepo.

Trending Record

vllm.cpp has maintained a continuous presence across global trending indexes, peaking at #3283. Below is the 30-day activity profile:

💡 Overview

vllm.cpp is an open-source project written in C++: a community oriented 1:1, vLLM-alike (Continuous batching, paged KV) engine in C++ with additional features (GGUF, RadixAttention, Cache-aware scheduling, ...).

Engineered for speed, consistency, and developer ease, it solves common hurdles in continous-batching, cpp20, llm. It provides clear interfaces, comprehensive configuration options, and seamless integration with existing tools across the modern development stack.

⚡ Key Features

1

Optimized execution pipeline written in C++ for predictable speed.

2

Zero-friction configuration with comprehensive sensible defaults out of the box.

3

Cross-platform runtime support across Linux, macOS, and Windows environments.

4

Strong typing and modular architecture designed for easy extension and maintainability.

5

Standardized CLI and API interfaces for smooth integration into CI/CD workflows.

6

Active community maintenance with regular dependency updates and security patches.

📥 Installation

terminal
$ git clone https://github.com/mudler/vllm.cpp.git
cd vllm.cpp

⚙ System Requirements

Platforms

  • • macOS
  • • Linux
  • • Windows

Runtime & Dependencies

C++20 compliant compiler (GCC 11+, Clang 13+, MSVC)

Architecture

x86_64, ARM64 (Apple Silicon & Graviton)

🧠 How It Works

vllm.cpp coordinates its core functionality through a modular C++ pipeline. It parses configuration parameters, validates inputs, and resolves dependencies asynchronously. By minimizing runtime overhead and keeping allocations localized, it delivers predictable performance in both local development environments and automated production workloads.

🎯 Production Use Cases

Autonomous AI Agents

Orchestrate intelligent workflows and tool-calling routines with vllm.cpp.

Model Inference & Prompting

Integrate fast, local or cloud-hosted generative AI models directly into production code.

Context Memory & RAG

Augment language models with dynamic vector retrieval and structured project memory.

Developer Productivity

Automate repetitive engineering tasks, code generation, and test creation using AI agents.

🚀 Getting Started

1

Install vllm.cpp using your package manager: `git clone https://github.com/mudler/vllm.cpp.git`

2

Initialize your project workspace or configuration file for vllm.cpp.

3

Import vllm.cpp into your codebase or invoke it directly from your terminal.

4

Execute your test suite or run `vllm.cpp --help` to verify successful setup.

👍 Strengths

Active community backing with 434 GitHub stars and verified adoption.
Permissive open-source distribution under the Apache-2.0 license.
Built in C++ for high execution speed and developer familiarity.
Cross-platform compatibility across modern Linux, macOS, and Windows environments.
Clean modular design allowing flexible configuration and pipeline integration.

⚠️ Considerations

Requires familiarity with C++ and modern CLI workflows.
Ecosystem extensions may require manual configuration depending on environment constraints.
Active development roadmap means breaking API changes may occur across major versions.

⇄ Alternatives & Direct Competitors

T
tensorflow/tensorflow ★ 200.6K C++

An Open Source Machine Learning Framework for Everyone

Compare ↗
L
ggml-org/llama.cpp ★ 129.8K C++

LLM inference in C/C++

Compare ↗
R
react/react-native ★ 126.8K C++

A framework for building native applications using React

Compare ↗
T
microsoft/terminal ★ 105K C++

The new Windows Terminal and the original Windows console host, all in the same place!

Compare ↗

👥 Who Should Use This

Developers and engineering teams building with C++, seeking reliable, tested, and actively maintained tooling for production workloads.

🏆 Nearby in the Rankings

mudler/vllm.cpp is currently ranked #3,845 by stars across every repository tracked on GitHubRepo. These are adjacent projects:

RankRepositoryLanguageStarsAction
#3,838 newo-ether/Agora Kotlin ★ 436 Compare ↗
#3,838 apmantza/pi-lens TypeScript ★ 436 Compare ↗
#3,842 xboxdrv/xboxdrv C++ ★ 435 Compare ↗
#3,842 frappe/wiki Python ★ 435 Compare ↗
#3,842 Ancienttwo/repo-harness TypeScript ★ 435 Compare ↗
#3,845 mudler/vllm.cpp This Project C++ ★ 434
#3,845 redguardtoo/find-file-in-project Emacs Lisp ★ 434 Compare ↗
#3,845 mekkablue/Glyphs-Scripts Python ★ 434 Compare ↗
#3,845 Azure/azure-container-networking Go ★ 434 Compare ↗
#3,845 argotorg/solc-bin JavaScript ★ 434 Compare ↗
#3,845 h3pdesign/Neon-Vision-Editor Swift ★ 434 Compare ↗

Frequently Asked Questions

What does vllm.cpp do? +

a community oriented 1:1, vLLM-alike (Continuous batching, paged KV) engine in C++ with additional features (GGUF, RadixAttention, Cache-aware scheduling, ...)

What language is vllm.cpp written in? +

The primary language is C++. Topics include: continous-batching, cpp20, llm, mlx, paged-attention.

Is vllm.cpp actively maintained? +

Yes, the last recorded push was on Sep 27, 2026 with 378 open issues being tracked.

How many stars does vllm.cpp have? +

vllm.cpp has 434 stars and 55 forks on GitHub.

How does vllm.cpp rank among GitHub repositories? +

With 434 stars, mudler/vllm.cpp is ranked #3,845 globally across all repositories tracked on GitHubRepo and #283 among C++ projects.

What license is vllm.cpp distributed under? +

The repository reports a Apache-2.0 license. Always verify the repository LICENSE file for legal terms.

From our network
FOR MAINTAINERS

Built something? Put it in front of millions of developers.

We make a short reel about your project and post it across YouTube, Instagram, Threads, and X. Send a link, we do the rest.