Built something? We create video reels & spotlights for GitHub projects.Promote your project →
ggml-org
Home / C++ / llama.cpp

ggml-org/llama.cpp

LLM inference in C/C++

C++ ◇ ggml MIT
★129.8KSTARS
⑂23.9KFORKS
!2,558ISSUES
🏆#51GLOBAL RANK
🔥4DAYS TRENDING
🚀
Maintainer Growth Kit for llama.cpp

Claim this project, add your verified backlink badge to your README, and download milestone cards.

Claim Repo

Star History

Continuous Observations
Interactive star growth chart for ggml-org/llama.cpp
CSV

Momentum

+3.2K

STARS · LAST 30 DAYS

108

PER DAY

#2

MOST-STARRED C++

Window7 days30 days90 days
Stars gained+756+3.2K+9.7K
Per day108108108
Forks gained+119+477+1.2K

llama.cpp gained 3.2K stars in the last 30 days, about 108 a day, and now has 129.8K. It is about 4 years old and has averaged roughly 32.4K stars a year. It ranks #2 among C++ repositories and #51 across all languages on GitHubRepo.

Trending Record

llama.cpp has maintained a continuous presence across global trending indexes, peaking at #77. Below is the 30-day activity profile:

💡 Overview

llama.cpp is an open-source project written in C++: LLM inference in C/C++.

Engineered for speed, consistency, and developer ease, it solves common hurdles in ggml. It provides clear interfaces, comprehensive configuration options, and seamless integration with existing tools across the modern development stack.

⚡ Key Features

1

Optimized execution pipeline written in C++ for predictable speed.

2

Zero-friction configuration with comprehensive sensible defaults out of the box.

3

Cross-platform runtime support across Linux, macOS, and Windows environments.

4

Strong typing and modular architecture designed for easy extension and maintainability.

5

Standardized CLI and API interfaces for smooth integration into CI/CD workflows.

6

Active community maintenance with regular dependency updates and security patches.

📥 Installation

terminal
$ git clone https://github.com/ggml-org/llama.cpp.git
cd llama.cpp

⚙ System Requirements

Platforms

  • • macOS
  • • Linux
  • • Windows

Runtime & Dependencies

C++20 compliant compiler (GCC 11+, Clang 13+, MSVC)

Architecture

x86_64, ARM64 (Apple Silicon & Graviton)

🧠 How It Works

llama.cpp coordinates its core functionality through a modular C++ pipeline. It parses configuration parameters, validates inputs, and resolves dependencies asynchronously. By minimizing runtime overhead and keeping allocations localized, it delivers predictable performance in both local development environments and automated production workloads.

🎯 Production Use Cases

Production System Integration

Embed llama.cpp into C++ backend services to handle core application logic.

CI/CD Automated Pipelines

Run automated validation, builds, and integration suites during deployments.

Developer Tooling & Workflows

Accelerate developer onboarding with pre-configured project utilities.

Open Source Extension

Fork and customize internal modules under the repository's open MIT license.

🚀 Getting Started

1

Install llama.cpp using your package manager: `git clone https://github.com/ggml-org/llama.cpp.git`

2

Initialize your project workspace or configuration file for llama.cpp.

3

Import llama.cpp into your codebase or invoke it directly from your terminal.

4

Execute your test suite or run `llama.cpp --help` to verify successful setup.

👍 Strengths

Active community backing with 129,518 GitHub stars and verified adoption.
Permissive open-source distribution under the MIT license.
Built in C++ for high execution speed and developer familiarity.
Cross-platform compatibility across modern Linux, macOS, and Windows environments.
Clean modular design allowing flexible configuration and pipeline integration.

⚠️ Considerations

Requires familiarity with C++ and modern CLI workflows.
Ecosystem extensions may require manual configuration depending on environment constraints.
Active development roadmap means breaking API changes may occur across major versions.

⇄ Alternatives & Direct Competitors

T
tensorflow/tensorflow ★ 200.6K C++

An Open Source Machine Learning Framework for Everyone

Compare ↗
R
react/react-native ★ 126.8K C++

A framework for building native applications using React

Compare ↗
T
microsoft/terminal ★ 105K C++

The new Windows Terminal and the original Windows console host, all in the same place!

Compare ↗
B
bitcoin/bitcoin ★ 90.3K C++

Bitcoin Core integration/staging tree

Compare ↗

👥 Who Should Use This

Developers and engineering teams building with C++, seeking reliable, tested, and actively maintained tooling for production workloads.

🏆 Nearby in the Rankings

ggml-org/llama.cpp is currently ranked #51 by stars across every repository tracked on GitHubRepo. These are adjacent projects:

RankRepositoryLanguageStarsAction
#46 farion1231/cc-switch Rust ★ 137.5K Compare ↗
#47 Comfy-Org/ComfyUI Python ★ 135K Compare ↗
#48 garrytan/gstack TypeScript ★ 134.3K Compare ↗
#49 excalidraw/excalidraw TypeScript ★ 133.1K Compare ↗
#50 nextlevelbuilder/ui-ux-pro-max-skill Python ★ 131K Compare ↗
#51 ggml-org/llama.cpp This Project C++ ★ 129.8K
#52 Chalarangelo/30-seconds-of-code JavaScript ★ 129.2K Compare ↗
#53 kubernetes/kubernetes Go ★ 128K Compare ↗
#54 react/react-native C++ ★ 126.8K Compare ↗
#55 openai/codex Rust ★ 126.7K Compare ↗
#56 shadcn-ui/ui TypeScript ★ 124.7K Compare ↗

Frequently Asked Questions

What does llama.cpp do? +

LLM inference in C/C++

What language is llama.cpp written in? +

The primary language is C++. Topics include: ggml.

Is llama.cpp actively maintained? +

Yes, the last recorded push was on Sep 28, 2026 with 2,558 open issues being tracked.

How many stars does llama.cpp have? +

llama.cpp has 129,763 stars and 23,851 forks on GitHub.

How does llama.cpp rank among GitHub repositories? +

With 129,763 stars, ggml-org/llama.cpp is ranked #51 globally across all repositories tracked on GitHubRepo and #2 among C++ projects.

What license is llama.cpp distributed under? +

The repository reports a MIT license. Always verify the repository LICENSE file for legal terms.

From our network
FOR MAINTAINERS

Built something? Put it in front of millions of developers.

We make a short reel about your project and post it across YouTube, Instagram, Threads, and X. Send a link, we do the rest.