MockServer is an HTTP(S) mock server and proxy for testing that lets you mock APIs, inspect and modify live traffic, and inject failures. It supports HTTP/1.1, HTTP/2, gRPC, WebSockets, TCP and more on a single port, with additional support for HTTP/3, message brokers, and AI/LLM APIs.
Best Local LLMs GitHub Repositories
Top 302 open source projects in Local LLMs. Ranked by stars, activity, and recorded community growth.
Achieve state of the art inference performance with modern accelerators on Kubernetes
Koog is a JVM (Java and Kotlin) framework for building predictable, fault-tolerant and enterprise-ready AI agents across all platforms – from backend services to Android and iOS, JVM, and even in-browser environments. Koog is based on our AI products expertise and provides proven solutions for complex LLM and AI problems
AI observability platform for production LLM and agent systems.
The Ruby-native AI framework. Chats, agents, tools, images, audio, and video through one consistent API, in plain Ruby or Rails.
Evaluation and Tracking for LLM Experiments and AI Agents
Distributed AI/LLM for the people. Share compute privately or publicly to power your agents and chat.
Switchyard lets LLM applications route traffic across models and providers while preserving native OpenAI and Anthropic API compatibility - enabling flexible model selection, benchmarking, and cost/performance optimization.
The Swiss Army Knife of Offline AI. Chat, see, speak, and generate images on your phone or Mac — GGUF LLMs, vision, Whisper speech-to-text, Stable Diffusion, tool calling, and local-network servers. Runs on your CPU, GPU, or NPU. No account, no API key, zero data leaves your device.
Add a real-time analytics node to your operational database. Spice is a portable, accelerated SQL query, search, and LLM-inference engine in Rust for data-grounded AI apps and agents.
Miles is an enterprise-facing reinforcement learning framework for LLM and VLM post-training, forked from and co-evolving with slime.
LLM speculative inference server for heterogeneous hardware & consumer GPUs
Use LLMs and LLM Vision (OCR) to handle paperless-ngx - Document Digitalization powered by AI
OpenLake is a high performance storage engine for efficient LLM inference and GPU Training
Atomic Agent is a local-first AI agent. Runs open-weight models on your own machine via llama.cpp.
🥨 Lobe Icons - Brings AI/LLM brand logos to your React & React Native apps — static SVG/PNG/WebP, no dependencies.
the terminal client for LLMs
A simple, performant, and scalable Jax LLM!
Open-source AI Security Operations Center: alert fusion, LLM-agent triage, MITRE ATT&CK investigation, and a replayable decision ledger for every agent step. Self-hostable, runs with no API keys, MIT licensed. Ships an MCP server for Claude, Cursor and Continue.