Infrastructure for continually self‑improving agents
Top INFERENCE GitHub Repositories & Tools (2026)
Discover the most starred and trending open source tools tagged with #inference.
Your Cheat Sheet For AI Engineering Interviews at Top AI Companies - Questions and Answers.
Native MLX runtime for Laya typed decision models — 7–14 ms short decisions on M3 Max. No text generation, PyTorch, or cloud API.
Port of OpenAI's Whisper model in C/C++
A high-throughput and memory-efficient inference and serving engine for LLMs
SGLang is a high-performance serving framework for large language models and multimodal models.
Standardized Distributed Generative and Predictive AI Inference Platform for Scalable, Multi-Framework Deployment on Kubernetes
NobodyWho is an inference engine that lets you run LLMs locally and efficiently on any device.
A local inference engine for Apple silicon, built around the model.
Mooncake is the serving platform for Kimi, a leading LLM service provided by Moonshot AI.
A framework for efficient model inference with omni-modality models
AI sovereignty for the world's most trusted runtime.
Ray is an AI compute engine. Ray consists of a core distributed runtime and a set of AI Libraries for accelerating ML workloads.
TypeScript-first schema validation with static type inference
Open-source inference server and production cluster for all the models your agent needs.
Lemonade helps users discover and run local AI apps by serving optimized LLMs right from their own GPUs and NPUs. Join our discord: https://discord.gg/5xXzkMu8Zk
Open-source local AI SDK - run AI on-device with no cloud, no API keys. Supports GGUF, RAG, image, music, and video generation, speech-to-text, P2P inference, and more. Cross-platform: Linux, macOS, Windows, Android, iOS.
Manages Unified Access to Generative AI Services built on Envoy Gateway
FlashRT is a high-performance realtime inference engine for small-batch, latency-sensitive AI workloads. The flagship integration is production VLA control for Pi0, Pi0.5, GROOT N1.6, and Pi0-FAST. Also support llm e.g, qwen3.6-27B
RuVector provides High Performance, Real-Time decisions and agent memory , Self-Learning Ai, Vector GNN DB built in Rust.