A high-throughput and memory-efficient inference and serving engine for LLMs
Top KIMI GitHub Repositories & Tools (2026)
Discover the most starred and trending open source tools tagged with #kimi.
A floating macOS monitor for how much Claude Code, Codex, Antigravity, OpenCode Go and Kimi Code you have left
Mission control for your AI agents
External-model router for Codex with guided Kimi OAuth/API, DeepSeek, safe migration, and rollback.
A coding agent for open models like Kimi K3 and GLM 5.3
Local-first macOS app to browse, search, analyze, and resume supported AI coding-agent session history across Codex, Claude Code, OpenCode, Cursor Agent, Antigravity, Hermes, OpenClaw, Copilot CLI, and more.
๐ Pytorch Distributed native training library for LLMs/VLMs with OOTB Hugging Face support
TokenSpeed is a speed-of-light LLM inference engine.
The most accurate, cheapest and fastest memory for coding agents: searchable session history from Claude Code, Codex, Cursor and 32 more agents, already on your disk. No LLM, one Go binary.
Pure Rust + CUDA LLM inference engine โ no PyTorch, OpenAI-compatible, serves Qwen3 to Kimi-K2
Open-source AI pair programming for desktop: a Mentor + Executor agent cross-check each other's code to catch AI hallucinations. Works with Claude Code, Codex, Gemini & opencode. macOS / Windows / Linux.
Use 30+ AI models (DeepSeek, Kimi, GLM, Claude, GPT-5.5, Gemini, Grok) in GitHub Copilot Chat - free via BYOK. No Copilot Pro needed.
Run the pi coding agent as a service โ triggered on demand, on a cron schedule, or by a GitHub or GitLab issue, comment or pull/merge request โ in a container you control, with a durable queue, a spend cap, and a live admin panel.
Kimi K3 Code Free desktop - kimi k3, is kimi k3 free, kimi k3 huggingface, kimi k3 open weights, kimi k3 benchmarks, kimi k3 vs fable 5, kimi k3 cost, kimi k3 reddit, chinese ai, 200K context, GitHub PR review, Moonshot. Install:๐ก
Open Model Engine (OME) โ Kubernetes operator for LLM serving, GPU scheduling, and model lifecycle management. Works with SGLang, vLLM, TensorRT-LLM, and Triton