FlashRT is a high-performance realtime inference engine for small-batch, latency-sensitive AI workloads. The flagship integration is production VLA control for Pi0, Pi0.5, GROOT N1.6, and Pi0-FAST. Also support llm e.g, qwen3.6-27B
Top INFERENCE GitHub Repositories & Tools (2026)
Discover the most starred and trending open source tools tagged with #inference.
RuVector provides High Performance, Real-Time decisions and agent memory , Self-Learning Ai, Vector GNN DB built in Rust.
Run a 105 GB AI model on a Mac that can't hold it. Slotstream streams Qwen3.8-Flash-Next (125B mixture of experts) from your SSD and caches the busiest experts in memory, so it runs on Macs with 16 to 64 GB. One native Swift binary on MLX and Metal, no Python, offline. Works with Claude Code, Codex and Ollama or OpenAI clients.
Cross-platform, customizable ML solutions for live and streaming media.
LLM inference server with continuous batching & SSD caching for Apple Silicon — managed from the macOS menu bar
Gemma 4 26B-A4B inference in ~2 GB of RAM on any M-series MacBook
Python Toolkit for Causal and Probabilistic Reasoning
The Agentic Framework of the PHP ecosystem to build production-ready AI driven applications. Connect components (LLMs, Tools, vector DBs, memory) to agents that interact with your data and UI.
Ultrafast serverless GPU inference, sandboxes, and background jobs
Nvidia GPU exporter for prometheus using nvidia-smi binary OR using NVML
A scalable inference server for models optimized with OpenVINO™
Run open decision models locally: pull and serve Laya, decider, NLI and GLiClass behind a TypeSafe-compatible API. Ollama for decision models.
⚡ Native MLX Swift LLM inference server for Apple Silicon. OpenAI-compatible API, SSD streaming for 100B+ MoE models, TurboQuant KV cache compression, MACOS + iOS iPhone app.
Pure Rust + CUDA LLM inference engine — no PyTorch, OpenAI-compatible, serves Qwen3 to Kimi-K2
🏗️ Fine-tune, build, and deploy open-source LLMs easily!
High Performance KV Cache for LLM Inference
Central repository for the Holoscan Ecosystem
Rails as a specification; the deployment target is a build flag
An open-source, cloud-native, high-performance gateway unifying multiple LLM providers, from local solutions like Ollama to major cloud providers such as OpenAI, Groq, Cohere, Anthropic, Cloudflare and DeepSeek.