Infrastructure for continually self‑improving agents
Top NFE GitHub Repositories & Tools (2026)
Discover the most starred and trending open source tools tagged with #nfe.
Your Cheat Sheet For AI Engineering Interviews at Top AI Companies - Questions and Answers.
Native MLX runtime for Laya typed decision models — 7–14 ms short decisions on M3 Max. No text generation, PyTorch, or cloud API.
A local inference engine for Apple silicon, built around the model.
Port of OpenAI's Whisper model in C/C++
A high-throughput and memory-efficient inference and serving engine for LLMs
SGLang is a high-performance serving framework for large language models and multimodal models.
Standardized Distributed Generative and Predictive AI Inference Platform for Scalable, Multi-Framework Deployment on Kubernetes
NobodyWho is an inference engine that lets you run LLMs locally and efficiently on any device.
A native .NET LLM inference engine and agent runtime for GGUF models. TensorSharp provides a console application, a web-based chatbot interface, iPhone App, and Ollama/OpenAI-compatible HTTP APIs for programmatic access. It supports Windows/MacOS/iOS/Linux with full GPU capability
Mooncake is the serving platform for Kimi, a leading LLM service provided by Moonshot AI.
A framework for efficient model inference with omni-modality models
Peer-to-peer LLM inference in the browser: pool your devices to run big open models, for chat and coding agents.
AI sovereignty for the world's most trusted runtime.
Ray is an AI compute engine. Ray consists of a core distributed runtime and a set of AI Libraries for accelerating ML workloads.
TypeScript-first schema validation with static type inference
Open-source inference server and production cluster for all the models your agent needs.
Lemonade helps users discover and run local AI apps by serving optimized LLMs right from their own GPUs and NPUs. Join our discord: https://discord.gg/5xXzkMu8Zk
Open-source local AI SDK - run AI on-device with no cloud, no API keys. Supports GGUF, RAG, image, music, and video generation, speech-to-text, P2P inference, and more. Cross-platform: Linux, macOS, Windows, Android, iOS.
Manages Unified Access to Generative AI Services built on Envoy Gateway