FlashRT is a high-performance realtime inference engine for small-batch, latency-sensitive AI workloads. The flagship integration is production VLA control for Pi0, Pi0.5, GROOT N1.6, and Pi0-FAST. Also support llm e.g, qwen3.6-27B
Top KERNELS GitHub Repositories & Tools (2026)
Discover the most starred and trending open source tools tagged with #kernels.
LLM speculative inference server for heterogeneous hardware & consumer GPUs
A lightweight, LXC-like container runtime for Android and Linux. Run full Linux distributions natively with zero performance penalty
FlagGems is an operator library for large language models implemented in the Triton Language.
Pure Rust + CUDA LLM inference engine — no PyTorch, OpenAI-compatible, serves Qwen3 to Kimi-K2
Yet Another Root Checker and Play Integrity API Application.
AgentKernelArena provides an end-to-end siloed-benchmarking environment where different LLM-powered agents—such as Cursor Agent, Claude Code, Codex, SWE-agent, and GEAK—can be evaluated side-by-side on the same GPU kernel tasks, using objective and reproducible metrics.
cuDNN Frontend is NVIDIA's modern, open-source entry point to the cuDNN library and a growing collection of high-performance open-source kernels.
:metal: TT-NN operator library, and TT-Metalium low level kernel programming model.