FlashRT is a high-performance realtime inference engine for small-batch, latency-sensitive AI workloads. The flagship integration is production VLA control for Pi0, Pi0.5, GROOT N1.6, and Pi0-FAST. Also support llm e.g, qwen3.6-27B
Topic Catalog
Top CUDA-KERNELS GitHub Repositories & Tools (2026)
Discover the most starred and trending open source tools tagged with #cuda-kernels.
SPONSORED BY
4 matching projects · updated Sep 28, 2026
How we rank ↗
↑ +4 today
LLM speculative inference server for heterogeneous hardware & consumer GPUs
↑ +1 today
Pure Rust + CUDA LLM inference engine — no PyTorch, OpenAI-compatible, serves Qwen3 to Kimi-K2
↑ +1 today
cuDNN Frontend is NVIDIA's modern, open-source entry point to the cuDNN library and a growing collection of high-performance open-source kernels.
↑ +1 today
From our network