Local AI song generator with an editable score — YuE2 on your GPU: full songs with vocals, sheet music, covers, exact replay. Native Windows app, no Python, installer with auto-update.
Top CUDA GitHub Repositories & Tools (2026)
Discover the most starred and trending open source tools tagged with #cuda.
Cross-vendor 3D Gaussian Splatting trainer - video to splat to mesh, Vulkan or CUDA.
A high-throughput and memory-efficient inference and serving engine for LLMs
SGLang is a high-performance serving framework for large language models and multimodal models.
Bend 2: a fast language that blocks AI mistakes via proof. Install: curl -fsSL https://bend-lang.com/install.sh | sh
A native .NET LLM inference engine and agent runtime for GGUF models. TensorSharp provides a console application, a web-based chatbot interface, iPhone App, and Ollama/OpenAI-compatible HTTP APIs for programmatic access. It supports Windows/MacOS/iOS/Linux with full GPU capability
AI sovereignty for the world's most trusted runtime.
Train and serve LLMs at extreme speed and massive throughput.
cuTile Rust provides a safe, tile-based kernel programming DSL for the Rust programming language. It features a safe host-side API for passing tensors to asynchronously executed kernel functions.
Multi-platform high-performance compute language extension for Rust.
FlashRT is a high-performance realtime inference engine for small-batch, latency-sensitive AI workloads. The flagship integration is production VLA control for Pi0, Pi0.5, GROOT N1.6, and Pi0-FAST. Also support llm e.g, qwen3.6-27B
A Python framework for GPU-accelerated simulation, robotics, and machine learning.
The set-and-forget LLM engine for Pascal and Volta: PXQ codec + kernels, auto-tuned per card. Ready-to-run PXQ models in MODELS.md; benchmarks vs llama.cpp in the README.
NVIDIA cuML: GPU-Accelerated Machine Learning
LLM speculative inference server for heterogeneous hardware & consumer GPUs
LUPINE is a GPU over IP bridge allowing GPUs on remote machines to be attached to CPU-only machines.
Ultrafast serverless GPU inference, sandboxes, and background jobs
🔥🔥🔥TensorRT for YOLOv8、YOLOv8-Pose、YOLOv8-Seg、YOLOv8-Cls、YOLOv7、YOLOv6、YOLOv5、YOLONAS......🚀🚀🚀CUDA IS ALL YOU NEED.🍎🍎🍎
High-Performance Rendering Framework on Stream Architectures