Jev-like family of decision models built on top of Qwen3.5/3.8 you can train and run on your own
Top QWEN3 GitHub Repositories & Tools (2026)
Discover the most starred and trending open source tools tagged with #qwen3.
A high-throughput and memory-efficient inference and serving engine for LLMs
FlashRT is a high-performance realtime inference engine for small-batch, latency-sensitive AI workloads. The flagship integration is production VLA control for Pi0, Pi0.5, GROOT N1.6, and Pi0-FAST. Also support llm e.g, qwen3.6-27B
Run a 105 GB AI model on a Mac that can't hold it. Slotstream streams Qwen3.8-Flash-Next (125B mixture of experts) from your SSD and caches the busiest experts in memory, so it runs on Macs with 16 to 64 GB. One native Swift binary on MLX and Metal, no Python, offline. Works with Claude Code, Codex and Ollama or OpenAI clients.
๐ Pytorch Distributed native training library for LLMs/VLMs with OOTB Hugging Face support
Pure Rust + CUDA LLM inference engine โ no PyTorch, OpenAI-compatible, serves Qwen3 to Kimi-K2
A native .NET LLM inference engine for GGUF models. TensorSharp provides a console application, a web-based chatbot interface, and Ollama/OpenAI-compatible HTTP APIs for programmatic access. It supports Windows/MacOS/iOS/Linux with full GPU capability
On-device Speech AI for Apple Silicon