#1 TRENDING pegainfer-project/pegainfer Rust NEW 2026 Pure Rust + CUDA LLM inference engine — no PyTorch, OpenAI-compatible, serves Qwen3 to Kimi-K2 ◇ cuda◇ cuda-kernels◇ deepseek◇ gpu 110 ↑ +1 today 712
#2 TRENDING mudler/vllm.cpp C++ NEW 2026 a community oriented 1:1, vLLM-alike (Continuous batching, paged KV) engine in C++ with additional features (GGUF, RadixAttention, Cache-aware scheduling, ...) ◇ continous-batching◇ cpp20◇ llm◇ mlx 55 ↑ +1 today 434