A local inference engine for Apple silicon, built around the model.
Top SPECULATIVE-DECODING GitHub Repositories & Tools (2026)
Discover the most starred and trending open source tools tagged with #speculative-decoding.
Peer-to-peer LLM inference in the browser: pool your devices to run big open models, for chat and coding agents.
LLM speculative inference server for heterogeneous hardware & consumer GPUs
An open-source virtual AI Infra team that discovers, benchmarks, and safely upgrades local LLM inference—with independent quality gates, a stable OpenAI-compatible API, and automatic rollback.
The set-and-forget LLM engine for Pascal and Volta: PXQ codec + kernels, auto-tuned per card. Ready-to-run PXQ models in MODELS.md; benchmarks vs llama.cpp in the README.