A local inference engine for Apple silicon, built around the model.
Decoding GitHub Repositories (2026) · Open Source Tools
Discover the best Decoding open source repositories on GitHub. Explore trending tools, live star counts, source code, and developer projects updated daily.
Your iPhone helps your Mac run a 27B model: faster prompt reading and more context over a USB-C cable
llama.cpp fork for significantly improved performance on Ampere (especially RTX 3090 / 3090 Ti): example: 95+ tok/s over a 100K-token generation at temperature 1 for qwen3.8, supporting full 262K ctx.
Peer-to-peer LLM inference in the browser: pool your devices to run big open models, for chat and coding agents.
The set-and-forget LLM engine for Pascal and Volta: PXQ codec + kernels, auto-tuned per card. Ready-to-run PXQ models in MODELS.md; benchmarks vs llama.cpp in the README.
llama.cpp/ik_llama.cpp launcher: loads big MoE models across mismatched multi-GPU rigs by exact VRAM math.
ECMWF's GRIB and BUFR decoding/encoding library
I/O interface and utilities for CCSDS binary spacecraft data in Python. Library used in flight missions at NASA, NOAA, and SWRI
An open-source virtual AI Infra team that discovers, benchmarks, and safely upgrades local LLM inference—with independent quality gates, a stable OpenAI-compatible API, and automatic rollback.
JEV-inspired parallel decisions for CUDA LLMs. One context, many decisions. vLLM API, game-agent examples, and reproducible benchmarks.
Claude Code Monitor : Monitor & audit every action Claude Code takes on your computer.
A support library for Ronin. Like activesupport, but for hacking!
LLM speculative inference server for heterogeneous hardware & consumer GPUs