Local Laya typed decisions on Apple Core ML and Neural Engine. Validated ports, ~5 ms short decisions on M3 Max, reproducible speed and energy benchmarks.
Top ON-DEVICE GitHub Repositories & Tools (2026)
Discover the most starred and trending open source tools tagged with #on-device.
Local computer use on Apple silicon: on-device typed decisions with laya and form filling with CUA-S1-FORMS, driven through the macOS Accessibility API
NobodyWho is an inference engine that lets you run LLMs locally and efficiently on any device.
Open-source local AI SDK - run AI on-device with no cloud, no API keys. Supports GGUF, RAG, image, music, and video generation, speech-to-text, P2P inference, and more. Cross-platform: Linux, macOS, Windows, Android, iOS.
Muesli: agent-native local meeting transcription + dictation for macOS (Granola + WisprFlow alternative)
The free AI already on your Mac. CLI tool, OpenAI-compatible server, and interactive chat โ all on-device via Apple Intelligence. No API keys, no cloud, no downloads.
Run a 105 GB AI model on a Mac that can't hold it. Slotstream streams Qwen3.8-Flash-Next (125B mixture of experts) from your SSD and caches the busiest experts in memory, so it runs on Macs with 16 to 64 GB. One native Swift binary on MLX and Metal, no Python, offline. Works with Claude Code, Codex and Ollama or OpenAI clients.
๐ Your face is the password. Face unlock for your Mac's lock screen, plus App Lock for chosen apps. 100% on-device, built with Swift and Core ML.
Gemma 4 26B-A4B inference in ~2 GB of RAM on any M-series MacBook
AI You Control: Choose your models. Own your data. Eliminate vendor lock-in.
Sudoless Apple Silicon system monitor (native SwiftUI GUI) with ANE / Media Engine / memory-bandwidth tracking
Shared Single-file memory layer for all your agents, sub mili-second RAG over text, photo and video on Apple Silicon.. No Server. No API. One File. Pure Swift
โก Native MLX Swift LLM inference server for Apple Silicon. OpenAI-compatible API, SSD streaming for 100B+ MoE models, TurboQuant KV cache compression, MACOS + iOS iPhone app.
Swiftlet is a Swift and Metal runtime that runs large Qwen Mixture-of-Experts models locally on Apple devices by streaming expert weights from storage, enabling 35B and 80B models to run with low RAM, including on iPhone.
An on-device LLM understands your entire life, then proactively offers to get your work done through computer use.
PennyWise automatically reads transaction SMS messages and transforms them into organized financial data with on-device AI assistance. No manual entry, no cloud processing, complete privacy.
Cross-platform on-device AI toolkit
TypeWhisper for Windows - Local speech-to-text with translation
On-device meeting transcriber for macOS โ auto-records Teams/Zoom/Webex, transcribes & separates speakers locally. No cloud. Open-source alternative to Otter/Granola/Fireflies.
Records and transcribes online meetings. Automatically