Turn any PDF or image document into structured data for your AI. A powerful, lightweight OCR toolkit that bridges the gap between images/PDFs and LLMs. Supports 100+ languages.
Top OCR GitHub Repositories & Tools (2026)
Discover the most starred and trending open source tools tagged with #ocr.
Computer use for about $0.0002 a step: OCR the screen, classify the next action with TypeSafe, click. macOS.
Fast Rust library for PDF inspection, classification, and text extraction. Intelligently detects scanned vs text-based PDFs to enable smart routing decisions.
Any file → clean Markdown for AI agents: PDF, Word, PowerPoint, Excel, EPUB, HTML and web pages, images (OCR), audio and video (metadata, subtitles, transcripts). A fast Rust MCP server and CLI that runs on your machine. No API key.
A community-supported supercharged document management system: scan, index and archive all your documents
Desktop app for managing BibTeX and BibLaTeX (.bib) libraries
Free, open-source menu bar toolkit for macOS. Eighteen tools in one panel: Pomodoro timer, time tracker, keep-awake, system monitor, clipboard history, file converter, window manager, torrents, archives, OCR, screenshots, draw on screen, color picker, VPN switcher, app launcher, uninstaller, keyboard lock, speed test. SwiftUI, no telemetry, MIT.
Use LLMs and LLM Vision (OCR) to handle paperless-ngx - Document Digitalization powered by AI
ShareX is a free and open-source application that enables users to capture or record any area of their screen with a single keystroke. It also supports uploading images, text, and various file types to a wide range of destinations.
The Privacy First PDF Toolkit
🖼️ Image Toolbox is a powerful app for advanced image manipulation. It offers dozens of features, from basic tools like crop and draw to filters, OCR, and a wide range of image processing options
PyMuPDF is a high performance Python library for data extraction, analysis, conversion & manipulation of PDF (and other) documents.
MCP server that lets Claude Code and other AI agents read and search large PDFs, one file or a whole folder: agentic RAG with hybrid semantic + keyword search, selective page reads, tables, images, OCR, chart data, and multi-column/CJK layouts.
Cross-device clipboard sync for macOS, Windows & Linux — end-to-end encrypted, LAN-only, no cloud. OCR, CLI and MCP server built in.
⌘⌘ - A lightweight, native macOS screenshot tool that lives in your menu bar. Double-tap ⌘ Command to capture any region of your screen — instantly copied to clipboard, or annotate first with pen and mosaic tools.
Quick, painless, intuitive OCR platform written in Rust and TypeScript. Modern UI with modern API, with an emphasis on intuitive user experience.
Hints, grids and vim keys for your whole desktop. Built for macOS, linux and windows, natively.
JavaScript OCR and text extraction for images and PDFs.
A program to recognize text on the screen
Your new favorite, free, and open source screenshot tool