The web data API to search, scrape, and interact at scale. ๐ฅ
Top EXTRACT GitHub Repositories & Tools (2026)
Discover the most starred and trending open source tools tagged with #extract.
๐ท๏ธ An adaptive Web Scraping framework that handles everything from a single request to a full-scale crawl! Don't be shy, join here: https://discord.gg/EMgGbDceNQ and follow here for daily tips and tricks: https://x.com/Scrapling_dev
Open-source web crawler and scraper for LLMs and AI agents: any website into clean, LLM-ready Markdown. Run it yourself, or use Crawl4AI Cloud with one key.
Turn any PDF or image document into structured data for your AI. A powerful, lightweight OCR toolkit that bridges the gap between images/PDFs and LLMs. Supports 100+ languages.
Fast Rust library for PDF inspection, classification, and text extraction. Intelligently detects scanned vs text-based PDFs to enable smart routing decisions.
Polyglot document intelligence with a Rust core: extract text, metadata, images, tables, and structured data from 106 formats across 140 file extensions, plus code intelligence for 371 languages. Fifteen bindings, with CLI, REST API, and MCP server.
SwiftSoup: Pure Swift HTML Parser, with best of DOM, CSS, and jquery (Supports Linux, iOS, Mac, tvOS, watchOS)
Database Subsetting and Relational Data Browsing Tool.
Free open-source extractor for AI coding assistant chat histories. Supports Claude Code, Cursor, Windsurf, Aider, Cline/Roo Code, and more.
Desbordante is a high-performance data profiler that is capable of discovering many different patterns in data using various algorithms. It also allows to run data cleaning scenarios using these algorithms. Desbordante has a console version and an easy-to-use web application.
Graby helps you extract article content from web pages
Any source (PDF, video, web, audio, text) to interactive learning package with quizzes, flashcards and spaced repetition. One command, 12-section study guide.
A fork of https://bitbucket.org/fivefilters/php-readability
Unified real-time data engine
MCP server that lets Claude Code and other AI agents read and search large PDFs, one file or a whole folder: agentic RAG with hybrid semantic + keyword search, selective page reads, tables, images, OCR, chart data, and multi-column/CJK layouts.
Pure Python HWPX automation: read, edit, generate, and validate documents without Hancom Office.
Adaptive video frame extraction for SfM, Gaussian Splatting, and photogrammetry.
The "Missing GitHub Status Page" -- a Flat Data attempt at historically documenting GitHub statuses
Archive manager and 7zip replacement for macOS. Preview (nested) archives without extracting them. Extract single files.
CLI tool for saving a faithful copy of a complete web page in a single HTML file (based on SingleFile)