PyMuPDF is a high performance Python library for data extraction, analysis, conversion & manipulation of PDF (and other) documents.
Topic Catalog
Top TABLE-EXTRACTION GitHub Repositories & Tools (2026)
Discover the most starred and trending open source tools tagged with #table-extraction.
SPONSORED BY
3 matching projects · updated Sep 29, 2026
How we rank ↗
↑ +1 today
Polyglot document intelligence with a Rust core: extract text, metadata, images, tables, and structured data from 106 formats across 140 file extensions, plus code intelligence for 371 languages. Fifteen bindings, with CLI, REST API, and MCP server.
↑ +1 today
MCP server that lets Claude Code and other AI agents read and search large PDFs, one file or a whole folder: agentic RAG with hybrid semantic + keyword search, selective page reads, tables, images, OCR, chart data, and multi-column/CJK layouts.
↑ +1 today
From our network