High-performance web crawling engine with bindings for 11 languages
Top SCRAPING GitHub Repositories & Tools (2026)
Discover the most starred and trending open source tools tagged with #scraping.
AI-native web scraper. Single binary with a bundled Claude Code skill. MIT-licensed alternative to Firecrawl.
โฌ๏ธ A simple all-in-one CLI tool to download EVERYTHING from a URL (like youtube-dl/yt-dlp, forum-dl, gallery-dl, simpler ArchiveBox). ๐ญ Uses headless Chrome to get HTML, JS, CSS, images/video/audio/subtitles, PDFs, screenshots, article text, git repos, and more...
Proxify is an automated tool that collects and updates fresh SOCKS4, SOCKS5, HTTP/HTTPS proxies and V2Ray configs from public sources. It maintains a regularly updated repository of working proxies, useful for privacy, bypassing restrictions, or network testing.
Free Proxy List Updated every 10 minutes
Cloudflare Turnstile challenge internals: the live request flow, the challenge bundle, and a capture toolkit.
High-performance browser automation bridge and multi-instance orchestrator with advanced stealth injection and real-time dashboard.
Free, self-hosted X (Twitter) scraper no API key, no credits. Uses your own browser session via Playwright. Export tweets, search results & timelines as LLM-ready JSON/CSV/Markdown. Ships an AI agent Skill for OpenClaw & Hermes.
CLI tool for saving a faithful copy of a complete web page in a single HTML file (based on SingleFile)
Mechanize is a ruby library that makes automated web interaction easy.