π₯ Supercharge your AI agents with data from the web and beyond. A web data API to search, scrape, and access more sources.
Top SCRAPER GitHub Repositories & Tools (2026)
Discover the most starred and trending open source tools tagged with #scraper.
Give your AI agent eyes to see the entire internet. Read & search Twitter, Reddit, YouTube, GitHub, Bilibili, XiaoHongShu β one CLI, zero API fees.
π·οΈ An adaptive Web Scraping framework that handles everything from a single request to a full-scale crawl! Don't be shy, join here: https://discord.gg/EMgGbDceNQ and follow here for daily tips and tricks: https://x.com/Scrapling_dev
Best headless browser for AI agents. Lite, Fast, High-Compatibility. Built in Rust
Open-source web crawler and scraper for LLMs and AI agents: any website into clean, LLM-ready Markdown. Run it yourself, or use Crawl4AI Cloud with one key.
Free, self-hosted X (Twitter) scraper no API key, no credits. Uses your own browser session via Playwright. Export tweets, search results & timelines as LLM-ready JSON/CSV/Markdown. Ships an AI agent Skill for OpenClaw & Hermes.
CrawleeβA web scraping and browser automation library for Node.js to build reliable crawlers. In JavaScript and TypeScript. Extract data for AI, LLMs, RAG, or GPTs. Download HTML, PDF, JPG, PNG, and other files from websites. Works with Puppeteer, Playwright, Cheerio, JSDOM, and raw HTTP. Both headful and headless mode. With proxy rotation.
Create agents that monitor and act on your behalf. Your agents are standing by!
β‘οΈ Free Verified HTTP, SOCKS5, & SOCKS4 Proxy List β° Updated every 30 minutes
Fast proxy scraper and checker written in Rust. Collects HTTP, SOCKS4 and SOCKS5 proxies from any number of sources, verifies each one really works, and writes JSON and plain text with response time, exit IP, ASN and geolocation. Single binary for Windows, Linux, macOS and Android, with an interactive TUI.
Client-side downloader app built with Capacitor and Tauri.
Free HTTP, HTTPS, SOCKS4 & SOCKS5 proxy list. ~22k proxies across 90+ countries, refreshed every 5 minutes from the ProxyScrape v4 API. TXT, JSON & CSV.
Free open-source desktop app that scrapes JAV metadata and generates NFO + cover art for Jellyfin, Emby & Kodi. No Docker, no CLI β one-click install on Windows & macOS. 8 built-in sources + optional Metatube federation (30+ providers), actress collections, cross-language tag aliases, and a REST API for AI agents.
CLI tool for saving a faithful copy of a complete web page in a single HTML file (based on SingleFile)
APIs for browser automation, testing, and bypassing bot-detection. Makes Chromium an anti-detect browser for handling CAPTCHAs. Supports MCP.
Jobs scraper library for LinkedIn, Indeed, Glassdoor, Google, ZipRecruiter & more
Open-source MCP server for LinkedIn. Give Claude and any MCP-compatible AI agent access to profiles, companies, jobs, and messages.
NewPipe's core library for extracting data from streaming sites
An ergonomic, privacy-aware Python HTTP Client
source for Open States scrapers