The web data API to search, scrape, and interact at scale. ๐ฅ
Top SCRAPING GitHub Repositories & Tools (2026)
Discover the most starred and trending open source tools tagged with #scraping.
๐ท๏ธ An adaptive Web Scraping framework that handles everything from a single request to a full-scale crawl! Don't be shy, join here: https://discord.gg/EMgGbDceNQ and follow here for daily tips and tricks: https://x.com/Scrapling_dev
Best headless browser for AI agents. Lite, Fast, High-Compatibility. Built in Rust
Open-source web crawler and scraper for LLMs and AI agents: any website into clean, LLM-ready Markdown. Run it yourself, or use Crawl4AI Cloud with one key.
Create agents that monitor and act on your behalf. Your agents are standing by!
A curated directory of ecommerce intelligence APIs for products, prices, reviews, sellers, and marketplaces.
The only browser automation that bypasses anti-bot systems. AI writes network hooks, clones UIs pixel-perfect via simple chat.
A native Android app that aggregates third-party app stores into one catalogue. Search across all of them at once, compare what each publishes for the same app, download and install APKs from the store you pick, and keep what you installed up to date.
A curated directory of job data APIs and scrapers for listings, hiring signals, salaries, and recruiting.
PHP Curl Class makes it easy to send HTTP requests and integrate with web APIs
The headless Chrome/Chromium driver on top of Puppeteer. Take screenshots, generate PDFs, extract text and HTML with a production-ready API.
โค๏ธ Fredy - [F]ind [R]eal [E]state [D]amn Eas[y] - Fredy keeps searching for new apartments, houses, and flats in Europe on platforms like ImmoScout24, Immowelt, eBay Kleinanzeigen and instantly delivers the results to you via Slack, Telegram, Email, Discord or ntfy, so you can focus on the more important things in life ;)
Self-hosted SERP API for AI, SEO & automation. Browser-rendered Google, Bing, Yandex, Baidu, DuckDuckGo and Ecosia search with page extraction ๐
A versatile Ruby web spidering library that can spider a site, multiple domains, certain links or infinitely. Spidr is designed to be fast and easy to use.
A high-performance proxy rotation engine with automated IP management and real-time health monitoring
Provides a list of fresh, working proxy servers (HTTP, HTTPS, SOCKS4 & SOCKS5) with multiple formats available for download.
(InstaTools 2.0 soon to come ๐) ๐งฐ A collection of automation tools for Instagram ๐ฑ| Written in Python ๐ | Don't forget to โญ the repo !
โก๏ธ Free Verified HTTP, SOCKS5, & SOCKS4 Proxy List โฐ Updated every 30 minutes
this shows how to use github actions to do periodic data scraping
Common Release Data for various projects in a consumable format, automatically updated.