socialpercon/selenium-crawler
Sometimes sites make crawling hard. Selenium-crawler uses selenium automation to fix that.
A generic Java scraper that can be easily extended.
This repository is cataloged as part of our automated global GitHub synchronization. Full telemetry, velocity snapshots, and code summaries are scheduled for continuous enrichment.
Sometimes sites make crawling hard. Selenium-crawler uses selenium automation to fix that.
Heritrix is the Internet Archive's open-source, extensible, web-scale, archival-quality web crawler project.
Bixo is an open source web mining toolkit that runs as a series of Cascading pipes on top of Hadoop. By building a customized Cascading pipe assembly, you can quickly create specialized web mining applications.
Example code for the book "Indexing Data in Apache Solr"