darthbear/goryproxy
Public repository.
Scrapy Plugin to use the simple http queue as the queue for the URLs in order to allow distributed crawling
This repository is cataloged as part of our automated global GitHub synchronization. Full telemetry, velocity snapshots, and code summaries are scheduled for continuous enrichment.
Public repository.
Simple scrapy downloader implemented using what is described in http://web.archive.org/web/20120316092048/http://dev.scrapy.org/ticket/153
MongoDB adapter for Hadoop. Small mongo hadoop pig patch to allow to use mongodb fields that starts with an underscore by prefixing them with u_ (e.g., u__id instead of _id).
Use scrapy with mongodb to store the request queues (FIFO or LIFO)