darthbear/goryproxy
Public repository.
Discovered public repositories for darthbear in the GitHub catalog.
Public repository.
Simple scrapy downloader implemented using what is described in http://web.archive.org/web/20120316092048/http://dev.scrapy.org/ticket/153
MongoDB adapter for Hadoop. Small mongo hadoop pig patch to allow to use mongodb fields that starts with an underscore by prefixing them with u_ (e.g., u__id instead of _id).
Use scrapy with mongodb to store the request queues (FIFO or LIFO)
Python HTTP Requests for Humans™. Small update to add source_address support
MongoDB pipeline for Scrapy. It allows to update existing entries (set new values or add elements to array) when item values are spread over multiple pages
Scrapy Plugin to use the simple http queue as the queue for the URLs in order to allow distributed crawling
Simple HTTP queue (FIFO and LIFO) implemented using Python, SQLite3 and Tornado. It supports multiple queues.
Use scrapy with a list of proxies generated from proxynova.com
Redis-based components for scrapy that allows distributed crawling. Small update to make it work for Scrapy 0.16+ and added QUEUE_TYPE and DUPE_FILTER options