LeadsPlus/scrapy
Scrapy, a fast high-level screen scraping and web crawling framework for Python.
LA-PDFText has been developed by members of the Biomedical Knowledge Engineering group @ the Information Sciences Institute. It is intended for use both scientists and NLP engineers interested in getting access to text within specific sections of research articles. The system is open-source and provides a simple baseline function for extracting text from primary research articles using rules that developers can customize. This means that the system works quite well for most applications (and might occasionally make mistakes and extract the wrong text), but it is always possible to 'hack' your own rules and improve performance.
This repository is cataloged as part of our automated global GitHub synchronization. Full telemetry, velocity snapshots, and code summaries are scheduled for continuous enrichment.
Scrapy, a fast high-level screen scraping and web crawling framework for Python.
Extensions for using Scrapy on Amazon AWS
A redis-based parallel application
Use scrapy with mongodb to store the request queues (FIFO or LIFO)