guoyunsky/heritrix3
Heritrix is the Internet Archive's open-source, extensible, web-scale, archival-quality web crawler project.
demo for my gitbook
This repository is cataloged as part of our automated global GitHub synchronization. Full telemetry, velocity snapshots, and code summaries are scheduled for continuous enrichment.
Heritrix is the Internet Archive's open-source, extensible, web-scale, archival-quality web crawler project.
Daemon intended to monitor a queue to which Heritrix will submit URLs. On receipt, the URL is submitted to a webservice (currently via django-phantomjs) and stores the response, a modified HAR record, in a WARC file.
Python library for reading and writing warc files
Introspected tunnels to localhost