rdhyee/cc-mrjob
Demonstration of using Python to process the Common Crawl dataset with the mrjob framework
tutorials around http://commoncrawl.org/
This repository is cataloged as part of our automated global GitHub synchronization. Full telemetry, velocity snapshots, and code summaries are scheduled for continuous enrichment.
Demonstration of using Python to process the Common Crawl dataset with the mrjob framework
gzipstream allows Python to process multi-part gzip files from a streaming source
Working out how to copy data to a docker container's filesystem
Docker container for the IPython notebook