lemurproject/clueweb12pp-core
Unifies all code written for processing clueweb12++
Discovered public repositories for lemurproject in the GitHub catalog.
Unifies all code written for processing clueweb12++
Tools for processing our nabble crawl
Processing the Yahoo groups crawl (part of clueweb12++)
Download the reddit dataset for ClueWeb12++
Collection of Scrapers I am writing
Command line utilities for working with WARC files
bindings for the Blekko search engine API
Tools for processing the Trec KBA dataset
Simple command line tool to export the ClueWeb dataset as HTML files.
Oauth Implementation in Racket (PLT Scheme)
Public repository.
Python library for reading and writing warc files
A collection of tools for running the crawler, including scripts and scrapers to collect seeds.
language identification module in racket
Code to sample forums and check stats on pages and their information content
Configuration files and scripts needed to run the Heritrix jobs
Code and other stuff for discussion-forum seeds.
Extension to the Clueweb12 dataset
Heritrix is the Internet Archive's open-source, extensible, web-scale, archival-quality web crawler project.