inman/distribute_crawler
使用scrapy,redis, mongodb,graphite实现的一个分布式网络爬虫,底层存储mongodb集群,分布式使用redis实现,爬虫状态显示使用graphite实现
Discovered public repositories for inman in the GitHub catalog.
使用scrapy,redis, mongodb,graphite实现的一个分布式网络爬虫,底层存储mongodb集群,分布式使用redis实现,爬虫状态显示使用graphite实现
scikit-learn: machine learning in Python
Materials for "Python for Data Analysis" by Wes McKinney, published by O'Reilly Media
Html Content / Article Extractor, web scrapping lib in Python - Fork form Goose scala project form Gravity Labs
Html Content / Article Extractor in Scala - open sourced from Gravity Labs - http://gravity.com
lightweight, multi-purpose library of recommender system algorithms
Scriptable Headless WebKit
Cloud9 is a Hadoop toolkit for working with big data
A horridly implemented scrapy app that will scrape all (?) of Delicious' bookmarks.
Scrapy, a fast high-level screen scraping and web crawling framework for Python.
Programmatic web browsing module with AJAX support for Python
中山大学东校区校园网认证的客户端(非官方)
Coursera materials downloader.