rjurney/incubator-datafu
Mirror of Apache DataFu
Discovered public repositories for rjurney in the GitHub catalog.
Mirror of Apache DataFu
Azkaban scheduler re-written from the ground up.
Public repository.
Public repository.
Hadoop library for large-scale data processing
Working example of bug in Pig 0.12/0.13 re: Macros
Yelp Data Challenge Entry
Python examples of the homework examples for Andrew Ng's Stanford Machine Learning class on Coursera
Ruby utilities for metamx druid
A Realtime Chart Web Application Development with Druid
Demonstration of druid, pyDruid, Flask and d3.js
A Python connector for Druid
Mirror of Apache Whirr
Metamarkets Druid Data Store
Harness for testing Druid
GitHub Archive is a project to record the public GitHub timeline, archive it, and make it easily accessible for further analysis.
Recommender system for Github projects using the github archive data
Mirror of Apache Pig
Ron Lee's DataFu hacking
Process your tweets in Apache Hive
Working with w3c log files
Working through the nltk book
Enron Emails -> Pig ->ToJson -> RedisStorer -> Node.js
An introduction to elephant-bird on the Hortonworks Blog
Machine learning and natural language processing with Apache Pig
Twitter's collection of LZO and Protocol Buffer-related Hadoop, Pig, Hive, and HBase code.
Chapter-wise code for Agile Data the O'Reilly book
Rewrite of Avro storage functions for Pig (with Trevni support)
Example of using Pig with Accumulo on the Berkely enron emails
Hortonworks demo of Enron emails with Pig, Cassandra, Python and Flask
A library of examples showing how to use the Common Crawl corpus.
A Pig to JSON UDF for Pig that converts tuples and bags to JSON strings
Redis bulk-loader for Apache Pig
CommonCrawl Project Repository
Hortonworks tutorial on using Pig to store records in Redis and serve them with Ruby
Pig ArcFileLoader examples for loading the Common Crawl internet data
Processing Atlanta Directories from Emory University to understand the demographics of race and class in Atlanta in the Late 19th and early 20th centuries
Hortonworks demo of Enron emails using Hadoop, Pig, HBase, JRuby, Sinatra
Building a simple Node application with Pig, MongoDB, Node.js and the Enron Emails
Titanium AddressBook extensions for iOS.
Pig/ElasticSearch/Wonderdog example with the Enron Emails
Bulk loading for elastic search
Public repository.
Using HCatalog with the Enron Avro dataset
A time series serde for HIVE
Working with the Enron emails in Pig and HIVE
A platform for visualization and real-time monitoring of data workflows
Simple syntax highlighting for writing Pig scripts (http://hadoop.apache.org/pig) in Textmate.
Public repository.
Code for creating and querying an Avro encoded repository of the UC Berkeley Enron email archive