jwills/mrjob
Run MapReduce jobs on Hadoop or Amazon Web Services
Discovered public repositories for jwills in the GitHub catalog.
Run MapReduce jobs on Hadoop or Amazon Web Services
A prototype of Hive UDFs/UDTFs that execute nested SQL queries within rows.
Word count done with the Scrunch (Apache Crunch for Scala) library.
A Scala API for Cascading
Hive UDF's for the data warehouse
Twitter's collection of LZO and Protocol Buffer-related Hadoop, Pig, Hive, and HBase code.
Impala SDK for UDF development
A demo application for getting started with Apache Crunch.
Mirror of Apache Spark
An R package for tracking the transformations applied to the vectors in a data frame.
Simple real-time large-scale machine learning infrastructure.
Streaming MapReduce with Scalding and Storm
Lightning-fast cluster computing in Java, Scala and Python.
Either[Hotel in Austin, Prototype of a Scala Distributed Collections API]
MapReduce job for creating multitouch attribution models.
Java implementation to use with Map-Reduce
BlueSNP
RHIPE: The R and Hadoop Integrated Processing Environment
The Spatial Framework for Hadoop allows developers and data scientists to use the Hadoop data processing system for spatial data analysis.
The Cloudera Data Science Team's Machine Learning Toolkit and APIs.
Experiments w/the Play Framework to create servers for performing lookups into SequenceFiles stored in HDFS.
Utilities for converting to and from JSON from Avro records via Hadoop streaming or Hive.
A wrapper that uses the Hive AvroSerDe to deserialize data as JSON for use with Hive Streaming
A starter template for Avro + Maven projects
Mirror of Apache Avro
Clojure wrapper for Apache Crunch
Luigi is a Python module that helps you build complex pipelines of batch jobs. It handles dependency resolution, workflow management, visualization etc. It also comes with Hadoop support built in.
Record Breaker
Public repository.
GATK Official Release Repository
System for performing seismic data processing on a Hadoop cluster.
The fast and fun way to write YARN applications.
a column file format
A Scala productivity framework for Hadoop.
Mirror of Apache Flume
Mirror of Apache Hive
Classes in the new mapreduce.* API that are not part of CDH3 yet.
Hadoop library for large-scale data processing
Me messing around with some Avro stuff