yuanke/cascading
Cascading is a feature rich API for defining and executing complex and fault tolerant data processing flows on a Hadoop cluster. See https://github.com/Cascading/cascading for the release repository.
Discovered public repositories for yuanke in the GitHub catalog.
Cascading is a feature rich API for defining and executing complex and fault tolerant data processing flows on a Hadoop cluster. See https://github.com/Cascading/cascading for the release repository.
cn-clojure聚会资料
Sparrow scheduling platform (U.C. Berkeley).
Public repository.
ADFS (Ali Distributed File System) is an evolutional version of Hadoop which delivers high availability, auick-restart and other features.
first use sbt and scala to write spark application
Public repository.
Open Source In-Memory Data Grid
Muppet
Netty project - an event-driven asynchronous network application framework
avro format for mapreduce
column storage on hadoop
Data and example code for Programming Pig, by Alan F. Gates
ZooKeeper client wrapper and rich ZooKeeper framework
Hive UDF's for the data warehouse
Hbase demonstration code as an addendum to the presentation
Public repository.
Storm for Yarn
This package provides an efficient implementation of locality-sensitve hashing (LSH)
R functions for fitting latent factor models with internal computation in C/C++
Public repository.
Patched, refactored version of code.google.com/hadoop-gpl-compression for hadoop 0.20
ORC File working repo
A rough prototype of a tool for discovering Apache Hive schemas from JSON documents.
Public repository.
Public repository.
A starter project for using the splout-hadoop API (Splout SQL)
Example project for building apps with Pangool
Pangool-Flow is an experimental module on top of Pangool (http://pangool.net) which adds automatic flow building and management, parallel execution and high-level constructs.
Tuple MapReduce for Hadoop: Hadoop API made easy
Public repository.
Samples for working with the Salesforce Mobile SDK
A web-latency SQL spout for Hadoop.
Public repository.
A list of helpful front-end related questions you can use to interview potential candidates, test yourself or completely ignore.
Pure Java Rpm Library
Public repository.
Multidimensional data storage with rollups for numerical data
A distributed publish/subscribe messaging service
A simple BFS exploration of Last FM's social graph
Akka Project
Hadopp Module For Puppet.
Mysql Puppet Module
HBase-based BI "OLAP-ish" solution
Mirror of Apache Cassandra (incubating)
Redis is an in-memory database that persists on disk. The data model is key-value, but many different kind of values are supported: Strings, Lists, Sets, Sorted Sets, Hashes
Redis Python Client
an online request replication tool, fit for online testing, stress testing, performance evaluation,etc
Puppet recipes for deploying Hstack
An indexing library for HBASE: repo created from the work of LilyCMS