tmalaska/SparkOnALog
Examples of Integrating Spark Streaming, Flume, and HBase to solve Streaming problems
Discovered public repositories for tmalaska in the GitHub catalog.
Examples of Integrating Spark Streaming, Flume, and HBase to solve Streaming problems
This is a tool for testing and managing many repeatedly and large bulk loads on HBase
This is a layer on top of the Flume NettyAvroRpcClient that allows for multiple connects to a server.
A simple MR job where you can declare the number of mappers and reducer and a sleep time that they will sleep for.
This is a simple example to show how a single HBase "get" can retrieve the top N {items,amount} in the order of amount decresing
A teaching example of KMeans implemented with Giraph running on CDH 4.3
This will run a map only job to determine if the correct number of columns are in each row.
A simple example of using Giraph to root nodes in a tree
A upgrade Extended FairScheduler that takes Sub-Groups into account.
Tool to read many small files in HDFS with MR while control allowing the caller to define the number of mappers.
Connecting the power of the D3 graphing library to CDH (HDFS, HBase and Impala)
This is a single map reduce job that will append a unique sequence number to the front of every row in a source file.
This is a FixedLengthInputFormat for Hadoop map reduce.
Generation tool that generates DDLs and simple data load scripts.
This is a working example of how to use Flume 1.1.0 to load files into hadoop.
The rules of tera sort say you can't compress the input and output. Well those rules are out of touch with how real use cases on hadoop.
Advanced common functionality for hadoop
An analysis of adverse drug event data using Hadoop, R, and Gephi