hodgesz/kafka-blur-consumer
Real Time Kafka Consumer to index Kafka messages into Apache Blur
Discovered public repositories for hodgesz in the GitHub catalog.
Real Time Kafka Consumer to index Kafka messages into Apache Blur
Scalable machine learning library running on Hive/Hadoop
Flexible and powerful data analysis / manipulation library for Python, providing labeled data structures similar to R data.frame objects, statistical functions, and much more
Mortar Development Framework
Flux Capacitor is a Java-based distributed application demonstrating many of the key Netflix Open Source components.
ADL's Open Source Learning Record Store (LRS) is used to store learning data collected with the Experience API.
Source-agnostic distributed change data capture system
A fork of cascading patterns, but implemented for trident
Optimistic punt: create redshift tables and sink data via S3
Enterprise-strength marketing and product analytics, powered by Hadoop, Hive and Redshift
Tool which generates Avro schemas and Java bindings from XML schemas.
Jackson dataformat module to support Avro-encoded data
Azkaban scheduler re-written from the ground up.
CQL Binary Protocol .NET client for Apache Cassandra
Rewrite of Avro storage functions for Pig (with Trevni support)
Tools for keeping your cloud operating in top form. Chaos Monkey is a resiliency tool that helps applications tolerate random instance failures.
Co-Process for backup/recovery, Token Management, and Centralized Configuration management for Cassandra.
Public repository.
A collection of spouts, bolts, serializers, DSLs, and other goodies to use with Storm
Capturing JVM- and application-level metrics. So you know what's going on.
Hastur's server
Hastur's Ruby client
Automated deploy for Kafka on AWS
Cassandra state implementation for Twitter Storm Trident API
NOTE: This project has been moved into storm-kafka in storm-contrib
Twitter's collection of LZO and Protocol Buffer-related Hadoop, Pig, Hive, and HBase code.
Experiments with Storm
A Storm spout that tails a file (or collection of files)
A Kafka-like interface for accepting high-volume traffic and writing to topic files or directories with various strategies.