dvryaboy/awesome-bigdata
A curated list of awesome big data frameworks, ressources and other awesomeness.
Discovered public repositories for dvryaboy in the GitHub catalog.
A curated list of awesome big data frameworks, ressources and other awesomeness.
Mirror of Apache Parquet
As we are moving to Apache, please open your pull requests on: https://github.com/apache/incubator-parquet-format
Mirror of Apache Parquet
Java library relying on semver.org principles to check binary code compatibility
Apache Incubator Proposal for Parquet Format
Cascading is a feature rich API for defining and executing complex and fault tolerant data processing workflows on a Hadoop cluster.
A Scala API for Cascading
an anagram
source examples to support the "Cascading for the Impatient" blog post series
Vertica Hadoop Connector
Elephant Twin LZO uses Elephant Twin to create LZO block indexes
Elephant Twin is a framework for creating indexes in Hadoop
Patched, refactored version of code.google.com/hadoop-gpl-compression for hadoop 0.20
Mirror of Apache Giraph
Eclipse plugin for Apache Pig
Prototype Bud runtime (Bloom Under Development)
Flume is a distributed, reliable, and available service for efficiently collecting, aggregating, and moving large amounts of log data. It has a simple and flexible architecture based on streaming data flows. It is robust and fault tolerant with tunable reliability mechanisms and many failover and recovery mechanisms. The system is centrally managed and allows for intelligent dynamic management. It uses a simple extensible data model that allows for online analytic applications.
Scribe is a server for aggregating log data streamed in real time from a large number of servers. It is designed to be scalable, extensible without client-side modification, and robust to failure of the network or any specific machine.
Twitter's collection of LZO and Protocol Buffer-related Hadoop, Pig, and HBase code.
Mirror of Apache Pig
PigLatin mode for Emacs.