gdtm86/flot-downsample
Downsample plugin for Flot charts.
Discovered public repositories for gdtm86 in the GitHub catalog.
Downsample plugin for Flot charts.
Public repository.
ScalOps - A Scala DSL for large scale analytics
Oozie - workflow engine for Hadoop
HBase operations tools
Linear Algebra for Java
Utilities to use Avro files from Hadoop Map/Reduce jobs and Streaming
Distributed Graph Analytics (DGA) is a compendium of graph analytics written for Bulk-Synchronous-Parallel (BSP) processing frameworks such as Giraph and GraphX. The analytics included are High Betweenness Set Extraction, Weakly Connected Components, Page Rank, Leaf Compression, and Louvain Modularity.
HBase.MCC (HBase Multi Cluster Client). The goal is to support aways up solutions with HBase through multiple clusters
A Maven-based example of using Cloudera Impala's JDBC driver
Java client to connect directly to Impala using thrift
Distributed SQL query engine for big data
A macro-based PEG parser generator for Scala 2.10+
Example MapReduce jobs in Java, Hive, Pig, and Hadoop Streaming that work on Avro data.
Public repository.
Dockerfiles and scripts for Spark and Shark Docker images
An example of using Avro and Parquet in Spark SQL
SeqPig is a library for Apache Pig for the distributed analysis of large sequencing datasets. It provides import and export functions for file formats commonly used for sequencing data, as well as a collection of Pig user-defined-functions (UDF’s) to help process aligned and unaligned sequence data.
Scala DSL on top of Oozie XML
The Hadoop GP Toolbox provides tools to exchange features between a Geodatabase and Hadoop and run Hadoop workflow jobs.
The GIS Tools for Hadoop are a collection of GIS tools for spatial analysis of big data.
Example project to show how to use Spark to read and write Avro/Parquet files
A Thrift parser/generator
TPC-DS Kit for Impala
A simple integer compression library in Java
A platform for visualization and real-time monitoring of data workflows
Mirror of Apache Parquet
As we are moving to Apache, please open your pull requests on: https://github.com/apache/incubator-parquet-format
Open Source Code of Conduct at Twitter
The fast and fun way to write YARN applications.
Oryx 2 (incubating): Lambda architecture on Spark for real-time large scale machine learning
Simple real-time large-scale machine learning infrastructure.
Real-time Query for Hadoop
Python client and Numba-based UDFs for Impala
Pig on Spark Streaming
Spark GCE Script Helps you deploy Spark cluster on Google Cloud.
Pig on Apache Spark
Miscellaneous short shell scripts.
Community-driven snippets for presentations in LaTeX, Beamer, and TikZ.
A C++ client library for Apache Kafka v0.8+. Also includes C API.
Automates Spark standalone cluster tasks with Puppet and Fabric.
Example Spark project using Parquet as a columnar store with Thrift objects.
Async Scala-Akka-Netty based Stress Tool
Learning and Using ØMQ
Pandas-like DataFrame DSL on Spark
I simple API to interact with HBase with Spark
azure-kubernetes-visualizer
Mirror of Apache HBase
Mirror of Apache Phoenix
Utilities for converting to and from JSON from Avro records via Hadoop streaming or Hive.