zeeshanali/spark-workshop
Data and code for "Fast Data Applications with Spark and Python"
Docker containers for the IPython notebook (+SciPy Stack)
This repository is cataloged as part of our automated global GitHub synchronization. Full telemetry, velocity snapshots, and code summaries are scheduled for continuous enrichment.
Data and code for "Fast Data Applications with Spark and Python"
Interactive Scala REPL in a browser
Automated Installation of Hadoop 2.4.0, Scala 2.10.3, Spark 1.1.0 and sparkR (Ubuntu 12.04 +)
SeqPig is a library for Apache Pig for the distributed analysis of large sequencing datasets. It provides import and export functions for file formats commonly used for sequencing data, as well as a collection of Pig user-defined-functions (UDF’s) to help process aligned and unaligned sequence data.