zeeshanali/spark-workshop
Data and code for "Fast Data Applications with Spark and Python"
Interactive Scala REPL in a browser
This repository is cataloged as part of our automated global GitHub synchronization. Full telemetry, velocity snapshots, and code summaries are scheduled for continuous enrichment.
Data and code for "Fast Data Applications with Spark and Python"
Automated Installation of Hadoop 2.4.0, Scala 2.10.3, Spark 1.1.0 and sparkR (Ubuntu 12.04 +)
SeqPig is a library for Apache Pig for the distributed analysis of large sequencing datasets. It provides import and export functions for file formats commonly used for sequencing data, as well as a collection of Pig user-defined-functions (UDF’s) to help process aligned and unaligned sequence data.
Hadoop-BAM is a Java library for the manipulation of files in common bioinformatics formats using the Hadoop MapReduce framework with the Picard SAM JDK, and command line tools similar to SAMtools.