rxin/stream-lib
Stream summarizer and cardinality estimator.
Discovered public repositories for rxin in the GitHub catalog.
Stream summarizer and cardinality estimator.
ByteBuffer utilities using Unsafe for fast reads.
JavaScript Code Style checker
Mirror of Apache Spark
Apache Hive with patches from the Shark team
berkeley-dbgroup-site
GraphX development repository (which will eventually be merged into Apache Spark)
Examples of using Shark API
Performance tests for Spark, Shark, etc.
scalastyle
Modified benchmark from http://database.cs.brown.edu/projects/mapreduce-vs-dbms/
Public repository.
Collection of some database benchmarks, along with tools to parallelize the data generation.
Star Schema Benchmark dbgen
scripts used for ampcamp
Hive on Spark
Public repository.
Readings in Databases
Spark project template
(naive bayes and svm-based) spam detection for a proprietary social network with scalanlp
More kryo serializers
Scala framework for iterative and interactive cluster computing.
Clustering Wikipedia articles. This only contains feature generation.
Assignment 3 for the Data Science course at UC Berkeley
Assignment 1 for the Data Science course at UC Berkeley
Linear regression based sentiment scorer
Naive Bayes Sentiment Analaysis
Running TPC-H on Apache Hive
My dot files.
A distributed machine translation aligner implemented in Spark.
Public repository.