tmalaska/spark
Mirror of Apache Spark
Just some example of using GraphX
This repository is cataloged as part of our automated global GitHub synchronization. Full telemetry, velocity snapshots, and code summaries are scheduled for continuous enrichment.
Mirror of Apache Spark
This is an example of how to make Unique Sequences in a distributed way with Spark (No dups, No Skips)
This will do a Merge Join of absolute Sorted data any number of files of ether side.
Using JRecord to build a mapred and mapreduce inputformat for HDFS, MAPREDUCE, PIG, HIVE, Spark, ...