rogersmarin/presto
Distributed SQL query engine for running interactive analytic queries against big data sources.
Extracts and cleans text from Wikipedia database dump and stores output in a number of files of similar size in a given directory. This is a mirror of the script by Giuseppe Attardi.
This repository is cataloged as part of our automated global GitHub synchronization. Full telemetry, velocity snapshots, and code summaries are scheduled for continuous enrichment.
Distributed SQL query engine for running interactive analytic queries against big data sources.
Elasticsearch real-time search and analytics natively integrated with Hadoop
Azkaban workflow manager.
Utility to perform feature extraction via spark-word2vec on the wikipedia (en) dataset