Sotera/mitie-trainer
Model Training tool for MITIE
Discovered public repositories for Sotera in the GitHub catalog.
Model Training tool for MITIE
newman vm
This project is superseded by the current Datawake project but is maintained here for existing users. Browser extension and backend services aimed at enhancing Internet search with domain specific knowledge, collaboration, and analysis.
Public repository.
Quickly analyze and explore email with advanced analytics and visualization.
Stores and updates the bitcoin blockchain and historical bitcoin market data into a mysql database.
Java Client for xlang/thunderdome
Public repository.
Public repository.
Public repository.
Spark / graphX implementation of the distributed louvain modularity algorithm
Distributed Graph Analytics (DGA) is a compendium of graph analytics written for Bulk-Synchronous-Parallel (BSP) processing frameworks such as Giraph and GraphX. The analytics included are High Betweenness Set Extraction, Weakly Connected Components, Page Rank, Leaf Compression, and Louvain Modularity.
Meta information for the DARPA open catalog project.
Approximate Betweenness Centrality computation for big graph data.
For creating kml to visualize aggregate micro-path output.
Tools to mine nba data
Twitter Scraper
Meta information about the XData project
A series of analytics for creating networks from geo-temporal track data based on time/space co-occurrence. Includes UI for visualization of communities and tracks.
Vagrant-Ubuntu VM serving as a platform for XDATA performer software integration
Infer movement patterns from large amounts of geo-temporal data in a cloud environment.
A collection of common Apache Hive UDFs
An R/Hadoop Arima analytic using Rhipe to submit mapreduce jobs.
Community Detection and Compression Analytic for Big Graph Data
Public repository.
Spark implementation of the Google Correlate algorithm to quickly find highly correlated vectors in huge datasets
Public repository.
A sample project (or, rather, sample projects) to show various ways of using Zephyr - generally a good starting point for your own Zephyr implementations.
Useful classes for functions outside the scope of Zephyr's ETL, but still used in many scenarios (generally with extensive dependencies that probably shouldn't be in the core API).
Public repository.
Zephyr is a big data, platform agnostic ETL API, with Hadoop MapReduce, Storm, and other big data bindings.