jack19861225/TreeSplitWord
一个用.tire树的数据结构构成的词典.方便简单.提供了便利的词典格式.可以进行一些文本分析.
Discovered public repositories for jack19861225 in the GitHub catalog.
一个用.tire树的数据结构构成的词典.方便简单.提供了便利的词典格式.可以进行一些文本分析.
Official mirror of the Myrrix open source recommender's Subversion repository
Python based REST service which converts SQL queries into ES format. The QueryService.py is a tornado Web service which converts SQL like queries into a corresponding ElasticSearch Format. It returns a dictionary of status and result.
a SQL-like command line client for elasticsearch
Monitoring and Management Web Application for ElasticSearch instances and clusters.
A web front end for an elastic search cluster
elasticsearch中文发行版,针对中文集成了相关插件,并带有Demo,方便新手学习,或者在生产环境中直接使用
Public repository.
MySql storage engine for the cloud
Secondary Index for HBase
The CommonCrawl Crawler Engine and Related MapReduce code
Behemoth is an open source platform for large scale document analysis based on Apache Hadoop.
Katta - distributed Lucene
Hadoop Platform as a Service
forked version of Solr to support Embedded sort/filter fields in Solbase
forked version of Lucene to support Embedded sort/filter fields in Solbase
Cloudera Development Kit
Cloudera Manager API Client
hadoop test tool from generic data
HiBench is a Hadoop benchmark suite.
Storm for Yarn
Public repository.
hbase replication
NexR Hive UDFs
Public repository.
Mirror of Apache Ambari
Hue is a Web application for interacting with Apache Hadoop. It supports a file browser, job tracker interface, Hive, Pig, Impala, Oozie editors, and more.
Public repository.
Blur is an open source search engine capable of querying massive amounts of data at incredible speeds. Rather than using the flat, document-like data model used by most search solutions, Blur allows you to build rich data models and search them in a relational manner similar to querying a relational database.
open source search platform based on Lucene, Solr, HBase
RHadoop - rhadoop@revolutionanalytics.com
Mirror of Apache Hive
sbt, a build tool for Scala
simple, distributed message queue system
Library to use Kestrel as a spout within Storm
Automate Clojure projects without setting your hair on fire.
storm example for internal use
Distributed and fault-tolerant realtime computation: stream processing, continuous computation, distributed RPC, and more
WE HAVE MOVED to Apache Incubator. https://cwiki.apache.org/FLUME/ . Flume is a distributed, reliable, and available service for efficiently collecting, aggregating, and moving large amounts of log data. It has a simple and flexible architecture based on streaming data flows. It is robust and fault tolerant with tunable reliability mechanisms and many failover and recovery mechanisms. The system is centrally managed and allows for intelligent dynamic management. It uses a simple extensible data model that allows for online analytic applications.