xuyanhui/free-programming-books-zh_CN
免费的计算机编程类中文书籍,欢迎投稿
Discovered public repositories for xuyanhui in the GitHub catalog.
免费的计算机编程类中文书籍,欢迎投稿
大数据/数据挖掘/推荐系统/机器学习相关资源
tsocks 1.8 with the mac osx patch
A Powerful Spider System with Web UI
Self-written notes that may be useful
Notes talking about the design and implementation of Apache Spark
A Maven-based example of using Cloudera Impala's JDBC driver
A Spark WordCountJob example as a standalone SBT project with Specs2 tests, runnable on Amazon EMR
A curated list of awesome big data frameworks, ressources and other awesomeness.
A Typesafe Activator tutorial for Apache Spark.
A simple Java API and command line interface for importing, managing and retrieving data from HBase.
Big Cloud Bulk Synchronous Parallel
Minos is beyond a hadoop deployment system.
Apache hadoop management system
Secondary Index for HBase
Mirror of Apache Lucene & Solr
Mirror of Apache Solr
结巴中文分词
compatibility tests to make sur C and Java implementations can read each other
The Parquet site.
Columnar file format for hadoop
Java readers/writers for Parquet columnar file formats to use with Map-Reduce
Distributed SQL query engine for big data
Hoop, Hadoop HDFS over HTTP
my personal learning code
elasticsearch中文发行版,针对中文集成了相关插件,并带有Demo,方便新手学习,或者在生产环境中直接使用
A web front end for an elastic search cluster
CSV river for ElasticSearch
Kibana Dashboard Preview
A log analyzing web interface for logstash and elasticsearch. More info at http://www.kibana.org
An import river similar to the elasticsearch mysql river
logstash - logs/event transport, processing, management, search.
Mirror of Apache Kafka
Cobub Razor - Open Source Mobile Analytics Solution
Hadoop Platform as a Service
Read and write data to/from ElasticSearch within Hadoop
Kettle plugin that provides support for interacting within many "big data" projects including Hadoop, Hive, HBase, Cassandra, MongoDB, and others.
A library of examples showing how to use the Common Crawl corpus.
The CommonCrawl Crawler Engine and Related MapReduce code
Public repository.
distributed realtime searchable database
Job scheduler
Behemoth is an open source platform for large scale document analysis based on Apache Hadoop.
Fabric scripts to install Solr 4 in distributed mode (SolrCloud)
RHadoop - rhadoop@revolutionanalytics.com
forked version of Solr to support Embedded sort/filter fields in Solbase
forked version of Lucene to support Embedded sort/filter fields in Solbase
open source search platform based on Lucene, Solr, HBase
Alerts for the 21st century #hubspot-open-source
Tiny bootstrap-compatible WISWYG rich text editor