dgkris/XPathExtractor
Extracts contents from all the html files in a given folder using XPath. Ideal when there is large amounts of crawled content
An attempt to measure the sentiment of each media publication to a set of predefined entities.
This repository is cataloged as part of our automated global GitHub synchronization. Full telemetry, velocity snapshots, and code summaries are scheduled for continuous enrichment.
Extracts contents from all the html files in a given folder using XPath. Ideal when there is large amounts of crawled content
An attempt to create a repository of all the machine learning algorithms
Java clone for python term extractor topia.termextract
An RSS feed importer using apache flume. Tried and tested on CDH 4.6.