westei/n3-collection
N3 - A Collection of Datasets for Named Entity Recognition and Disambiguation in the NLP Interchange Format
Public repository discovered through GitHub real-time crawl.
This repository is cataloged as part of our automated global GitHub synchronization. Full telemetry, velocity snapshots, and code summaries are scheduled for continuous enrichment.
N3 - A Collection of Datasets for Named Entity Recognition and Disambiguation in the NLP Interchange Format
working repository for experimenting with idea of providing a commons library for RDF 1.1 that could be implemented by the upcoming versions of the main Java toolkits
TBD enhancement engine uses Sphinix library to convert the captured audio. Media (audio/video) data file is parsed with the ContentItem and formatted to proper audio format by Xuggler libraries. Audio speech is than extracted by Sphinix to 'plain/text' with the annotation of temporal position of the extracted text. Sphinix uses acoustic model and language model to map the utterances with the text, so the engine will also provide support of uploading acoustic model and language model.
A text tagger based on Lucene / Solr, using FST technology