ai-ku/langvis-resources
Supporting code and data for the langvis project.
Discovered public repositories for ai-ku in the GitHub catalog.
Supporting code and data for the langvis project.
Word vectors
Demo run of the S-CODE algorithm on a 3D-Sphere
Unsupervised multilingual part of speech induction system (2014 version)
glookup - reads ngram patterns with wildcards from stdin and prints their counts from the Web1T Google ngram data.
Public repository.
Public repository.
Run SRILM with different options to find the best language model given the training and test data.
Unsupervised word sense disambiguation
CONNL-X Turkish data set of upos repository
CONNL-X Swedish data set of upos repository
CONNL-X Spanish data set of upos repository
CONNL-X Slovene data set of upos repository
CONNL-X Portuguese data set of upos repository
CONNL-X German data set of upos repository
CONNL-X Dutch data set of upos repository
CONNL-X Danish data set of upos repository
CONNL-X Czech data set of upos repository
Multext East Hungarian data set of upos repository
CONNL-X Bulgarian data set of upos repository
Multext-East Serbian data set of upos repository
Multext-East Slovene data set of upos repository
Multext-East Estonian data set of upos repository
Multext-East Czech data set of upos repository
Multext-East Bulgarian data set of upos repository
Multext-East Romanian data set of upos repository
Multext East English data set of upos repository
WSJ English data set of upos repository
Semeval 2013 | Task 13 WSI and WSD
Word Sense Induction
Paradigmatic approach to Childes child data.
Public repository.
Mex file to read large sparse files to Matlab
Feature Decay Algorithm
Public repository.
Protein dynamics research.
Unsupervised part of speech induction.
Calculates a variety of distances between vectors.
Unsupervised parsing techniques for natural languages
k-means algorithm with (optional) instance weights.
Generate most likely substitutes for words in a given text based on an n-gram language model.
Sphere embedding (s-code) is a variation of Euclidean embedding of co-occurence data (code).