tedunderwood/pagelevelHMM
Java code I used to train hidden Markov models on top of page-level classification. Weka is a dependency. Needs refactoring.
Discovered public repositories for tedunderwood in the GitHub catalog.
Java code I used to train hidden Markov models on top of page-level classification. Weka is a dependency. Needs refactoring.
Java code that uses existing metadata to train classifiers that then make predictions for cases where metadata is missing / suspected.
Contains Java code for a page-tagging interface.
Code and data to support a collaboration between Andrew Goldstone and Ted Underwood.
Code and documentation associated with "Understanding Genre in a Collection of a Million Volumes"
Files for eMOP.
Public repository.
Scripts that clean up OCR and munge Hathi metadata.
Python scripts used to wrangle collection from Hathi, mostly on a cluster.
Data for 1924-2006 pmla model, plus scripts to turn into Gephi network.
Java package that partitions a corpus and runs LDA in parallel on it
Public repository.
Python scripts for collating HathiTrust page files.
Public repository.
Python modules that evaluate OCR quality.
A Java package that does basic LDA, without hyperparameter optimization. Folder settings are local. Ymmv.
folder storing current rulesets, scripts, and metadata for tokenizing / collection building
Python scripts for tokenizing text files
R scripts that browse the results of LDA