Early-Modern-OCR/hOCR-De-Noising
code to remove "noise" from hOCR output of Tesseract OCR.
Discovered public repositories for Early-Modern-OCR in the GitHub catalog.
code to remove "noise" from hOCR output of Tesseract OCR.
Scala code to correct Tesseract OCR output and generate ALTO XML and text files. Uses dictionary files, rules and a google-3gram DB to make corrections.
Java code to examine the output of Tesseract OCR and generate scores for general page quality and correctabiliby (see page-corrector repo).
Github organization page for the Early Modern OCR Project (eMOP)
Training files produced for and by the Tesseract OCR engine for work on the Early Modern OCR Project (eMOP)
Public repository.
Forked version of the Juxta Command Line tool created for eMOP by Performant Software Solutions. Will be official after eMOP is complete (10/1/14).
Part of eMOP: the Recursive Text Alignment Tool compares OCR text results to groundtruth by character and computes a score.
Part of eMOP: Franken+ tool for creating font training for Tesseract OCR engine from page images.