Early-Modern-OCR/hOCR-De-Noising
code to remove "noise" from hOCR output of Tesseract OCR.
Training files produced for and by the Tesseract OCR engine for work on the Early Modern OCR Project (eMOP)
This repository is cataloged as part of our automated global GitHub synchronization. Full telemetry, velocity snapshots, and code summaries are scheduled for continuous enrichment.
code to remove "noise" from hOCR output of Tesseract OCR.
Scala code to correct Tesseract OCR output and generate ALTO XML and text files. Uses dictionary files, rules and a google-3gram DB to make corrections.
Java code to examine the output of Tesseract OCR and generate scores for general page quality and correctabiliby (see page-corrector repo).
Github organization page for the Early Modern OCR Project (eMOP)