Early-Modern-OCR/page-corrector
Scala code to correct Tesseract OCR output and generate ALTO XML and text files. Uses dictionary files, rules and a google-3gram DB to make corrections.
Github organization page for the Early Modern OCR Project (eMOP)
This repository is cataloged as part of our automated global GitHub synchronization. Full telemetry, velocity snapshots, and code summaries are scheduled for continuous enrichment.
Scala code to correct Tesseract OCR output and generate ALTO XML and text files. Uses dictionary files, rules and a google-3gram DB to make corrections.
Java code to examine the output of Tesseract OCR and generate scores for general page quality and correctabiliby (see page-corrector repo).
Training files produced for and by the Tesseract OCR engine for work on the Early Modern OCR Project (eMOP)
Public repository.