Built something? We create video reels & spotlights for GitHub projects.Promote your project →
Home / Developers / Early-Modern-OCR
Developer Profile

Early-Modern-OCR

Discovered public repositories for Early-Modern-OCR in the GitHub catalog.

↗

Early-Modern-OCR/page-corrector

Scala code to correct Tesseract OCR output and generate ALTO XML and text files. Uses dictionary files, rules and a google-3gram DB to make corrections.

↗

Early-Modern-OCR/page-evaluator

Java code to examine the output of Tesseract OCR and generate scores for general page quality and correctabiliby (see page-corrector repo).

↗

Early-Modern-OCR/Juxta-cl

Forked version of the Juxta Command Line tool created for eMOP by Performant Software Solutions. Will be official after eMOP is complete (10/1/14).

↗

Early-Modern-OCR/RETAS

Part of eMOP: the Recursive Text Alignment Tool compares OCR text results to groundtruth by character and computes a score.

FOR MAINTAINERS

Built something? Put it in front of millions of developers.

We make a short reel about your project and post it across YouTube, Instagram, Threads, and X. Send a link, we do the rest.