Built something? We create video reels & spotlights for GitHub projects.Promote your project →
Catalog / Early-Modern-OCR / TesseractTraining
Public GitHub Catalog Discovered Oct 2, 2026

Early-Modern-OCR / TesseractTraining

Training files produced for and by the Tesseract OCR engine for work on the Early Modern OCR Project (eMOP)

View repository on GitHub ↗ View creator profile Browse directory

About this discovery

This repository is cataloged as part of our automated global GitHub synchronization. Full telemetry, velocity snapshots, and code summaries are scheduled for continuous enrichment.

#24719950GitHub System ID
Early-Modern-OCROrganization / User
PublicVisibility
ActiveCatalog Status

More from Early-Modern-OCR

↗

Early-Modern-OCR/page-corrector

Scala code to correct Tesseract OCR output and generate ALTO XML and text files. Uses dictionary files, rules and a google-3gram DB to make corrections.

Discovered
↗

Early-Modern-OCR/page-evaluator

Java code to examine the output of Tesseract OCR and generate scores for general page quality and correctabiliby (see page-corrector repo).

Discovered
FOR MAINTAINERS

Built something? Put it in front of millions of developers.

We make a short reel about your project and post it across YouTube, Instagram, Threads, and X. Send a link, we do the rest.