kraken
Turn-key OCR engine trained for historical and non-Latin script documents, not modern print.
Desktopfree
Kraken is an open-source, turn-key OCR engine built for historical and non-Latin script material, where mainstream OCR tools tend to fail. It offers fully trainable layout analysis, reading-order detection, and character recognition, with support for right-to-left, bidirectional, and top-to-bottom scripts and output in ALTO, PageXML, and hOCR formats. It's aimed at digital humanities researchers and archivists digitizing historical or multilingual document collections.
Categories
OCRdocument processingdigital humanities
Something wrong with this listing — dead link, not a real product, wrong info?

