Matchboxmatchbox
← Back to match

kraken

Turn-key OCR engine trained for historical and non-Latin script documents, not modern print.

Desktopfree

Kraken is an open-source, turn-key OCR engine built for historical and non-Latin script material, where mainstream OCR tools tend to fail. It offers fully trainable layout analysis, reading-order detection, and character recognition, with support for right-to-left, bidirectional, and top-to-bottom scripts and output in ALTO, PageXML, and hOCR formats. It's aimed at digital humanities researchers and archivists digitizing historical or multilingual document collections.

Categories
OCRdocument processingdigital humanities

Full match profile

Behind the summary, Matchbox keeps a richer profile of kraken - the signals our matcher actually reads to decide when to surface it. It stays private; claim the listing to see and control it.

  • Problem & pain-point mapping
  • Who we surface it to (audience fit)
  • What it's a strong alternative to
  • Trust & credibility signals

Try Matchbox with your own problem

Describe what is not working - we’ll show you whether kraken (or something else) actually fits.