Kraken

Turn-key OCR system for historical and non-Latin script documents.

Open SourceSelf HostedOffline Capable
0.0 (0)

About

Historical documents and non-Latin scripts are the focus of Kraken, a turn key OCR system developed at the Ecole Pratique des Hautes Etudes in Paris. It provides fully trainable layout analysis, reading order detection, and character recognition, with support for right-to-left, bidirectional, and top-to-bottom writing such as Arabic and Hebrew. Recognition output comes in ALTO, PageXML, abbyyXML, and hOCR formats with word bounding boxes and character cuts. Pretrained models are shared through a public repository on Zenodo, and users can train their own segmentation and recognition models on custom material. Kraken installs via pip on Linux and macOS, including ARM machines, and pairs closely with eScriptorium, a web interface for annotating data, training models, and running transcription. Digital humanities researchers, libraries, and archives digitizing manuscripts and early printed books are its main users. Released under the Apache 2.0 license.

Reviews (0)

Leave a Review

No reviews yet. Be the first to review!

Details

Price
Free
Platform
Local/Desktop
Difficulty
Intermediate (3/5)
License
Apache-2.0
Added
Apr 3, 2026

Related Tools

Featured

Document parsing library by IBM for converting PDFs and documents to structured data.

Open SourceSelf HostedOffline
Easy
0.0 (0)

Deep learning based OCR library in Python and TensorFlow/PyTorch.

Open SourceSelf HostedOffline
Easy
0.0 (0)

One-stop tool for high-quality PDF extraction to Markdown or JSON.

Open SourceSelf HostedOffline
Easy
0.0 (0)

Python bindings for MuPDF library for fast PDF text and image extraction.

Open SourceSelf HostedOffline
Beginner
0.0 (0)

Tool for extracting tables from PDF files into CSV or DataFrame format.

Open SourceSelf HostedOffline
Beginner
0.0 (0)

Python library for extracting tables from PDF files.

Open SourceSelf HostedOffline
Beginner
0.0 (0)
Browse all OCR & Document Processing tools