OCR Studio & Controls

No file chosen
100%

Extracted Text

Source details

Saved with every page of this file in all exports, so the text can be cited, credited and used with the right permissions. The page number is added automatically.

Save this page

Collection

Pages you add are kept on this device. Export them together as one corpus: every page's text, images, recordings, source details and training lines, plus a combined corpus.txt.

No pages yet. Recognise and correct a page, then press Add.

Language models

Bengali, Assamese and English are built in and work offline. You can add a model trained for your own language (a Tesseract .traineddata file, optionally .gz), for example one trained from the training lines this app exports.