Назад к выпуску

Optical Character Recognition on Nineteenth-Century Chagatai Manuscripts: An Error Typology

Аннотация

Digitisation programmes across Central Asian collections increasingly rely on automated transcription, yet recognition models trained on printed Arabic script perform poorly on Chagatai manuscript hands. We assembled a corpus of 1,240 annotated folios from four repositories and classified recognition failures into six recurring types, of which ligature segmentation and diacritic omission together account for 61 per cent of errors. A fine-tuned model reduced character error rate from 24.1 to 9.7 per cent. We argue that error typologies, rather than aggregate accuracy figures, should guide procurement decisions in heritage digitisation.

Ключевые слова