Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Chahan Vidal-Gorène

ENC

New Results for the Text Recognition of Arabic Maghrib{ī} Manuscripts -- Managing an Under-resourced Script

Nov 29, 2022

Lucas Noëmie, Clément Salah, Chahan Vidal-Gorène

Abstract:HTR models development has become a conventional step for digital humanities projects. The performance of these models, often quite high, relies on manual transcription and numerous handwritten documents. Although the method has proven successful for Latin scripts, a similar amount of data is not yet achievable for scripts considered poorly-endowed, like Arabic scripts. In that respect, we are introducing and assessing a new modus operandi for HTR models development and fine-tuning dedicated to the Arabic Maghrib{\=i} scripts. The comparison between several state-of-the-art HTR demonstrates the relevance of a word-based neural approach specialized for Arabic, capable to achieve an error rate below 5% with only 10 pages manually transcribed. These results open new perspectives for Arabic scripts processing and more generally for poorly-endowed languages processing. This research is part of the development of RASAM dataset in partnership with the GIS MOMM and the BULAC.

Via

Access Paper or Ask Questions

Handling Heavily Abbreviated Manuscripts: HTR engines vs text normalisation approaches

Jul 07, 2021

Jean-Baptiste Camps, Chahan Vidal-Gorène, Marguerite Vernet

Figure 1 for Handling Heavily Abbreviated Manuscripts: HTR engines vs text normalisation approaches

Figure 2 for Handling Heavily Abbreviated Manuscripts: HTR engines vs text normalisation approaches

Figure 3 for Handling Heavily Abbreviated Manuscripts: HTR engines vs text normalisation approaches

Figure 4 for Handling Heavily Abbreviated Manuscripts: HTR engines vs text normalisation approaches

Abstract:Although abbreviations are fairly common in handwritten sources, particularly in medieval and modern Western manuscripts, previous research dealing with computational approaches to their expansion is scarce. Yet abbreviations present particular challenges to computational approaches such as handwritten text recognition and natural language processing tasks. Often, pre-processing ultimately aims to lead from a digitised image of the source to a normalised text, which includes expansion of the abbreviations. We explore different setups to obtain such a normalised text, either directly, by training HTR engines on normalised (i.e., expanded, disabbreviated) text, or by decomposing the process into discrete steps, each making use of specialist models for recognition, word segmentation and normalisation. The case studies considered here are drawn from the medieval Latin tradition.

* Accompanying data available at: https://zenodo.org/record/5071964

Via

Access Paper or Ask Questions