ABBYY, Tess and Claude

Lessons learned from digitising printed Old and Middle Irish – Jane Lawson, University of Sydney · Asynchronous talk · 16 minutes · Captions available

Taking part. Watch at any time, then bring your questions, comments and disagreements to the discussion thread below. I’ll be checking in daily until the symposium, twice a day on 28–29 September, and the thread stays open until 6 October. First-time commenters are held briefly for approval.

About this talk. While many secondary texts in Celtic Studies are digitised and available online, poor quality optical character recognition of Celtic languages results limits their accessibility. Searches for specific language terms result in few – or even zero – hits due to recognition errors, and corpus linguistic techniques are impossible to apply without full scale transcription efforts.

Many modern off-the-shelf digitisation programs offer medieval language support for languages such as Latin and Old English, or modern dictionaries of Breton, Irish, Welsh and Scottish Gaelic. However, for the reader of their medieval equivalents, there is little available.

This paper discusses a low-cost ‘ensemble-OCR’ approach to digitising Old and Middle Irish law texts, adaptable to other medieval Celtic languages and genres with few linguistic resources.  An ensemble OCR approach uses two or more OCR programs to produce text transcriptions, then compares the two transcriptions to identify potential errors. The comparison can be augmented through a range of techniques, including voting or weighted models, and the analysis and application of systemic errors unique to each program. AI support was employed to develop bespoke coding scripts that identified and corrected errors, and tools to compare images with transcripts for human-in-the loop review.

I will discuss the many lessons learned from challenges that arose in each stage of the project, including generating and selecting images, setting up OCR engines for an ensemble approach, analysis and quality assurance of the final product. I will share what worked (and more importantly: what didn’t), highlight opportunities for further improvement, and share some of the project resources with the research community.  

Resources. Links to the resources mentioned in the talk:

  • OCR Tip Sheet: docx or pdf
  • O’Davoren’s Glossary (Stokes, 1904): Searchable pdf (approx. 75 MB)
  • O’Davoren’s Glossary (Stokes, 1904): Image and OCR text on same page (approx. 75 MB)
  • Combined wordlist: A text file containing lemmas and variants from MOLOR, CorPH, Goidelex and constructed from the OCR version of O’Davoren’s Glossary (Stokes, 1904). Suitable for use in ABBYY as a custom dictionary.
  • O’Davoren wordlist: Building your own law or glossary-specific wordlist? A text file containing the unique words from the OCR version of O’Davoren’s Glossary (Stokes, 1904). Suitable for use in ABBYY as a custom dictionary.

Questions outside the thread: celticlanguageresearch@gmail.com

Leave a Reply

Discover more from Celtic Languages Symposium

Subscribe now to keep reading and get access to the full archive.

Continue reading