Post-correction of Historical Text Transcripts with Large Language Models: An Exploratory Study

Frédéric Kaplan, Maud Ehrmann, Matteo Romanello, Emanuela Boros, Sven-Nicolas Yoann Najem
2024
Article de conférence

Résumé

The quality of automatic transcription of heritage documents, whether from printed, manuscripts or audio sources, has a decisive impact on the ability to search and process historical texts. Although significant progress has been made in text recognition (OCR, HTR, ASR), textual materials derived from library and archive collections remain largely erroneous and noisy. Effective post-transcription correction methods are therefore necessary and have been intensively researched for many years. As large language models (LLMs) have recently shown exceptional performances in a variety of text-related tasks, we investigate their ability to amend poor historical transcriptions. We evaluate fourteen foundation language models against various post-correction benchmarks comprising different languages, time periods and document types, as well as different transcription quality and origins. We compare the performance of different model sizes and different prompts of increasing complexity in zero and few-shot settings. Our evaluation shows that LLMs are anything but efficient at this task. Quantitative and qualitative analyses of results allow us to share valuable insights for future work on post-correcting historical texts with LLMs.

Source officielle

https://infoscience.epfl.ch/record/307961?ln=fr

À propos de ce résultat

Cette page est générée automatiquement et peut contenir des informations qui ne sont pas correctes, complètes, à jour ou pertinentes par rapport à votre recherche. Il en va de même pour toutes les autres pages de ce site. Veillez à vérifier les informations auprès des sources officielles de l'EPFL.

Post-correction of Historical Text Transcripts with Large Language Models: An Exploratory Study

Graph Chatbot

Chattez avec Graph Search

Graph generative deep learning models with an application to circuit topologies

Hybrid ground-state quantum algorithms based on neural Schrödinger forging

Inhalation of Microplastics—A Toxicological Complexity

Graph generative deep learning models with an application to circuit topologies

Hybrid ground-state quantum algorithms based on neural Schrödinger forging

Inhalation of Microplastics—A Toxicological Complexity