Hierarchical Text Segmentation for Medieval Manuscripts - Ecole Centrale de Nantes Accéder directement au contenu
Communication Dans Un Congrès Année : 2020

Hierarchical Text Segmentation for Medieval Manuscripts

Résumé

In this paper, we address the segmentation of books of hours, Latin devotional manuscripts of the late Middle Ages, that exhibit challenging issues: a complex hierarchical entangled structure, variable content, noisy transcriptions with no sentence markers, and strong correlations between sections for which topical information is no longer sufficient to draw segmentation boundaries. We show that the main state-of-the-art segmentation methods are either inefficient or inapplicable for books of hours and propose a bottom-up greedy approach that considerably enhances the segmentation results. We stress the importance of such hierarchical segmentation of books of hours for historians to explore their overarching differences underlying conception about Church.
Fichier principal
Vignette du fichier
2020.coling-main.549.pdf (2.46 Mo) Télécharger le fichier
Origine : Fichiers éditeurs autorisés sur une archive ouverte

Dates et versions

hal-03100170 , version 1 (06-01-2021)

Licence

Paternité

Identifiants

  • HAL Id : hal-03100170 , version 1

Citer

Amir Hazem, Béatrice Daille, Louis Chevalier, Dominique Stutzmann, Christopher Kermorvant. Hierarchical Text Segmentation for Medieval Manuscripts. COLING'2020 The 28th International Conference on Computational Linguistics, Dec 2020, Barcelona, Spain. pp.6240-6251. ⟨hal-03100170⟩
510 Consultations
106 Téléchargements

Partager

Gmail Facebook X LinkedIn More