Enhancing the PARSEME Turkish Corpus of Verbal Multiword Expressions - Laboratoire d'Informatique pour la Mécanique et les Sciences de l'Ingénieur Accéder directement au contenu
Communication Dans Un Congrès Année : 2022

Enhancing the PARSEME Turkish Corpus of Verbal Multiword Expressions

Résumé

The PARSEME (Parsing and Multiword Expressions) project proposes multilingual corpora annotated for multiword expressions (MWEs). In this case study, we focus on the Turkish corpus of PARSEME. Turkish is an agglutinative language and shows high inflection and derivation in word forms. This can cause some issues in terms of automatic morphosyntactic annotation. We provide an overview of the problems observed in the morphosyntactic annotation of the Turkish PARSEME corpus. These issues are mostly observed on the lemmas, which is important for the approximation of a type of an MWE. We propose modifications of the original corpus with some enhancements on the lemmas and parts of speech. The enhancements are then evaluated with an identification system from the PARSEME Shared Task 1.2 to detect MWEs, namely Seen2Seen. Results show increase in the F-measure for MWE identification, emphasizing the necessity of robust morphosyntactic annotation for MWE processing, especially for languages that show high surface variability.
Fichier principal
Vignette du fichier
2022-OZTURK-et-al-MWE-Turkish-on-ACL-anthology.pdf (177.47 Ko) Télécharger le fichier
Origine : Fichiers éditeurs autorisés sur une archive ouverte

Dates et versions

hal-03925083 , version 1 (05-01-2023)

Identifiants

  • HAL Id : hal-03925083 , version 1

Citer

Yagmur Ozturk, Najet Hadj Mohamed, Adam Lion-Bouton, Agata Savary. Enhancing the PARSEME Turkish Corpus of Verbal Multiword Expressions. 18th Workshop on Multiword Expressions (MWE 2022) @LREC2022, Jun 2022, Marseille, France. ⟨hal-03925083⟩
122 Consultations
18 Téléchargements

Partager

Gmail Facebook X LinkedIn More