Please use this identifier to cite or link to this item: https://www.um.edu.mt/library/oar/handle/123456789/125233
Title: Towards a corpus of spoken Maltese : Korpus tal-Malti Mitkellem, KMM
Authors: Vella, Alexandra
Agius, Sarah
Williams, Aiden
Borg, Claudia
Keywords: Natural language processing (Computer science)
Transliteration
Computational linguistics
Translating and interpreting
Issue Date: 2024-05
Publisher: ELRA and ICCL
Citation: Vella, A., Agius, S., Williams, A., & Borg, C. (2024). Towards a corpus of spoken Maltese : Korpus tal-Malti Mitkellem, KMM. In Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024), pp. 16343–16352. Torino, Italia. ELRA and ICCL.
Abstract: This paper presents the rationale for a “dedicated” corpus of spoken Maltese, Korpus tal-Malti Mitkellem, KMM, ‘Corpus of Spoken Maltese’, based on the concept of a gold-standard Core collection. The Core collection is designed to cater to as wide a variety of user needs as possible whilst respecting basic principles governing corpus design, such as representativeness and balance, and delivering high quality in terms of both audio quality and annotations. An overview is provided of the composition of the current Core corpus of around 20 hours of data and of the human annotation effort involved. We also carry out a small qualitative analysis of the output of a Maltese ASR system and compare it to the human annotators’ output. Initial results are promising, showing that the ASR is robust enough to generate first-pass texts for annotators to work on, thus reducing the human effort, and consequently, the cost involved.
URI: https://www.um.edu.mt/library/oar/handle/123456789/125233
Appears in Collections:Scholarly Works - FacICTAI

Files in This Item:
File Description SizeFormat 
2024.lrec-main.1420_kmm.pdf570.85 kBAdobe PDFView/Open


Items in OAR@UM are protected by copyright, with all rights reserved, unless otherwise indicated.