Please use this identifier to cite or link to this item:
https://www.um.edu.mt/library/oar/handle/123456789/125233| Title: | Towards a corpus of spoken Maltese : Korpus tal-Malti Mitkellem, KMM |
| Authors: | Vella, Alexandra Agius, Sarah Williams, Aiden Borg, Claudia |
| Keywords: | Natural language processing (Computer science) Transliteration Computational linguistics Translating and interpreting |
| Issue Date: | 2024-05 |
| Publisher: | ELRA and ICCL |
| Citation: | Vella, A., Agius, S., Williams, A., & Borg, C. (2024). Towards a corpus of spoken Maltese : Korpus tal-Malti Mitkellem, KMM. In Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024), pp. 16343–16352. Torino, Italia. ELRA and ICCL. |
| Abstract: | This paper presents the rationale for a “dedicated” corpus of spoken Maltese, Korpus tal-Malti Mitkellem, KMM, ‘Corpus of Spoken Maltese’, based on the concept of a gold-standard Core collection. The Core collection is designed to cater to as wide a variety of user needs as possible whilst respecting basic principles governing corpus design, such as representativeness and balance, and delivering high quality in terms of both audio quality and annotations. An overview is provided of the composition of the current Core corpus of around 20 hours of data and of the human annotation effort involved. We also carry out a small qualitative analysis of the output of a Maltese ASR system and compare it to the human annotators’ output. Initial results are promising, showing that the ASR is robust enough to generate first-pass texts for annotators to work on, thus reducing the human effort, and consequently, the cost involved. |
| URI: | https://www.um.edu.mt/library/oar/handle/123456789/125233 |
| Appears in Collections: | Scholarly Works - FacICTAI |
Files in This Item:
| File | Description | Size | Format | |
|---|---|---|---|---|
| 2024.lrec-main.1420_kmm.pdf | 570.85 kB | Adobe PDF | View/Open |
Items in OAR@UM are protected by copyright, with all rights reserved, unless otherwise indicated.
