NWU Institutional Repository

Pre-training a Transformer-Based Generative Model Using a Small Sepedi Dataset

dc.contributor.authorRamalepe, Simonen.ZA
dc.contributor.authorModipa, Thipe Ien.ZA
dc.contributor.authorDavel, Marelie Hen.ZA
dc.date.accessioned2025-10-21T08:25:10Zen.ZA
dc.date.issued2024en.ZA
dc.descriptionJournal Article, Faculty of Engineering, Multilingual Speech Technologies (MUST)-- Potchefstroom Campusen.ZA
dc.description.abstractDue to the scarcity of data in low-resourced languages, the development of language models for these languages has been very slow. Currently, pre-trained language models have gained popularity in natural language processing, especially, in developing domain-specific models for low-resourced languages. In this study, we experiment with the impact of using occlusion-based techniques when training a language model for a text generation task. We curate 2 new datasets, the Sepedi monolingual (SepMono) dataset from several South African resources and the Sepedi radio news (SepNews) dataset from the radio news domain. We use the SepMono dataset to pre-train transformer-based models using the occlusion and non-occlusion pre-training techniques and compare performance. The SepNews dataset is specifically used for fine-tuning. Our results show that the non-occlusion models perform better compared to the occlusion-based models when measuring validation loss and perplexity. However, analysis of the generated text using the BLEU score metric, which measures the quality of the generated text, shows a slightly higher BLEU score for the occlusion-based models compared to the nonocclusion models.en.ZA
dc.description.sponsorshipAcknowledgments We would like to acknowledge the Telkom Centre of Excellence for Speech Technology at the University of Limpopo and the MUST deep learning research group at the Northwest University (Potchefstroom) for their continued support. This work is based on research supported in part by the National Research Foundation of South Africa (Ref Number RA211019646111).en.ZA
dc.identifier.citationDavel, Marelie H. et al. 2024. Pre-training a Transformer-Based Generative Model Using a Small Sepedi Dataset. Artificial Intelligence Research. SACAIR 2024. Communications in Computer and Information Science, Volume 2326. Springer, Cham, (2024), [https://doi.org/10.1007/978-3-031-78255-8_19]en.ZA
dc.identifier.urihttps://doi.org/10.1007/978-3-031-78255-8_19en.ZA
dc.identifier.urihttp://hdl.handle.net/10394/43718en.ZA
dc.language.isoenen.ZA
dc.publisherCommunications in Computer and Information Scienceen.ZA
dc.subjectTransformersen.ZA
dc.subjectText Generationen.ZA
dc.subjectPre-Trainingen.ZA
dc.subjectOcclusion Based Trainingen.ZA
dc.subjectDatasetsen.ZA
dc.titlePre-training a Transformer-Based Generative Model Using a Small Sepedi Dataseten.ZA
dc.typeArticleen.ZA

Files

Original bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
Davel, Marelie H. et al. 2024.pdf
Size:
1.15 MB
Format:
Adobe Portable Document Format

License bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
license.txt
Size:
1.71 KB
Format:
Item-specific license agreed upon to submission
Description:

Collections