NWU Institutional Repository

Investigating supervised surface segmentation of isiZulu text using synthetic data generation

dc.contributor.advisorMarais, Len_ZA
dc.contributor.advisorGoede, Roelienen_ZA
dc.contributor.authorMkhwanazi, Sthembisoen_ZA
dc.contributor.researchIDGoede, Roelien-10085971en_ZA
dc.date.accessioned2025-09-10T08:25:20Z
dc.date.issued2025
dc.descriptionMaster of Science in Computer Science, North-West University, Potchefstroom Campus
dc.description.abstractIsiZulu, one of South Africa's most widely spoken languages, is classified as a low-resource language, especially regarding digital tools. As part of the Nguni family, isiZulu exhibits complex morphology and conjunctive orthography. These features result in data sparsity, as a single root or stem may appear in numerous morphological variants, complicating language modelling. This underscores the importance of morphological segmentation, a natural language processing NLP task that decomposes words into their smallest meaningful units (morphemes). Rule-based methods yield high accuracy in low-resource contexts but typically lack robustness and are costly to develop. Conversely, machine learning approaches require large, high-quality datasets, often unavailable for low-resource languages. To address these challenges, this study employs a hybrid approach using a rule-based system, the isiZulu Resource Grammar (ZRG), to generate synthetic datasets with varying segmentation granularities. These datasets underwent data augmentation through syntactic tree manipulation, significantly increasing their size and diversity. Subsequently, this data trained supervised machine learning models: Conditional Random Fields (CRF), Long Short-Term Memory (LSTM), and Transformer-based models--for morphological segmentation. The effectiveness of these models was assessed intrinsically, using precision, recall, F1 score, BLEU, and chrF, and extrinsically, by evaluating their impact on Neural Machine Translation (NMT) quality for isiZulu-English translation. Intrinsic evaluation showed that the Transformer model consistently outperformed the CRF and LSTM models, achieving segmentation accuracy above 0.9 across all metrics and granularity styles. Additionally, the hybrid approach demonstrated superior robustness, effectively handling out-of-vocabulary (OOV) words and performing segmentation 30 times faster than ZRG alone. Extrinsic evaluation confirmed that segmentation improved translation quality, with Segmenter Two achieving the highest BLEU score (0.235), representing a 25.0% improvement over the unsegmented baseline (0.188). These findings highlight the effectiveness of integrating rule-based and machine learning approaches for morphological segmentation, offering a scalable solution for processing low-resource languages with complex morphologies such as isiZulu in NLP applications.
dc.description.sponsorship-Council for Scientific and Industrial Research (CSIR)
dc.description.thesistypeMastersen
dc.identifier.urihttps://orcid.org/0009-0003-3050-4744
dc.identifier.urihttp://hdl.handle.net/10394/43369
dc.language.isoen
dc.publisherNorth-West University (South Africa)
dc.subjectAgglutinative Languages
dc.subjectMorphological Segmentation
dc.subjectisiZulu
dc.subjectSupervised Segmenter Learning
dc.subjectRule-Based Segmenter
dc.titleInvestigating supervised surface segmentation of isiZulu text using synthetic data generation
dc.typeThesis

Files

Original bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
Mkhwanazi_SN_2025.pdf
Size:
3.91 MB
Format:
Adobe Portable Document Format

License bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
license.txt
Size:
1.71 KB
Format:
Item-specific license agreed upon to submission
Description: