NWU Institutional Repository

Investigating supervised surface segmentation of isiZulu text using synthetic data generation

Loading...
Thumbnail Image

Date

Journal Title

Journal ISSN

Volume Title

Publisher

North-West University (South Africa)

Record Identifier

Abstract

IsiZulu, one of South Africa's most widely spoken languages, is classified as a low-resource language, especially regarding digital tools. As part of the Nguni family, isiZulu exhibits complex morphology and conjunctive orthography. These features result in data sparsity, as a single root or stem may appear in numerous morphological variants, complicating language modelling. This underscores the importance of morphological segmentation, a natural language processing NLP task that decomposes words into their smallest meaningful units (morphemes). Rule-based methods yield high accuracy in low-resource contexts but typically lack robustness and are costly to develop. Conversely, machine learning approaches require large, high-quality datasets, often unavailable for low-resource languages. To address these challenges, this study employs a hybrid approach using a rule-based system, the isiZulu Resource Grammar (ZRG), to generate synthetic datasets with varying segmentation granularities. These datasets underwent data augmentation through syntactic tree manipulation, significantly increasing their size and diversity. Subsequently, this data trained supervised machine learning models: Conditional Random Fields (CRF), Long Short-Term Memory (LSTM), and Transformer-based models--for morphological segmentation. The effectiveness of these models was assessed intrinsically, using precision, recall, F1 score, BLEU, and chrF, and extrinsically, by evaluating their impact on Neural Machine Translation (NMT) quality for isiZulu-English translation. Intrinsic evaluation showed that the Transformer model consistently outperformed the CRF and LSTM models, achieving segmentation accuracy above 0.9 across all metrics and granularity styles. Additionally, the hybrid approach demonstrated superior robustness, effectively handling out-of-vocabulary (OOV) words and performing segmentation 30 times faster than ZRG alone. Extrinsic evaluation confirmed that segmentation improved translation quality, with Segmenter Two achieving the highest BLEU score (0.235), representing a 25.0% improvement over the unsegmented baseline (0.188). These findings highlight the effectiveness of integrating rule-based and machine learning approaches for morphological segmentation, offering a scalable solution for processing low-resource languages with complex morphologies such as isiZulu in NLP applications.

Sustainable Development Goals

Description

Master of Science in Computer Science, North-West University, Potchefstroom Campus

Citation

Endorsement

Review

Supplemented By

Referenced By