NWU Institutional Repository

Machine translation training data form English–Tshivenḓa

Loading...
Thumbnail Image

Date

Researcher ID

Supervisors

Journal Title

Journal ISSN

Volume Title

Publisher

Data in Brief

Record Identifier

Abstract

This data article describes a machine translation training data set for translation between English and Tshiven ḓa. The data set contains parallel, aligned English-Tshiven ḓa data as well as monolingual Tshiven ḓa data. The data was collected from both web crawling of multilingual South African government sites and matched documents from translators or publishing sources. Additional unique data was translated from English into Tshiven ḓa by professional translators to increase the to- tal corpus size. This article contains information about the collection and translation of the data as well as how align- ments and corpus cleanup were done. The wordcounts of the corpus are also given. In addition to training machine transla- tion systems this data can also be used for the development of other Tshiven ḓa core technologies as well as for linguistic studies.

Sustainable Development Goals

Description

Journal Article, Faculty of Humanities, Unit for Languages and Literature In The SA Context-- Potchefstroom Campus

Citation

Puttkammer, Martin J. et al. 2024. Machine translation training data form English–Tshivenḓa. Data in Brief, Volume 57, (2024), 110898, [https://doi.org/10.1016/j.dib.2024.110898]

Collections

Endorsement

Review

Supplemented By

Referenced By