Machine translation training data form English–Tshivenḓa
Loading...
Date
Researcher ID
Supervisors
Journal Title
Journal ISSN
Volume Title
Publisher
Data in Brief
Record Identifier
Abstract
This data article describes a machine translation training data set for translation between English and Tshiven ḓa. The data set contains parallel, aligned English-Tshiven ḓa data as well as monolingual Tshiven ḓa data. The data was collected from both web crawling of multilingual South African government sites and matched documents from translators or publishing sources. Additional unique data was translated from English into Tshiven ḓa by professional translators to increase the to- tal corpus size. This article contains information about the collection and translation of the data as well as how align- ments and corpus cleanup were done. The wordcounts of the
corpus are also given. In addition to training machine transla- tion systems this data can also be used for the development of other Tshiven ḓa core technologies as well as for linguistic studies.
Sustainable Development Goals
Description
Journal Article, Faculty of Humanities, Unit for Languages and Literature In The SA Context-- Potchefstroom Campus
Citation
Puttkammer, Martin J. et al. 2024. Machine translation training data form English–Tshivenḓa. Data in Brief, Volume 57, (2024), 110898, [https://doi.org/10.1016/j.dib.2024.110898]
