NWU Institutional Repository

Machine translation training data form English–Tshivenḓa

dc.contributor.authorGaustad, Tanjaen.ZA
dc.contributor.authorMcKellar, Cindy A.en.ZA
dc.contributor.authorPuttkammer, Martin Jen.ZA
dc.date.accessioned2025-10-13T10:01:58Zen.ZA
dc.date.issued2024en.ZA
dc.descriptionJournal Article, Faculty of Humanities, Unit for Languages and Literature In The SA Context-- Potchefstroom Campusen.ZA
dc.description.abstractThis data article describes a machine translation training data set for translation between English and Tshiven ḓa. The data set contains parallel, aligned English-Tshiven ḓa data as well as monolingual Tshiven ḓa data. The data was collected from both web crawling of multilingual South African government sites and matched documents from translators or publishing sources. Additional unique data was translated from English into Tshiven ḓa by professional translators to increase the to- tal corpus size. This article contains information about the collection and translation of the data as well as how align- ments and corpus cleanup were done. The wordcounts of the corpus are also given. In addition to training machine transla- tion systems this data can also be used for the development of other Tshiven ḓa core technologies as well as for linguistic studies.en.ZA
dc.description.sponsorshipAcknowledgments This research was made possible with the support from the South African Department of Arts and Culture (DSAC) as well as the support from the South African Centre for Digital Lan- guage Resources (SADiLaR), a research infrastructure established by the Department of Science and Innovation (DSI) of the South African government as part of the South African Research In- frastructure Roadmap (SARIR). The creation of the dataset was funded as part of the ongoing Autshumato project.en.ZA
dc.identifier.citationPuttkammer, Martin J. et al. 2024. Machine translation training data form English–Tshivenḓa. Data in Brief, Volume 57, (2024), 110898, [https://doi.org/10.1016/j.dib.2024.110898]en.ZA
dc.identifier.urihttps://doi.org/10.1016/j.dib.2024.110898en.ZA
dc.identifier.urihttp://hdl.handle.net/10394/43645en.ZA
dc.language.isoenen.ZA
dc.publisherData in Briefen.ZA
dc.subjectMachine Translationen.ZA
dc.subjectNatural Language Processingen.ZA
dc.subjectHuman Language Technologyen.ZA
dc.subjectParallel Corporaen.ZA
dc.subjectBilingual Dataen.ZA
dc.subjectEnglishen.ZA
dc.subjectTshivenḓaen.ZA
dc.titleMachine translation training data form English–Tshivenḓaen.ZA
dc.typeArticleen.ZA

Files

Original bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
Puttkammer, Martin J. et al. 2024.pdf
Size:
348.93 KB
Format:
Adobe Portable Document Format

License bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
license.txt
Size:
1.71 KB
Format:
Item-specific license agreed upon to submission
Description:

Collections