Building a Dataset for Misinformation Detection in the Low-Resource Language
| dc.contributor.author | MUKWEVHO Mulweli | |
| dc.contributor.author | RANANGA Seani | |
| dc.contributor.author | S MBOOI Mahlatse | |
| dc.contributor.author | ISONG Bassey | |
| dc.contributor.author | MARIVATE Vukosi | |
| dc.date.accessioned | 2025-10-31T07:57:43Z | |
| dc.date.issued | 2024 | |
| dc.description | Department of Computer Science, North-West University, South Africa | |
| dc.description.abstract | In the modern digital age, the widespread dissemination of misinformation has become a serious issue. Most focus in identifying misinformation online has been targeted at the English language in contrast to lowresource languages like Tshivenda. In this paper, we create a new dataset for news in the Tshivenda language to assist in developing resources for misinformation in the language. In our proposed methodology, we leveraged conditional random fields (CRF), gated recurrent unit (GRU), and long short-term memory (LSTM) to collect and annotate social media content. By applying these deep learning approaches to existing Tshivenda posts, we can assess their effectiveness for identifying false news in a low-resource language setting. This paper emphasises the vital need to combat misinformation in languages with limited resources, such as Tshivenda. Through the creation of a specialised dataset and the use of advanced techniques, it aims to address the problem of the spread of misinformation in low represented language communities. | |
| dc.identifier.citation | Mukwevho, M., Rananga, S., Mbooi, M.S., Isong, B. and Marivate, V., 2024. Building a dataset for misinformation detection in the low-resource language. In 2024 IST-Africa Conference (IST-Africa). | |
| dc.identifier.uri | http://hdl.handle.net/10394/43803 | |
| dc.language.iso | en | |
| dc.publisher | IEEE | |
| dc.subject | Misinformation | |
| dc.subject | Natural Language Processing (NLP) | |
| dc.subject | social media | |
| dc.subject | lowresource language | |
| dc.subject | Conditional Random Fields (CRF) | |
| dc.subject | Gated Recurrent Unit (GRU) | |
| dc.subject | Long Short-Term Memory (LSTM) | |
| dc.title | Building a Dataset for Misinformation Detection in the Low-Resource Language | |
| dc.type | Article |
