Building a Dataset for Misinformation Detection in the Low-Resource Language
Loading...
Date
Researcher ID
Supervisors
Journal Title
Journal ISSN
Volume Title
Publisher
IEEE
Record Identifier
Abstract
In the modern digital age, the widespread dissemination of
misinformation has become a serious issue. Most focus in identifying
misinformation online has been targeted at the English language in contrast to lowresource languages like Tshivenda. In this paper, we create a new dataset for news in
the Tshivenda language to assist in developing resources for misinformation in the
language. In our proposed methodology, we leveraged conditional random fields
(CRF), gated recurrent unit (GRU), and long short-term memory (LSTM) to collect
and annotate social media content. By applying these deep learning approaches to
existing Tshivenda posts, we can assess their effectiveness for identifying false news
in a low-resource language setting. This paper emphasises the vital need to combat
misinformation in languages with limited resources, such as Tshivenda. Through the
creation of a specialised dataset and the use of advanced techniques, it aims to
address the problem of the spread of misinformation in low represented language
communities.
Sustainable Development Goals
Description
Department of Computer Science, North-West University, South Africa
Citation
Mukwevho, M., Rananga, S., Mbooi, M.S., Isong, B. and Marivate, V., 2024. Building a dataset for misinformation detection in the low-resource language. In 2024 IST-Africa Conference (IST-Africa).
