Collecting and evaluating speech recognition corpora for 11 South African languages

Badenhorst, Jaco; Van Heerden, Charl; Barnard, Etienne; Davel, Marelie H.

Collecting and evaluating speech recognition corpora for 11 South African languages

Date

2011

Authors

Badenhorst, Jaco

Van Heerden, Charl

Barnard, Etienne

Davel, Marelie H.

Researcher ID

11539151 - Van Heerden, Carel Jacobus
23607955 - Davel, Marelie Hattingh
21021287 - Barnard, Etienne

Publisher

Springer

Abstract

We describe the Lwazi corpus for automatic speech recognition (ASR), a new telephone speech corpus which contains data from the eleven official languages of South Africa. Because of practical constraints, the amount of speech per language is relatively small compared to major corpora in world languages, and we report on our investigation of the stability of the ASR models derived from the corpus. We also report on phoneme distance measures across languages, and describe initial phone recognisers that were developed using this data. We find that a surprisingly small number of speakers (fewer than 50) and around 10 to 20 h of speech per language are sufficient for the purposes of acceptable phone-based recognition.

Keywords

Speech recognition, Lwazi corpus, Resource-scarce languages, South African languages

Citation

Badenhorst, J. & Van Heerden, C., et al. 2011. Collecting and evaluating speech recognition corpora for 11 South African languages. Language resources and evaluation, 45(3):289-309. [http://link.springer.com/journal/10579]

URI

http://hdl.handle.net/10394/13095

Collections

Faculty of Engineering
Faculty of Natural and Agricultural Sciences

Full item page

Collecting and evaluating speech recognition corpora for 11 South African languages

Date

Authors

Researcher ID

Supervisors

Journal Title

Journal ISSN

Volume Title

Publisher

Record Identifier

Abstract

Sustainable Development Goals

Description

Keywords

Citation

URI

Collections

Endorsement

Review

Supplemented By

Referenced By