NWU Institutional Repository

Efficient harvesting of Internet audio for resource-scarce ASR

Loading...
Thumbnail Image

Date

Authors

Davel, Marelie H.
van Heerden, Charl
Kleynhans, Neil
Barnard, Etienne

Researcher ID

Supervisors

Journal Title

Journal ISSN

Volume Title

Publisher

Interspeech 2011

Record Identifier

Abstract

Spoken recordings that have been transcribed for human reading (e.g. as captions for audiovisual material, or to provide alternative modes of access to recordings) are widely available in many languages. Such recordings and transcriptions have proven to be a valuable source of ASR data in well-resourced languages, but have not been exploited to a significant extent in under-resourced languages or dialects. Techniques used to harvest such data typically assume the availability of a fairly accurate ASR system, which is generally not available when working with resourcescarce languages. In this work, we define a process whereby an ASR corpus is bootstrapped using unmatched ASR models in conjunction with speech and approximate transcriptions sourced from the Internet. We introduce a new segmentation technique based on the use of a phone-internal garbage model, and demonstrate how this technique (combined with limited filtering) can be used to develop a large, high-quality corpus in an underresourced dialect with minimal effort.

Sustainable Development Goals

Description

Citation

Marelie H Davel, Charl Van Heerden, Neil Kleynhans and Etienne Barnard, “Efficient harvesting of Internet audio for resource-scarce ASR”, in Proc. Interspeech, pp 3153-3156, Florence, Italy, 2011. [http://engineering.nwu.ac.za/multilingual-speech-technologies-must/publications]

Endorsement

Review

Supplemented By

Referenced By