Automatic speech segmentation with limited data

Van Niekerk, Daniel Rudolph

Automatic speech segmentation with limited data

Files

vanniekerk_danielr.pdf (3.01 MB)

Date

2009

Authors

Van Niekerk, Daniel Rudolph

Publisher

North-West University

Abstract

The rapid development of corpus-based speech systems such as concatenative synthesis systems for under-resourced languages requires an efﬁcient, consistent and accurate solution with regard to phonetic speech segmentation. Manual development of phonetically annotated corpora is a time consuming and expensive process which suffers from challenges regarding consistency and reproducibility, while automation of this process has only been satisfactorily demonstrated on large corpora of a select few languages by employing techniques requiring extensive and specialised resources. In this work we considered the problem of phonetic segmentation in the context of developing small prototypical speech synthesis corpora for new under-resourced languages. This was done through an empirical evaluation of existing segmentation techniques on typical speech corpora in three South African languages. In this process, the performance of these techniques were characterised under different data conditions and the efﬁcient application of these techniques were investigated in order to improve the accuracy of resulting phonetic alignments. We found that the application of baseline speaker-speciﬁc Hidden Markov Models results in relatively robust and accurate alignments even under extremely limited data conditions and demonstrated how such models can be developed and applied efﬁciently in this context. The result is segmentation of sufﬁcient quality for synthesis applications, with the quality of alignments comparable to manual segmentation efforts in this context. Finally, possibilities for further automated reﬁnement of phonetic alignments were investigated and an efﬁcient corpus development strategy was proposed with suggestions for further work in this direction.

Description

Thesis (M.Ing. (Computer Engineering))--North-West University, Potchefstroom Campus, 2009.

Keywords

Phonetic speech segmentation, Phonetic alignment, Speech synthesis, Text-to-speech, Speech corpus development, Resource scarce languages, Hidden Markov models, Dynamic time warping

URI

http://hdl.handle.net/10394/3978

Collections

Engineering

Full item page

Automatic speech segmentation with limited data

Files

Date

Authors

Researcher ID

Supervisors

Journal Title

Journal ISSN

Volume Title

Publisher

Record Identifier

Abstract

Sustainable Development Goals

Description

Keywords

Citation

URI

Collections

Endorsement

Review

Supplemented By

Referenced By