NWU Institutional Repository

Part-of-speech effects on text-to-speech synthesis

Loading...
Thumbnail Image

Date

Authors

Schlunz, Georg I.
Barnard, Etienne
van Huyssteen, Gerhard B.

Researcher ID

Supervisors

Journal Title

Journal ISSN

Volume Title

Publisher

Pattern Recognition Association of South Africa and Mechatronics International Conference

Record Identifier

Abstract

One of the goals of text-to-speech (TTS) systems is to produce natural-sounding synthesized speech. Towards this end various natural language processing (NLP) tasks are performed to model the prosodic aspects of the TTS voice. One of the fundamental NLP tasks being used is the part-of-speech (POS) tagging of the words in the text. This paper investigates the effects of POS information on the naturalness of a hidden Markov model (HMM) based TTS voice when additional resources are not available to aid in the modeling of prosody. It is found that, when a minimal feature set is used for the HMM context labels, the addition of POS tags does improve the naturalness of the voice. However, the same effect can be accomplished by including segmental counting and positional information instead of the POS tags.

Sustainable Development Goals

Description

Citation

Georg Schlünz, Etienne Barnard and Gerhard van Huyssteen, “Part-of-speech effects on text-to-speech synthesis”, in Proc. Annual Symp. Pattern Recognition Association of South Africa (PRASA), pp 257-262, Stellenbosch, South Africa, 2010. [http://engineering.nwu.ac.za/multilingual-speech-technologies-must/publications]

Endorsement

Review

Supplemented By

Referenced By