Papers

Applying A Speaker-dependent Speech Compression Technique to Concatenative TTS Synthesizers

International Journal

2006~2010

작성자

이진영

작성일

2007-02-01 13:32

조회

1060

Authors : Chang-Heon Lee, Sung-Kyo Jung, Hong-Goo Kang

Year : 2007

Publisher / Conference : IEEE Transactions on Audio, Speech, and Language Processing

Volume : 15, 2

Page : 632-640

This paper proposes a new speaker-dependent coding algorithm to efficiently compress a large speech database for corpus-based concatenative text-to-speech (TTS) engines while maintaining high fidelity. To achieve a high compression ratio and meet the fundamental requirements of concatenative TTS synthesizers, such as partial segment decoding and random access capability, we adopt a nonpredictive analysis-by-synthesis scheme for speaker-dependent parameter estimation and quantization. The spectral coefficients are quantized by using a memoryless split vector quantization (VQ) approach that does not use frame correlation. Considering that excitation signals of a specific speaker show low intra-variation especially in the voiced regions, the conventional adaptive codebook for pitch prediction is replaced by a speaker-dependent pitch-pulse codebook trained by a corpus of single-speaker speech signals. To further improve the coding efficiency, the proposed coder flexibly combines nonpredictive and predictive type method considering the structure of the TTS system. By applying the proposed algorithm to a Korean TTS system, we could obtain comparable quality to the G.729 speech coder and satisfy all the requirements that TTS system needs. The results are verified by both objective and subjective quality measurements. In addition, the decoding complexity of the proposed coder is around 55% lower than that of G.729 annex A

« Optimum beamformer in correlated source environments

Perceptual relevance of the temporal envelope to the speech signal in the 4–7kHz band »

목록보기

전체 355

24	International Journal	Jae-Seong Lee, Chang-Joon Lee, Young-Cheol Park, Dae Hee Youn "Efficient FFT Algorithm for Psychoacoustic Model of the MPEG-4 AAC" in IEICE Transactions on Information and Systems, vol.E92-D, No.12, pp.2535-2539, 2009
23	International Journal	Chang-Heon Lee, Hyen-O Oh, Hong-Goo Kang "On the Study of Noise Allocation for Speech Signal in Low Bit-Rate Audio Coding" in IEEE Signal Processing Letters, vol.16, issue 10, pp.849-852, 2009
22	International Journal	Bong-Jin Lee, Jeung-Yoon Choi, Hong-Goo Kang "Phonetically optimized speaker modeling for robust speaker recognition" in The Journal of the Acoustical Society of America, vol.126, issue 3, 2009
21	International Journal	Tacksung Choi, Sunkuk Moon, Young-Cheol Park, Dea Hee Youn, Seokpil Lee "A GMM-Based Feature Selection Algorithm for Multi-Class Classification" in IEICE Transactions on Information and Systems, vol.E92-D. No.8, pp.1584-1587, 2009
20	International Journal	Jung-Won Lee, Jeung-Yoon Choi "Acoustic‐phonetic features for stop consonant place detection in clean and telephone speech" in The Journal of the Acoustical Society of America, vol.123, issue 5, 2008
19	International Journal	Jungin Lee, Jeung-Yoon Choi "Detection of obstruent consonant landmark for knowledge based speech recognition" in The Journal of the Acoustical Society of America, vol.123, issue 8, 2008
18	International Journal	Sukmyung Lee, Jeung-Yoon Choi "Vowel place detection for a knowledge‐based speech recognition system" in The Journal of the Acoustical Society of America, vol.123, issue 5, 2008
17	International Journal	Kyung-Tae Kim, Min-Ki Lee, Hong-Goo Kang "Speech Bandwidth Extension using Temporal Envelope Modeling" in IEEE Signal Processing Letters, vol.15, pp.429-432, 2008
16	International Journal	Junho Lee, Eunjung Song, Young-Cheol Park, Dae Hee Youn "Effective Bass Enhancement Using Second-Order Adaptive Notch Filter" in IEEE Transactions on Consumer Electronics, vol.54, issue 2, pp.663-668, 2008
15	International Journal	Hee-Young Park, Ki-Man Kim, Hyun-Woo Kang, Dae Hee Youn, Chungyong Lee "A Simplified Subspace Fitting Method for Estimating Shape of a Towed Array" in IEEE Journal of Oceanic Engineering, vol.33, issue 2, pp.215-223, 2008

Applying A Speaker-dependent Speech Compression Technique to Concatenative TTS Synthesizers

Previous

Sister Lab.

Yonsei University

Academic Website