Papers

Applying A Speaker-dependent Speech Compression Technique to Concatenative TTS Synthesizers

International Journal

2006~2010

작성자

이진영

작성일

2007-02-01 13:32

조회

3103

Authors : Chang-Heon Lee, Sung-Kyo Jung, Hong-Goo Kang

Year : 2007

Publisher / Conference : IEEE Transactions on Audio, Speech, and Language Processing

Volume : 15, 2

Page : 632-640

This paper proposes a new speaker-dependent coding algorithm to efficiently compress a large speech database for corpus-based concatenative text-to-speech (TTS) engines while maintaining high fidelity. To achieve a high compression ratio and meet the fundamental requirements of concatenative TTS synthesizers, such as partial segment decoding and random access capability, we adopt a nonpredictive analysis-by-synthesis scheme for speaker-dependent parameter estimation and quantization. The spectral coefficients are quantized by using a memoryless split vector quantization (VQ) approach that does not use frame correlation. Considering that excitation signals of a specific speaker show low intra-variation especially in the voiced regions, the conventional adaptive codebook for pitch prediction is replaced by a speaker-dependent pitch-pulse codebook trained by a corpus of single-speaker speech signals. To further improve the coding efficiency, the proposed coder flexibly combines nonpredictive and predictive type method considering the structure of the TTS system. By applying the proposed algorithm to a Korean TTS system, we could obtain comparable quality to the G.729 speech coder and satisfy all the requirements that TTS system needs. The results are verified by both objective and subjective quality measurements. In addition, the decoding complexity of the proposed coder is around 55% lower than that of G.729 annex A

« Optimum beamformer in correlated source environments

A Soft-Decision Adaptation Mode Controller for an Efficient Frequency-Domain Generalized Sidelobe Canceller »

목록보기

전체 372

62	International Conference	Yoomi Hur, Young-Choel Park, Dae Hee Youn "A New Structure for Stereo Acoustic Echo Cancellation Based on Binaural Cue Coding" in 122th Convention of Audio Engineering Society, pp.7094, 2007
61	International Conference	Jaseong Lee, Young-Cheol Park, Dae Hee Youn, Hong-Goo Kang "A Hybrid Warped Linear Prediction (WLP) AAC Audio Coding Algorithm" in 122th Convention of Audio Engineering Society, pp.7002, 2007
60	International Conference	Keun-Sup Lee, Young-Cheol Park, Dae Hee Youn "Fixed-Point Processing Optimization of Mpeg Audio Encoder Using Statistical Model" in 122th Convention of Audio Engineering Society, pp.7023, 2007
59	Domestic Conference	이동금, 박영철, 윤대희 "저 전송율 환경에서의 음질 개선에 대한 연구" in 2007년도 한국음향학회 춘계학술발표대회, 2007
58	International Conference	Min-Seok Choi, Chang-Hyun Baik, Young-Cheol Park, Hong-Goo Kang "A Soft-Decision Adaptation Mode Controller for an Efficient Frequency-Domain Generalized Sidelobe Canceller" in ICASSP, 2007
57	International Journal	Chang-Heon Lee, Sung-Kyo Jung, Hong-Goo Kang "Applying A Speaker-dependent Speech Compression Technique to Concatenative TTS Synthesizers" in IEEE Transactions on Audio, Speech, and Language Processing, vol.15, 2, pp.632-640, 2007
56	International Journal	Seungil Kim, Chungyong Lee, Hong-Goo Kang "Optimum beamformer in correlated source environments" in Journal of Acoustical Society of America, vol.120, issue 6, 2006
55	Domestic Conference	현동일, 박영철, 윤대희 "파라메트릭 스테레오 오디오 부호화를 위한 향상된 채널간 위상차 합성" in 2006년도 추계종합학술발표회, 2006
54	International Journal	Kyoung Ho Bang, Young-Cheol Park, Jeongil Seo "Audio Transcoding for Audio Streams from a T-DTV Broadcasting Station to a T-DMB Receiver" in ETRI Journal, vol.28, issue5, pp.664-667, 2006
53	International Conference	Myung-Suk Song, Chang-Heon Lee, Hong-Goo Kang "Performance Analysis of Various Single Channel Speech Enhancement Algorithms for Automatic Speech Recognition" in INTERSPEECH, 2006

Applying A Speaker-dependent Speech Compression Technique to Concatenative TTS Synthesizers

Previous

Sister Lab.

Yonsei University

Academic Website