Papers

Applying A Speaker-dependent Speech Compression Technique to Concatenative TTS Synthesizers

International Journal

2006~2010

작성자

이진영

작성일

2007-02-01 13:32

조회

1296

Authors : Chang-Heon Lee, Sung-Kyo Jung, Hong-Goo Kang

Year : 2007

Publisher / Conference : IEEE Transactions on Audio, Speech, and Language Processing

Volume : 15, 2

Page : 632-640

This paper proposes a new speaker-dependent coding algorithm to efficiently compress a large speech database for corpus-based concatenative text-to-speech (TTS) engines while maintaining high fidelity. To achieve a high compression ratio and meet the fundamental requirements of concatenative TTS synthesizers, such as partial segment decoding and random access capability, we adopt a nonpredictive analysis-by-synthesis scheme for speaker-dependent parameter estimation and quantization. The spectral coefficients are quantized by using a memoryless split vector quantization (VQ) approach that does not use frame correlation. Considering that excitation signals of a specific speaker show low intra-variation especially in the voiced regions, the conventional adaptive codebook for pitch prediction is replaced by a speaker-dependent pitch-pulse codebook trained by a corpus of single-speaker speech signals. To further improve the coding efficiency, the proposed coder flexibly combines nonpredictive and predictive type method considering the structure of the TTS system. By applying the proposed algorithm to a Korean TTS system, we could obtain comparable quality to the G.729 speech coder and satisfy all the requirements that TTS system needs. The results are verified by both objective and subjective quality measurements. In addition, the decoding complexity of the proposed coder is around 55% lower than that of G.729 annex A

« Optimum beamformer in correlated source environments

A Soft-Decision Adaptation Mode Controller for an Efficient Frequency-Domain Generalized Sidelobe Canceller »

목록보기

전체 364

74	International Journal	Kyung Tae Kim, Jeung-Yoon Choi, Hong-Goo Kang "Perceptual relevance of the temporal envelope to the speech signal in the 4–7kHz band" in The Journal of the Acoustical Society of America, vol.122, issue 3, 2007
73	International Conference	Chang-Heon Lee, Hong-Goo Kang "Improvement of Artificial Onset Reconstruction in VMR-WB Standard under Packet-loss Environments" in ITC-CSCC, 2007
72	International Conference	Sun-kuk Moon, Tack-sung Choi, Young-Cheol Park, Dae Hee Youn "An Efficient Feature Selection Algorithm Based on Kullback-Leibler Divergence for Music Information Retrieval" in ITC-CSCC, pp.524-525, 2007
71	International Conference	Min-Ki Lee, Sung-Wan Youn, Kyung-Tae Kim, Hong-Goo Kang "Speech Quality Degradation in Packet Loss Environment at Specific Speech Class" in ITC-CSCC, pp.781-782, 2007
70	International Conference	Chi-Sang Jung, Bong-Jin Lee, Jeung-Yoon Choi, Hong-Goo Kang "AN ADAPTIVE SELECTION OF FRAME SHIFT FOR SPEAKER RECOGNITION SYSTEMS" in The Second Beijing-Hong Kong International Doctoral Forum, 2007
69	International Conference	Min-Seok Choi, Young-Cheol Park, Hong-Goo Kang "A Soft-Decision Adaptation Mode Controller for an Efficient Frequency-Domain Generalized Sidelobe Canceller" in ICASSP, 2007
68	International Conference	Yoomi Hur, Young-Choel Park, Dae Hee Youn "A New Structure for Stereo Acoustic Echo Cancellation Based on Binaural Cue Coding" in 122th Convention of Audio Engineering Society, pp.7094, 2007
67	Domestic Conference	정치상, 이봉진, 최정윤, 강홍구, 윤대희 "화자인식 시스템을 위한 다양한 프레임 이동 길이 선택 방법에 관한 연구" in 2007년도 하계종합학술발표회, 2007
66	Domestic Conference	문선국, 최택성, 박영철, 윤대희 "다양한 해상도의 필터뱅크에 따른 음악 장르 분류를 위한 특징벡터의 성능 비교" in 2007년도 하계종합학술발표회, 2007
65	Domestic Conference	최택성, 문선국, 박영철 "Relative Specific Loudness Histogram 특징벡터를 이용한 음악 장르 분류" in 2007년도 하계종합학술발표회, 2007

Applying A Speaker-dependent Speech Compression Technique to Concatenative TTS Synthesizers

Previous

Sister Lab.

Yonsei University

Academic Website