Papers

Effective Spectral and Excitation Modeling Techniques for LSTM-RNN-Based Speech Synthesis Systems

International Journal
2016~2020
작성자
이진영
작성일
2017-11-01 22:07
조회
7968
Authors : Eunwoo Song, Frank K. Soong, Hong-Goo Kang

Year : 2017

Publisher / Conference : IEEE/ACM Transactions on Audio, Speech, and Language Processing

Volume : 25, issue 11

Page : 2152-2161

In this paper, we report research results on modeling the parameters of an improved time-frequency trajectory excitation (ITFTE) and spectral envelopes of an LPC vocoder with a long short-term memory (LSTM)-based recurrent neural network (RNN) for high-quality text-to-speech (TTS) systems. The ITFTE vocoder has been shown to significantly improve the perceptual quality of statistical parameter-based TTS systems in our prior works. However, a simple feed-forward deep neural network (DNN) with a finite window length is inadequate to capture the time evolution of the ITFTE parameters. We propose to use the LSTM to exploit the time-varying nature of both trajectories of the excitation and filter parameters, where the LSTM is implemented to use the linguistic text input and to predict both ITFTE and LPC parameters holistically. In the case of LPC parameters, we further enhance the generated spectrum by applying LP bandwidth expansion and line spectral frequency-sharpening filters. These filters are not only beneficial for reducing unstable synthesis filter conditions but also advantageous toward minimizing the muffling problem in the generated spectrum. Experimental results have shown that the proposed LSTM-RNN system with the ITFTE vocoder significantly outperforms both similarly configured band aperiodicity-based systems and our best prior DNN-trainecounterpart, both objectively and subjectively.
전체 387
387 International Journal Jihyun Kim, Doyeon Kim, Hong-Goo Kang "Speaker-Discriminative Attractors for Robust Continuous Speech Separation" in TASLP, vol.34, pp.3916-3929, 2026
386 Domestic Conference 김효민, 이지현, 장인선, 강홍구 "사후 의미 증류를 이용한 유한 스칼라 양자화 기반 이단계 음성 토크나이저" in 한국방송·미디어공학회 2026년 하계학술대회, 2026
385 Domestic Conference 신재훈, 장인선, 강홍구 "효율적인 신경망 기반 오디오 코덱을 위한 잔차 오토인코딩 및 연속형 오토인코더의 잠재 표현 증류" in 한국방송·미디어공학회 2026년 하계학술대회, 2026
384 International Conference Jihyun Lee, Jiahao Li, Woojin Chung, Yan Lu, Hong-Goo Kang "AudioSketch: Controllable Image-to-Audio Generation via Semantic-Temporal Energy Modulation" in EUSIPCO, 2026
383 International Conference Sangmin Lee, Woojin Chung, Woongjib Choi, Hong-Goo Kang "MoLGE: Mixture of Language Group Experts for Efficient Scaling of Massively Multilingual Speech Recognition" in Conference On Language Modeling (COLM), 2026
382 International Conference Sangmin Lee, Eekgyun Ahn, Woongjib Choi, Hong-Goo Kang "UR-BERT: Scaling Text Encoders for Massively Multilingual TTS Through Universal Romanization and Speech Token Prediction" in INTERSPEECH, 2026
381 International Conference Seyun Um, Doyeon Kim, Hong-Goo Kang "HANUI: Harnessing Distributional Discrepancies for Singing Voice Deepfake Detection" in in IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP), 2026
380 International Conference Miseul Kim, Soo jin Park, Kyungguen Byun, Hyeon-Kyeong Shin, Sunkuk Moon, Shuhua Zhang, Erik Visser "Mitigating Intra-Speaker Variability in Diarization with Style-Controllable Speech Augmentation" in in IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP), 2026
379 International Conference Woongjib Choi, Sangmin Lee, Hyungseob Lim, Hong-Goo Kang "UniverSR: Unified and Versatile Audio Super-Resolution via Vocoder-Free Flow Matching" in IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP), 2026
378 International Journal Hyeonjin Cha, Seyun Um, Miseul Kim, Changhwan Kim, Seungshin Lee, Hong-Goo Kang "Content-Aware Style Augmentation for Zero-Shot Voice Conversion With Short Target Speech" in IEEE Signal Processing Letters, vol.33, pp.66-70, 2025