Papers

Multi-class learning algorithm for deep neural network-based statistical parametric speech synthesis

International Conference
2016~2020
작성자
한혜원
작성일
2016-08-01 16:23
조회
265
Authors : Eunwoo Song, Hong-Goo Kang

Year : 2016

Publisher / Conference : EUSIPCO

This paper proposes a multi-class learning (MCL) algorithm for a deep neural network (DNN)-based statistical parametric speech synthesis (SPSS) system. Although the DNN-based SPSS system improves the modeling accuracy of statistical parameters, its synthesized speech is often muffled because the training process only considers the global characteristics of the entire set of training data, but does not explicitly consider any local variations. We introduce a DNN-based context clustering algorithm that implicitly divides the training data into several classes, and train them via a shared hidden layer-based MCL algorithm. Since the proposed MCL method efficiently models both the universal and class-dependent characteristics of various phonetic information, it not only avoids the model over-fitting problem but also reduces the over-smoothing effect. Objective and subjective test results also verify that the proposed algorithm performs much better than the conventional method.
전체 326
256 International Conference Haemin Yang, Kyungguen Byun, Youngsu Kwak, Hong-Goo Kang "Parametric-based non-intrusive speech quality assessment by deep neural network" in 21th International Conference on Digital Signal Processing (DSP), 2016
255 International Conference Jin-Seob Kim, Young-Sun Joo, Inseon Jang, ChungHyun Ahn, Jeongil Seo, Hong-Goo Kang "A pitch-synchronous speech analysis and synthesis method for DNN-SPSS system" in 21th International Conference on Digital Signal Processing (DSP), 2016
254 International Conference Eunwoo Song, Frank K. Soong, Hong-Goo Kang "Improved Time-Frequency Trajectory Excitation Vocoder for DNN-Based Speech Synthesis" in INTERSPEECH, 2016
253 Domestic Conference Hyeonjoo Kang, Young-sun Joo, Wonsuk Jun, Hong-goo Kang "다층신경망 기반 다중 화자 음성변환 시스템" in 한국음향학회 제 33회 음성통신 및 신호처리 학술대회, 2016
252 International Conference Eunwoo Song, Hong-Goo Kang "Multi-class learning algorithm for deep neural network-based statistical parametric speech synthesis" in EUSIPCO, 2016
251 Domestic Conference Min-jae Hwang, JeeSok Lee, Misuk Lee, and Hong-Goo Kang "사전 분석법을 통한 스프레드 스펙트럼 기반 오디오 워터마킹 알고리즘의 성능 향상" in 한국음향학회 제 33회 음성통신 및 신호처리 학술대회, 2016
250 Domestic Conference 문현기, 박영철, 윤대희 "위상 일치와 가변 지수 감쇄 가중치 방법이 적용된 가상 저음 시스템" in 2016 년 한국방송·미디어공학회 하계학술대회, 2016
249 Domestic Conference 문현기, 박영철, 윤대희 "주파수 종속 믹싱 타임 추정 기법 분석" in 한국음향학회 춘계학술대회, 2016
248 Domestic Conference 서지호, 박영철, 윤대희 "헤드폰/이어폰 환경에서의 제약 최적화 기반의 효율적인 피드백 능동 소음 제어 필터 설계 알고리즘" in 한국음향학회 춘계학술대회, 2016
247 Domestic Conference 박규태, 박영철, 윤대희 "가변 가중치 곡선을 적용한 가상저음시스템" in 한국음향학회 춘계학술대회, 2016