Papers

Two-Stage Refinement of Magnitude and Complex Spectra for Real-Time Speech Enhancement

International Journal

2021~

작성자

dsp

작성일

2022-11-09 14:37

조회

1230

Authors : Jinyoung Lee, Hong-Goo Kang

Year : 2022

Publisher / Conference : IEEE Signal Processing Letters

Volume : 29

Page : 2188-2192

Research area : Speech Signal Processing, Speech Enhancement

Presentation/Publication date : 17 October 2022

Presentation : None

In this letter, we propose a two-stage network for performing speech enhancement that predicts magnitude spectra in the first stage and complex spectra in the second stage. To maximize the model’s performance at each stage, we propose two convolutional modules: magnitude spectral masking (MSM) and complex spectra refinement (CSR). Each module is designed to take
into account the specific characteristics of the signal type it handles. The MSM estimates multiplicative masks to remove noise in the magnitude component of the convolutional features, and the CSR refines the complex component of the convolutional features using additive features. By using these modules, our proposed two-stage enhancement model shows higher performance than previously proposed state-of-the-art algorithms. In addition, the number of parameters of our model is only 2.63 million, and it can operate in real time thanks to its causal characteristics and low computational complexity.

« 엔트로피 모델을 활용한 심층 신경망 기반 오디오 압축 모델 최적화

Ultrathin crystalline-silicon-based strain gauges with deep learning algorithms for silent speech interfaces »

목록보기

전체 355

5	International Journal	Taemin Kim, Yejee Shin, Kyowon Kang, Kiho Kim, Gwanho Kim, Yunsu Byeon, Hwayeon Kim, Yuyan Gao, Jeong Ryong Lee, Geonhui Son, Taeseong Kim, Yohan Jun, Jihyun Kim, Jinyoung Lee, Seyun Um, Yoohwan Kwon, Byung Gwan Son, Myeongki Cho, Mingyu Sang, Jongwoon Shin, Kyubeen Kim, Jungmin Suh, Heekyeong Choi, Seokjun Hong, Huanyu Cheng, Hong-Goo Kang, Dosik Hwang & Ki Jun Yu "Ultrathin crystalline-silicon-based strain gauges with deep learning algorithms for silent speech interfaces" in Nature Communications, vol.13, 2022
4	International Conference	Zainab Alhakeem, Yoohwan Kwon, Hong-Goo Kang "Disentangled Representations for Arabic Dialect Identification based on Supervised Clustering with Triplet Loss" in EUSIPCO, 2021
3	International Conference	Yoohwan Kwon, Hee-Soo Heo, Bong-Jin Lee, Joon Son Chung "The ins and outs of speaker recognition: lessons from VoxSRC 2020" in ICASSP, 2021
2	International Conference	Seong Min Kye, Yoohwan Kwon, Joon Son Chung "Cross Attentive Pooling for Speaker Verification" in IEEE Spoken Language Technology Workshop (SLT), 2020
1	International Conference	Yoohwan Kwon, Soo-Whan Chung, Hong-Goo Kang "Intra-Class Variation Reduction of Speaker Representation in Disentanglement Framework" in INTERSPEECH, 2020

Two-Stage Refinement of Magnitude and Complex Spectra for Real-Time Speech Enhancement

Previous

Sister Lab.

Yonsei University

Academic Website