Papers

SC-ERM: Speaker-Centric Learning for Speech Emotion Recognition

International Conference

2021~

작성자

dsp

작성일

2024-01-22 16:12

조회

3017

Authors : Juhwan Yoon, Seyun Um, Woo-Jin Chung, Hong-Goo Kang

Year : 2024

Publisher / Conference : International Conference on Electronics, Information, and Communication (ICEIC)

Research area : Speech Signal Processing, Etc

Presentation/Publication date : 2024.01.29

Presentation : Poster

We propose a novel deep learning-based model for speech emotion recognition, SC-ERM, which focuses on speakercentric learning. This model effectively estimates emotions and demonstrates the ability to generalize to unseen speakers. Our proposed model utilizes speaker-specific emotion characteristics in two steps: first, it extracts emotion representations using an emotion encoder, and second, it employs speaker-centric learning by incorporating speaker style embeddings as a condition through a speaker mask generator. We evaluate our model’s performance using an emotional dataset and find that it demonstrates outstanding performance in recognizing emotional states. Notably, it achieves a 9.2% relative improvement in accuracy compared to the baseline when classifying emotions for speakers not seen during training. Overall, our model demonstrates promising performance in accurately identifying emotions across a range of emotional expressions, irrespective of the speakers involved.

« On Fine-Tuning Pre-Trained Speech Models With EMA-Target Self-Supervised Loss

Contextual Learning for Missing Speech Automatic Speech Recognition »

목록보기

전체 372

362	Domestic Conference	김병현, 강홍구, 장인선 "저지연 조건하의 심층신경망 기반 음성 압축" in 한국방송·미디어공학회 2024년 하계학술대회, 2024
361	International Conference	Miseul Kim, Soo-Whan Chung, Youna Ji, Hong-Goo Kang, Min-Seok Choi "Speak in the Scene: Diffusion-based Acoustic Scene Transfer toward Immersive Speech Generation" in INTERSPEECH, 2024
360	International Conference	Seyun Um, Doyeon Kim, Hong-Goo Kang "PARAN: Variational Autoencoder-based End-to-End Articulation-to-Speech System for Speech Intelligibility" in INTERSPEECH, 2024
359	International Conference	Jihyun Kim, Stijn Kindt, Nilesh Madhu, Hong-Goo Kang "Enhanced Deep Speech Separation in Clustered Ad Hoc Distributed Microphone Environments" in INTERSPEECH, 2024
358	International Conference	Woo-Jin Chung, Hong-Goo Kang "Speaker-Independent Acoustic-to-Articulatory Inversion through Multi-Channel Attention Discriminator" in INTERSPEECH, 2024
357	International Conference	Juhwan Yoon, Woo Seok Ko, Seyun Um, Sungwoong Hwang, Soojoong Hwang, Changhwan Kim, Hong-Goo Kang "UNIQUE : Unsupervised Network for Integrated Speech Quality Evaluation" in INTERSPEECH, 2024
356	International Conference	Yanjue Song, Doyeon Kim, Hong-Goo Kang, Nilesh Madhu "Spectrum-aware neural vocoder based on self-supervised learning for speech enhancement" in EUSIPCO, 2024
355	International Conference	Hyewon Han, Naveen Kumar "A cross-talk robust multichannel VAD model for multiparty agent interactions trained using synthetic re-recordings" in Hands-free Speech Communication and Microphone Arrays (HSCMA, Satellite workshop in ICASSP), 2024
354	International Conference	Yanjue Song, Doyeon Kim, Nilesh Madhu, Hong-Goo Kang "On the Disentanglement and Robustness of Self-Supervised Speech Representations" in International Conference on Electronics, Information, and Communication (ICEIC) (*awarded Best Paper), 2024
353	International Conference	Yeona Hong, Miseul Kim, Woo-Jin Chung, Hong-Goo Kang "Contextual Learning for Missing Speech Automatic Speech Recognition" in International Conference on Electronics, Information, and Communication (ICEIC), 2024

SC-ERM: Speaker-Centric Learning for Speech Emotion Recognition

Previous

Sister Lab.

Yonsei University

Academic Website