Papers

Contextual Learning for Missing Speech Automatic Speech Recognition

International Conference

2021~

작성자

dsp

작성일

2024-01-22 16:15

조회

204

Authors : Yeona Hong, Miseul Kim, Woo-Jin Chung, Hong-Goo Kang

Year : 2024

Publisher / Conference : International Conference on Electronics, Information, and Communication (ICEIC)

Research area : Speech Signal Processing, Speech Recognition

Presentation/Publication date : 2024.01.29

Presentation : Poster

—In this paper, we present an automatic speech recognition (ASR) system that is capable of decoding complete transcriptions from speech even in cases where there are missing segments in the audio. To predict complete transcriptions from speech that may have missing segments, we utilize a contextual learning approach inspired by recent language model training approaches, in which our model leverages surrounding speech segments as cues for the prediction. Our model consists of two modules: a contextual feature extractor designed with the structure of wav2vec 2.0, and a projection layer. We further explore various masking lengths for model training so as to optimally benefit the ASR system without compromising its performance. Our proposed methodology demonstrates highquality ASR performance on missing speech segments of various lengths, ranging from a word error rate (WER) of 4.7% on 0.25 seconds segments to 18.5% on 1 second segments.

« SC-ERM: Speaker-Centric Learning for Speech Emotion Recognition

On the Disentanglement and Robustness of Self-Supervised Speech Representations »

목록보기

전체 355

38	International Conference	Hyewon Han, Naveen Kumar "A cross-talk robust multichannel VAD model for multiparty agent interactions trained using synthetic re-recordings" in Hands-free Speech Communication and Microphone Arrays (HSCMA, Satellite workshop in ICASSP), 2024
37	International Conference	Yanjue Song, Doyeon Kim, Nilesh Madhu, Hong-Goo Kang "On the Disentanglement and Robustness of Self-Supervised Speech Representations" in International Conference on Electronics, Information, and Communication (ICEIC) (*awarded Best Paper), 2024
36	International Conference	Yeona Hong, Miseul Kim, Woo-Jin Chung, Hong-Goo Kang "Contextual Learning for Missing Speech Automatic Speech Recognition" in International Conference on Electronics, Information, and Communication (ICEIC), 2024
35	International Conference	Juhwan Yoon, Seyun Um, Woo-Jin Chung, Hong-Goo Kang "SC-ERM: Speaker-Centric Learning for Speech Emotion Recognition" in International Conference on Electronics, Information, and Communication (ICEIC), 2024
34	International Conference	Hejung Yang, Hong-Goo Kang "On Fine-Tuning Pre-Trained Speech Models With EMA-Target Self-Supervised Loss" in ICASSP, 2024
33	International Journal	Zainab Alhakeem, Se-In Jang, Hong-Goo Kang "Disentangled Representations in Local-Global Contexts for Arabic Dialect Identification" in Transactions on Audio, Speech, and Language Processing, 2024
32	International Conference	Hong-Goo Kang, W. Bastiaan Kleijn, Jan Skoglund, Michael Chinen "Convolutional Transformer for Neural Speech Coding" in Audio Engineering Society Convention, 2023
31	International Conference	Hong-Goo Kang, Jan Skoglund, W. Bastiaan Kleijn, Andrew Storus, Hengchin Yeh "A High-Rate Extension to Soundstream" in IEEE Workshop on Applications of Signal Processing to Audio and Acoustics (WASPAA), 2023
30	International Conference	Zhenyu Piao, Hyungseob Lim, Miseul Kim, Hong-goo Kang "PDF-NET: Pitch-adaptive Dynamic Filter Network for Intra-gender Speaker Verification" in APSIPA ASC, 2023
29	International Journal	Hyungchan Yoon, Changhwan Kim, Seyun Um, Hyun-Wook Yoon, Hong-Goo Kang "SC-CNN: Effective Speaker Conditioning Method for Zero-Shot Multi-Speaker Text-to-Speech Systems" in IEEE Signal Processing Letters, vol.30, pp.593-597, 2023

Contextual Learning for Missing Speech Automatic Speech Recognition

Previous

Sister Lab.

Yonsei University

Academic Website