Papers

Consideration of Varying Training Lengths for Short-Duration Speaker Verification

International Conference
작성자
dsp
작성일
2023-09-12 14:23
조회
358
Authors : WooSeok Ko, Seyun Um, Zhenyu Piao, Hong-goo Kang

Year : 2023

Publisher / Conference : APSIPA ASC

Research area : Speech Signal Processing, Speaker Recognition

Presentation : Poster

We present an efficient training scheme for speaker verification (SV) networks in short-duration speech input scenarios. We analyze the effects of varying training lengths on SV performance, with a particular focus on short utterances. Despite the high demand for short-duration SV in real-world applications, state-of-the-art SV systems have primarily been evaluated on long utterances, and little research has been conducted on shortduration SV. By considering the innate characteristics of SV architectures and the performance discrepancies associated with varying training data lengths, we propose a training scheme that accounts for varying length conditions. We categorize speaker characteristics as coarse-grained and fine-grained features and demonstrate that training models to learn both features can result in length-robust speaker embeddings. Our proposed training scheme improves model performance by 28.7% and 37.9% in terms of equal error rate on short-duration speech scenarios compared to baseline models.
전체 355
152 International Conference Hyewon Han, Naveen Kumar "A cross-talk robust multichannel VAD model for multiparty agent interactions trained using synthetic re-recordings" in Hands-free Speech Communication and Microphone Arrays (HSCMA, Satellite workshop in ICASSP), 2024
151 International Conference Yanjue Song, Doyeon Kim, Nilesh Madhu, Hong-Goo Kang "On the Disentanglement and Robustness of Self-Supervised Speech Representations" in International Conference on Electronics, Information, and Communication (ICEIC) (*awarded Best Paper), 2024
150 International Conference Yeona Hong, Miseul Kim, Woo-Jin Chung, Hong-Goo Kang "Contextual Learning for Missing Speech Automatic Speech Recognition" in International Conference on Electronics, Information, and Communication (ICEIC), 2024
149 International Conference Juhwan Yoon, Seyun Um, Woo-Jin Chung, Hong-Goo Kang "SC-ERM: Speaker-Centric Learning for Speech Emotion Recognition" in International Conference on Electronics, Information, and Communication (ICEIC), 2024
148 International Conference Hejung Yang, Hong-Goo Kang "On Fine-Tuning Pre-Trained Speech Models With EMA-Target Self-Supervised Loss" in ICASSP, 2024
147 International Conference Hong-Goo Kang, W. Bastiaan Kleijn, Jan Skoglund, Michael Chinen "Convolutional Transformer for Neural Speech Coding" in Audio Engineering Society Convention, 2023
146 International Conference Hong-Goo Kang, Jan Skoglund, W. Bastiaan Kleijn, Andrew Storus, Hengchin Yeh "A High-Rate Extension to Soundstream" in IEEE Workshop on Applications of Signal Processing to Audio and Acoustics (WASPAA), 2023
145 International Conference Zhenyu Piao, Hyungseob Lim, Miseul Kim, Hong-goo Kang "PDF-NET: Pitch-adaptive Dynamic Filter Network for Intra-gender Speaker Verification" in APSIPA ASC, 2023
144 International Conference WooSeok Ko, Seyun Um, Zhenyu Piao, Hong-goo Kang "Consideration of Varying Training Lengths for Short-Duration Speaker Verification" in APSIPA ASC, 2023
143 International Conference Miseul Kim, Zhenyu Piao, Jihyun Lee, Hong-Goo Kang "BrainTalker: Low-Resource Brain-to-Speech Synthesis with Transfer Learning using Wav2Vec 2.0" in The IEEE-EMBS International Conference on Biomedical and Health Informatics (BHI), 2023