Papers

Speaker-Discriminative Attractors for Robust Continuous Speech Separation

International Journal
2021~
작성자
김지현
작성일
2026-08-21 18:30
조회
16
Authors : Jihyun Kim, Doyeon Kim, Hong-Goo Kang

Year : 2026

Publisher / Conference : TASLP

Volume : 34

Page : 3916-3929

Research area : Speech Signal Processing, Speech Recognition, Source Separation

Presentation/Publication date : 20206.07.31

Presentation : None

Continuous speech separation (CSS) aims to separate unbounded audio streams under low-latency, chunk-wise streaming constraints. It further requires the separated outputs to remain speaker-consistent over time despite dynamically changing speaker activity. Most existing approaches that achieve strong speaker-attributed performance on long-form recordings rely on modular pipelines that combine separation with external speaker embedding models and offline clustering, which precludes real-time deployment. In this paper, we propose a single-channel CSS frame- work that performs end-to-end speaker tracking within a unified streaming network without requiring any external components, by explicitly enforcing cross-chunk speaker consistency using speaker- discriminative attractors. Our approach adopts a two-stage training strategy. In the first stage, the model is trained on full utterances to learn attractor representations that jointly perform separation and speaker existence modeling. In the second stage, the attractor network is frozen and refined for streaming inference by introducing identity-aware cross-segment alignment. To achieve robust inter-chunk consistency, we introduce Cross-Segment Contrastive Regularization (CSCR), which injects speaker identity supervision by aligning chunk-level embeddings with persistent speaker prototypes stored in an external memory queue. We further propose the Relational Tracking Resolver (RTR), a relational assignment algorithm that robustly matches current speakers to registered identities based on relational geometry, preventing identity overwriting during silence and turn-taking. Experimental results on WSJ0-mix and LibriCSS demonstrate that the proposed framework significantly improves both separation quality and long-term permutation stability under low-latency streaming conditions.
전체 387
387 International Journal Jihyun Kim, Doyeon Kim, Hong-Goo Kang "Speaker-Discriminative Attractors for Robust Continuous Speech Separation" in TASLP, vol.34, pp.3916-3929, 2026
386 Domestic Conference 김효민, 이지현, 장인선, 강홍구 "사후 의미 증류를 이용한 유한 스칼라 양자화 기반 이단계 음성 토크나이저" in 한국방송·미디어공학회 2026년 하계학술대회, 2026
385 Domestic Conference 신재훈, 장인선, 강홍구 "효율적인 신경망 기반 오디오 코덱을 위한 잔차 오토인코딩 및 연속형 오토인코더의 잠재 표현 증류" in 한국방송·미디어공학회 2026년 하계학술대회, 2026
384 International Conference Jihyun Lee, Jiahao Li, Woojin Chung, Yan Lu, Hong-Goo Kang "AudioSketch: Controllable Image-to-Audio Generation via Semantic-Temporal Energy Modulation" in EUSIPCO, 2026
383 International Conference Sangmin Lee, Woojin Chung, Woongjib Choi, Hong-Goo Kang "MoLGE: Mixture of Language Group Experts for Efficient Scaling of Massively Multilingual Speech Recognition" in Conference On Language Modeling (COLM), 2026
382 International Conference Sangmin Lee, Eekgyun Ahn, Woongjib Choi, Hong-Goo Kang "UR-BERT: Scaling Text Encoders for Massively Multilingual TTS Through Universal Romanization and Speech Token Prediction" in INTERSPEECH, 2026
381 International Conference Seyun Um, Doyeon Kim, Hong-Goo Kang "HANUI: Harnessing Distributional Discrepancies for Singing Voice Deepfake Detection" in in IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP), 2026
380 International Conference Miseul Kim, Soo jin Park, Kyungguen Byun, Hyeon-Kyeong Shin, Sunkuk Moon, Shuhua Zhang, Erik Visser "Mitigating Intra-Speaker Variability in Diarization with Style-Controllable Speech Augmentation" in in IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP), 2026
379 International Conference Woongjib Choi, Sangmin Lee, Hyungseob Lim, Hong-Goo Kang "UniverSR: Unified and Versatile Audio Super-Resolution via Vocoder-Free Flow Matching" in IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP), 2026
378 International Journal Hyeonjin Cha, Seyun Um, Miseul Kim, Changhwan Kim, Seungshin Lee, Hong-Goo Kang "Content-Aware Style Augmentation for Zero-Shot Voice Conversion With Short Target Speech" in IEEE Signal Processing Letters, vol.33, pp.66-70, 2025