Papers

Perfect Match: Self-Supervised Embeddings for Cross-Modal Retrieval

International Journal

2016~2020

작성자

이진영

작성일

2020-03-01 22:17

조회

1630

Authors : Soo-Whan Chung, Joon Son Chung, Hong Goo Kang

Year : 2020

Publisher / Conference : IEEE Journal of Selected Topics in Signal Processing

Volume : 14, issue 3

This paper proposes a new strategy for learning effective cross-modal joint embeddings using self-supervision. We set up the problem as one of cross-modal retrieval, where the objective is to find the most relevant data in one domain given input in another. The method builds on the recent advances in learning representations from cross-modal self-supervision using contrastive or binary cross-entropy loss functions. To investigate the robustness of the proposed learning strategy across multi-modal applications, we perform experiments for two applications - audio-visual synchronisation and cross-modal biometrics. The audio-visual synchronisation task requires temporal correspondence between modalities to obtain joint representation of phonemes and visemes, and the cross-modal biometrics task requires common speakers representations given their face images and audio tracks. Experiments show that the performance of systems trained using proposed method far exceed that of existing methods on both tasks, whilst allowing significantly faster training.

« A Study on Acoustic Parameter Selection Strategies to Improve Deep Learning-Based Speech Synthesis

Improving LPCNet-based Text-to-Speech with Linear Prediction-structured Mixture Density Network »

목록보기

전체 355

International Conference

Seyun Um, Sangshin Oh, Kyungguen Byun, Inseon Jang, ChungHyun Ahn, Hong-Goo Kang "Emotional Speech Synthesis with Rich and Granularized Control" in ICASSP, 2020

Perfect Match: Self-Supervised Embeddings for Cross-Modal Retrieval

Previous

Sister Lab.

Yonsei University

Academic Website