Skip to content

Opening book details…

Can I read Music Similarity Representation Learning Focusing on Individual Instruments with Source Separation and Human Preference on EtoBox?

Music Similarity Representation Learning Focusing on Individual Instruments with Source Separation and Human Preference by Imamura, Takehiro; Hashizume, Yuka; Huang, Wen-Chin; Toda, Tomoki is a scholarly article available to read on EtoBox.

What is Music Similarity Representation Learning Focusing on Individual Instruments with Source Separation and Human Preference about?

This paper proposes music similarity representation learning (MSRL) based on individual instrument sounds (InMSRL) utilizing music source separation (MSS) and human preference without requiring clean instrument sounds during inference. We propose three methods that effectively improve performance. First, we introduce end-to-end fine-tuning (E2E-FT) for the Cascade approach that sequentially performs MSS and music similarity feature extraction. E2E-FT allows the model to minimize the adverse effects of a separation error on the feature extraction. Second, we propose multi-task learning for the Direct approach that directly extracts disentangled music similarity features using a single music similarity feature extractor. Multi-task learning, which is based on the disentangled music similarity feature extraction and MSS based on reconstruction with disentangled music similarity features, further enhances instrument feature disentanglement. Third, we employ perception-aware fine-tuning (PAFT). PAFT utilizes human preference, allowing the model to perform InMSRL aligned with human perceptual similarity. We conduct experimental evaluations and demonstrate that 1) E2E-FT for Cascade signi

Author
Imamura, Takehiro; Hashizume, Yuka; Huang, Wen-Chin; Toda, Tomoki
Published
2025
Language
EN

More by Imamura, Takehiro; Hashizume, Yuka; Huang, Wen-Chin; Toda, Tomoki

Browse all works by Imamura, Takehiro; Hashizume, Yuka; Huang, Wen-Chin; Toda, Tomoki