CORTEXA
← Browse
arxivcs.SDcs.LG2026-07-01

Evaluating Pretrained Music Embeddings for Cross-Performance Jazz Standard Recognition

Çağrı Eser

Recognizing jazz standards from audio is a challenging form of tune-level music retrieval: different performances of the same standard may vary in tempo, key, arrangement, instrumentation, improvisational content, and even whether the head melody is present. We study this problem using a curated subset of the Jazz Trio Database designed for cross-performance standard recognition. We compare a from-scratch trained Harmonic CNN baseline against frozen pretrained music representations from recent music understanding foundation models, using both supervised probing and nearest-neighbor retrieval. Our results suggest that from-scratch spectrogram models overfit strongly to training performances, while pretrained embeddings provide better top-$k$ results but are sensitive to performer identity, which can be partially reduced with a lightweight contrastive projection. Our findings motivate jazz standard recognition as a useful stress test for music representation models and as a step toward retrieval-based standard identification. Project page: https://github.com/cagries/tipofmyear.

View free PDFSource page

Related papers

arxivcs.LGcs.AIcs.SD2026-07-02

Decomposer: Learning to Decompile Symbolic Music to Programs

Yewon Kim, Apurva Gandhi, David Chung, Graham Neubig, Chris Donahue

Musical performance involves executing a set of high-level musical instructions, yet recovering those instructions from the performance is a challenging inverse problem. We present Decomposer, a post-training framework for symbolic music decompilation: the task of recovering exec…

View free PDFSource page
arxivcs.SDcs.LG2026-07-09

MulTTiPop: A Multitrack Transcription Dataset for Pop Music

Nathan Pruyne, Benjamin Stoler, William Chen, Chien-yu Huang, Shinji Watanabe, Chris Donahue

We present MulTTiPop, a dataset of pop music segments and their associated multitrack MIDI recordings for the evaluation of automatic music transcription models. MulTTiPop contains 572 segments of popular music totaling 3.5 hours of audio, and contains songs from diverse genres a…

View free PDFSource page
arxivcs.SDcs.AIcs.LGeess.ASeess.SP2026-06-29

BEST-RQ-2: Contextualize-Then-Predict, a Two-Step Approach for Self-Supervised Audio Representations

Ludovic K. Tuncay, Etienne Labbé, Thomas Pellegrini

Self-supervised learning enables audio representations that transfer across domains and tasks. We present BEST-RQ-2, an evolution of BEST-RQ that retains frozen randomprojection-based discrete targets while introducing a two-step contextualize-then-predict pretraining scheme. A V…

View free PDFSource page
arxivcs.SDcs.LGeess.ASq-bio.QM2026-07-03

Adaptive Loss Balancing for Multi-Task Bioacoustic Classification of Bird Species and Call Types

Paria Vali Zadeh, Sven Tomforde

Reliable analysis of bird vocalisations in passive acoustic monitoring requires models handling multiple, imbalanced annotation targets. We extend BirdCallNet for joint species and call-type classification on the long-tailed WiWa dataset and investigate how task-loss balancing in…

View free PDFSource page