Browse State-of-the-Art › Instrument Recognition
Instrument Recognition
26 papers with code · 3 benchmarks · 4 datasets archive 2025-07-28
Benchmarks archive 2025-07-28
3 leaderboard tables shown for this task, 3 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted.
| Dataset | Best model (first row in archive order) | Paper | Code | Syntology | Compare |
|---|---|---|---|---|---|
| NSynth (7 rows) | M2D-CLAP | M2D2: Exploring General-purpose Audio-Language Representations Beyond CLAP | code | — | Compare |
| OpenMIC-2018 (5 rows) | DyMN-L | Dynamic Convolutional Neural Networks as Efficient Pre-trained Audio Models | code | — | Compare |
| IRMAS (1 row) | SVM | Predominant Musical Instrument Classification based on Spectral Features | code | — | Compare |
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
4 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
No subtask under this task in the archive's task tree.
Most implemented papers archive 2025-07-28
26 shown of 26 papers with code (39 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
26 Apr 2022 4 repositories listedSelf-supervised learning (SSL) learns knowledge from a large amount of unlabeled data, and then transfers the knowledge to a specific problem with a limited number of labeled data.
-
22 Apr 2021 3 repositories listedTo utilize automated methods in clinical settings, it is crucial to design lightweight models with low latency such that they can be integrated with low-end endoscope hardware devices.
-
28 Mar 2025 2 repositories listedIn the second stage, it learns CLAP features using the audio features learned from the LLM-based embeddings.
-
7 Jun 2023 2 repositories listedIn order to tackle both clip-level and frame-level tasks, this paper proposes Audio Teacher-Student Transformer (ATST), with a clip-level version (named ATST-Clip) and a frame-level version (named ATST-Frame),…
-
11 Oct 2021 2 repositories listed Syntology ran 3 of 3 samples · 0 unverifiedHowever, one of the main shortcomings of transformer models, compared to the well-established CNNs, is the computational complexity.
-
26 Jun 2025 1 repository listedIdentifying instrument activities within audio excerpts is vital in music information retrieval, with significant implications for music cataloging and discovery.
-
17 Feb 2025 1 repository listedTherefore, we propose a new method: MAsked latenT Prediction And Classification (MATPAC), which is trained with two pretext tasks solved jointly.
-
1 Nov 2024 1 repository listedThis paper introduces an extendable modular system that compiles a range of music feature extraction models to aid music information retrieval research.
-
25 Jul 2024 1 repository listedWe present an evaluation and analysis of two-tower systems for zero-shot instrument recognition and a detailed analysis of the properties of the pre-joint and joint embeddings spaces.
-
26 Jun 2024 1 repository listed Syntology ran 4 of 4 samples · 0 unverifiedOf the very few current systems that support source separation beyond this setup, most continue to rely on an inflexible decoder setup that can only support a fixed pre-defined set of stems.
-
24 Oct 2023 1 repository listedAudio Spectrogram Transformers are excellent at exploiting large datasets, creating powerful pre-trained models that surpass CNNs when fine-tuned on downstream tasks.
-
20 Jul 2023 1 repository listedThis approach allows representations derived for one task to be applied to another, and can result in high accuracy with less stringent training data requirements for the downstream task.
-
30 Jun 2023 1 repository listedMusic classification has been one of the most popular tasks in the field of music information retrieval.
-
29 Jun 2023 1 repository listedIt focuses on the visualization of the occurrence of phases, phase transitions, instruments, and instrument combinations across sets.
-
12 Jul 2022 1 repository listedIn audio classification, differentiable auditory filterbanks with few parameters cover the middle ground between hard-coded spectrograms and raw audio.
-
19 Aug 2021 1 repository listedThis paper propose a traditional Chinese music dataset for training model and performance evaluation, named ChMusic.
-
24 Jul 2021 1 repository listedConstructing an embedding space for musical instrument sounds that can meaningfully represent new and unseen instruments is important for downstream music generation tasks such as multi-instrument synthesis and timbre…
-
14 Jul 2021 1 repository listedDeep learning work on musical instrument recognition has generally focused on instrument classes for which we have abundant data.
-
26 May 2021 1 repository listedAs state-of-the-art CNN architectures-in computer vision and other domains-tend to go deeper in terms of number of layers, their RF size increases and therefore they degrade in performance in several audio…
-
23 Oct 2020 1 repository listedAdditionally, we provide a baseline for the segmentation of the GI tools to promote research and algorithm development.
-
30 Nov 2019 1 repository listedThis work aims to examine one of the cornerstone problems of Musical Instrument Retrieval (MIR), in particular, instrument classification.
-
28 Nov 2019 1 repository listedInstrument classification is one of the fields in Music Information Retrieval (MIR) that has attracted a lot of research interest.
-
9 Jul 2019 1 repository listedWhile the automatic recognition of musical instruments has seen significant progress, the task is still considered hard for music featuring multiple instruments as opposed to single instrument recordings.
-
4 Dec 2018 1 repository listedResults: We build a baseline tracker on top of the CNN model and demonstrate that our approach based on the ConvLSTM outperforms the baseline in tool presence detection, spatial localization, and motion tracking by over…
-
31 May 2016 1 repository listedWe train our network from fixed-length music excerpts with a single-labeled predominant instrument and estimate an arbitrary number of predominant instruments from an audio signal with a variable length.
-
17 Nov 2015 1 repository listedTraditional methods to tackle many music information retrieval tasks typically follow a two-step architecture: feature engineering followed by a simple learning algorithm.
Syntology lines on 2 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections