Browse State-of-the-Art › Genre classification
Genre classification
51 papers with code · 2 benchmarks · 6 datasets archive 2025-07-28
Genre classification is the process of grouping objects together based on defined similarities such as shape, pixel, location, or intensity.
Description from the archive archive 2025-07-28.
Benchmarks archive 2025-07-28
2 leaderboard tables shown for this task, 2 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted.
| Dataset | Best model (first row in archive order) | Paper | Code | Syntology | Compare |
|---|---|---|---|---|---|
| Book Cover Dataset (2 rows) | AlexNet | Judging a Book By its Cover | code | — | Compare |
| FMA (1 row) | cnn | Multi-label Music Genre Classification from Audio, Text, and... | code | — | Compare |
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
6 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
No subtask under this task in the archive's task tree.
Parent tasks archive 2025-07-28
Most implemented papers archive 2025-07-28
30 shown of 51 papers with code (130 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
7 Feb 2017 9 repositories listed Syntology ran 1 of 1 samples · 0 unverifiedThe Gated Multimodal Unit (GMU) model is intended to be used as an internal unit in a neural network architecture whose purpose is to find an intermediate representation based on a combination of data from different…
-
28 Oct 2016 4 repositories listedBook covers communicate information to potential readers, but can that same information be learned by computers?
-
24 Oct 2023 2 repositories listedIn this paper, we study whether music source separation can be used as a pre-training strategy for music representation learning, targeted at music classification tasks.
-
26 Oct 2022 2 repositories listedThe QCNN is a circuit model inspired by the architecture of Convolutional Neural Networks (CNNs).
-
17 Oct 2021 2 repositories listedAlong with the evolution of music technology, a large number of styles, or "subgenres," of Electronic Dance Music(EDM) have emerged in recent years.
-
10 Jun 2021 2 repositories listed Syntology ran 3 of 7 samples · 4 unverifiedInspired by the success of pre-training models in natural language processing, in this paper, we develop MusicBERT, a large-scale pre-trained model for music understanding.
-
21 Oct 2019 2 repositories listedThe importance of repetitions in music is well-known.
-
28 May 2019 2 repositories listedIn this paper, we evaluate the impact of frame selection on automatic music genre classification in a bag of frames scenario.
-
27 Feb 2018 2 repositories listedHere, we propose a new method that combines knowledge of human perception study in music genre classification and the neurophysiology of the auditory system.
-
22 Dec 2017 2 repositories listedDeep learning has been demonstrated its effectiveness and efficiency in music genre classification.
-
29 Dec 2024 1 repository listedThis study demonstrates that the modern generation of Large Language Models (LLMs, such as GPT-4) suffers from the same out-of-domain (OOD) performance gap observed in prior research on pre-trained Language Models…
-
10 Oct 2024 1 repository listedThe proposed approach splits audio signals into 20 ms chunks and processes them through convolutional feature encoders, a transformer encoder, and additional layers for coding audio units and generating feature vectors.
-
1 Oct 2024 1 repository listed Syntology ran 1 of 1 samples · 0 unverifiedOur findings suggest that music theory concepts are discernible within foundation models and that the degree to which they are detectable varies by model size and layer.
-
13 Sep 2024 1 repository listedOver the years, Music Information Retrieval (MIR) has proposed various models pretrained on large amounts of music data.
-
27 Nov 2023 1 repository listedWhile performance of many text classification tasks has been recently improved due to Pre-trained Language Models (PLMs), in this paper we show that they still suffer from a performance gap when the underlying…
-
12 Oct 2023 1 repository listedFirstly we retrieve the relevant embedding from the knowledge graph by utilizing group relations in metadata and then integrate it with other modalities.
-
26 Jun 2023 1 repository listedArt forms such as movies and television (TV) dramas are reflections of the real world, which have attracted much attention from the multimodal learning community recently.
-
14 Feb 2023 1 repository listedContrastive learning constitutes an emerging branch of self-supervised learning that leverages large amounts of unlabeled data, by learning a latent space, where pairs of different views of the same sample are…
-
8 Dec 2022 1 repository listedRating a video based on its content is an important step for classifying video age categories.
-
Effective Audio Classification Network Based on Paired Inverse Pyramid Structure and Dense MLP Block5 Nov 2022 1 repository listedRecently, massive architectures based on Convolutional Neural Network (CNN) and self-attention mechanisms have become necessary for audio classification.
-
4 Nov 2022 1 repository listedTo overcome these limitations, the present study explores and applies efficient transfer learning methods in the audio domain.
-
2 Nov 2022 1 repository listedIn this work, we introduce a novel method for leveraging pre-trained models for low-resource (music) classification based on the concept of Neural Model Reprogramming (NMR).
-
20 Oct 2022 1 repository listed Syntology ran 2 of 11 samples · 9 unverifiedLongform media such as movies have complex narrative structures, with events spanning a rich variety of ambient visual scenes.
-
14 Oct 2022 1 repository listedIn particular, we present an extensive evaluation of the transferability of ConvNet and Transformer models pretrained on ImageNet and Kinetics to Trailers12k, a new manually-curated movie trailer dataset composed of 12,…
-
9 Sep 2022 1 repository listedImbalanced music genre classification is a crucial task in the Music Information Retrieval (MIR) field for identifying the long-tail, data-poor genre based on the related music audio segments, which is very prevalent in…
-
25 Aug 2022 1 repository listedDue to the increased demand for music streaming/recommender services and the recent developments of music information retrieval frameworks, Music Genre Classification (MGC) has attracted the community's attention.
-
25 Aug 2022 1 repository listedIn this work, we explore cross-modal learning in an attempt to bridge audio and language in the music domain.
-
3 Aug 2022 1 repository listedLarge pre-trained neural networks are ubiquitous and critical to the success of many downstream tasks in natural language processing and computer vision.
-
27 Jun 2022 1 repository listedIn this work, we investigate the uncertainty calibration for deep audio classifiers.
-
31 May 2022 1 repository listedMassively multilingual sentence representation models, e.
Syntology lines on 4 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections