Papers › Efficient Large-scale Audio Tagging via Transformer-to-CNN Knowledge Distillation

Efficient Large-scale Audio Tagging via Transformer-to-CNN Knowledge Distillation

9 Nov 2022arXiv:2211.04772archive 2025-07-28

Florian Schmid, Khaled Koutini, Gerhard Widmer

Audio Spectrogram Transformer models rule the field of Audio Tagging, outrunning previously dominating Convolutional Neural Networks (CNNs). Their superiority is based on the ability to scale up and exploit large-scale datasets such as AudioSet. However, Transformers are demanding in terms of model size and computational requirements compared to CNNs. We propose a training procedure for efficient CNNs based on offline Knowledge Distillation (KD) from high-performing yet complex transformers. The proposed training schema and the efficient CNN design based on MobileNetV3 results in models outperforming previous solutions in terms of parameter and computational efficiency and prediction performance. We provide models of different complexity levels, scaling from low-complexity models up to a new state-of-the-art performance of .483 mAP on AudioSet. Source Code available at: https://github.com/fschmid56/EfficientAT

PaperPDFCodeCode Syntology ran

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

For agents, Syntology's MCP tool lists every function and class Syntology harvested from this paper and whether it ran (how to connect): get_harvested_code_for_paper(arxiv_id="2211.04772")

Code

Syntology Ran 1 of 2 code samples harvested from 1 repository linked to this paper; 1 has no recorded run. Of those that ran: 1 ran · our draft was wrong.

By repository: community (archive-listed): 2 samples from 1 repository, 1 ran. The run record, sample by sample. “Ran” means executed on a synthesized input, not that the code is correct or reproduces the paper.

fschmid56/efficientat officialmentioned in papermentioned on GitHubpytorchMIT report
fschmid56/efficientat_hear mentioned on GitHubpytorchMIT report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

2 samples harvested; 1 ran; 0 honoured the contract we drafted; 1 has no recorded run. Read from Syntology's graph 2026-09-24; that is when this build read the record, not when the samples ran.

1ran · our draft was wrong
1unverified

Licence: 0 of the 2 samples are pointer only, meaning Syntology does not serve that copy's text. This page shows no code text for any sample; each one links to its file in the repository.

Harvested from fschmid56/efficientat_hear. “Ran” means the sample executed on a synthesized input. It does not mean the output is correct, and nothing here reproduces the paper's results. “Honoured” and “violated” refer to a contract Syntology drafted from the code itself; “our draft was wrong” and “fixture could not drive it” are failures of Syntology's instrument, not of the code.

Each sample ends with its code_sha256, Syntology's identity for that exact code. An agent fetches the stored sample with Syntology's MCP tool get_code(code_sha256="…") (how to connect); click an identity to copy that call.

Repository labels, per sample. official repository: The archive marks this repository official for the paper. named in the paper: The archive records that the paper mentions this repository; it is not marked official. community (archive-listed): In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper. found in paper text by Syntology: Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted. community: Not in the archive's code links for this paper; a community repository Syntology harvested. Samples from a repository marked official are listed first. Licence labels name the repository's licence as recorded at harvest. “Pointer only” means Syntology does not serve that copy's text, for one of four reasons: no licence file was found; the licence was not identified; the licence is recorded as permissive but that copy's record is not marked cleared; or the licence is outside the permissive list Syntology serves text under (MIT, Apache-2.0, BSD and similar). Some licences outside that list permit redistribution, such as WTFPL, and GPL-3.0 under its conditions; they are simply not on the list. Hover a licence label for the reason. File links open the file on GitHub at the default branch, which may have changed since the harvest.

get_scene_embeddings fschmid56/efficientat_hear/hear_mn/mn01_all_b_avg_max_time_pool.py community (archive-listed) ran · our draft was wrong MIT (permissive) · 4e582be02ea360fa · report
get_timestamp_embeddings fschmid56/efficientat_hear/hear_mn/mn01_all_b_avg_max_time_pool.py community (archive-listed) unverified MIT (permissive) · 0e956046855f5a5e · report

Tasks

Audio ClassificationAudio TaggingComputational EfficiencyKnowledge Distillation

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Audio Classification AudioSet mn40_as (Ensemble) Test mAP 0.498 #10 of 51 Archive leaderboard report
Audio Classification AudioSet mn40_as (Single) Test mAP 0.483 #24 of 51 Archive leaderboard report
Audio Classification ESC-50 mn40_as Accuracy (5-fold) 97.45 #6 of 29 Archive leaderboard report
Audio Classification ESC-50 mn40_as PRE-TRAINING DATASET AudioSet #6 of 29 Archive leaderboard report
Audio Classification ESC-50 mn40_as Top-1 Accuracy 97.45 #6 of 29 Archive leaderboard report
Audio Tagging AudioSet mn40_as (Ensemble) mean average precision 0.498 #2 of 11 Archive leaderboard report
Audio Tagging AudioSet mn40_as (Single) mean average precision 0.483 #6 of 11 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

1x1 ConvolutionAbsolute Position EncodingsAdamAttentionAverage PoolingBPEBatch NormalizationConvolutionDense ConnectionsDepthwise ConvolutionDepthwise Separable ConvolutionDropoutGlobal Average PoolingHard SwishInverted Residual BlockKnowledge DistillationLabel SmoothingLayer NormalizationLinear LayerMulti-Head AttentionPointwise ConvolutionPosition-Wise Feed-Forward LayerReLUReLU6Residual ConnectionSigmoid ActivationSoftmaxSqueeze-and-Excitation BlockTransformer

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections