Papers › ATST: Audio Representation Learning with Teacher-Student Transformer

ATST: Audio Representation Learning with Teacher-Student Transformer

26 Apr 2022arXiv:2204.12076archive 2025-07-28

Xian Li, Xiaofei Li

Self-supervised learning (SSL) learns knowledge from a large amount of unlabeled data, and then transfers the knowledge to a specific problem with a limited number of labeled data. SSL has achieved promising results in various domains. This work addresses the problem of segment-level general audio SSL, and proposes a new transformer-based teacher-student SSL model, named ATST. A transformer encoder is developed on a recently emerged teacher-student baseline scheme, which largely improves the modeling capability of pre-training. In addition, a new strategy for positive pair creation is designed to fully leverage the capability of transformer. Extensive experiments have been conducted, and the proposed model achieves the new state-of-the-art results on almost all of the downstream tasks.

PaperPDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

Audio-WestlakeU/ATST-SED mentioned on GitHubpytorchMIT report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Audio ClassificationInstrument RecognitionRepresentation LearningSelf-Supervised Audio ClassificationSelf-Supervised LearningSpeaker IdentificationSpoken Command Recognition

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Audio Classification Balanced Audio Set Base (ours) Mean AP 37.4 #5 of 8 Archive leaderboard report
Speaker Identification VoxCeleb1 ATST Base (ours) Accuracy 94.3 #6 of 12 Archive leaderboard report
Speaker Identification VoxCeleb1 ATST Base (ours) Top-1 (%) 94.3 #6 of 12 Archive leaderboard report
Spoken Command Recognition Speech Command v2 Base (ours) Accuracy 98.0 #2 of 4 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

BYOL

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections