{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/unsupervised-learning-of-semantic-audio","title":"Unsupervised Learning of Semantic Audio Representations","arxiv_id":"1711.02209","date":"2017-11-06","proceeding":null,"authors":["Aren Jansen","Manoj Plakal","Ratheet Pandya","Daniel P. W. Ellis","Shawn Hershey","Jiayang Liu","R. Channing Moore","Rif A. Saurous"],"abstract":"Even in the absence of any explicit semantic annotation, vast collections of\naudio recordings provide valuable information for learning the categorical\nstructure of sounds. We consider several class-agnostic semantic constraints\nthat apply to unlabeled nonspeech audio: (i) noise and translations in time do\nnot change the underlying sound category, (ii) a mixture of two sound events\ninherits the categories of the constituents, and (iii) the categories of events\nin close temporal proximity are likely to be the same or related. Without\nlabels to ground them, these constraints are incompatible with classification\nloss functions. However, they may still be leveraged to identify geometric\ninequalities needed for triplet loss-based training of convolutional neural\nnetworks. The result is low-dimensional embeddings of the input spectrograms\nthat recover 41% and 84% of the performance of their fully-supervised\ncounterparts when applied to downstream query-by-example sound retrieval and\nsound event classification tasks, respectively. Moreover, in\nlimited-supervision settings, our unsupervised embeddings double the\nstate-of-the-art classification performance.","url_abs":"http://arxiv.org/abs/1711.02209v1","url_pdf":"http://arxiv.org/pdf/1711.02209v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"audio-classification","task_name":"Audio Classification"},{"task_slug":"classification-1","task_name":"Classification"},{"task_slug":"classification","task_name":"General Classification"},{"task_slug":"retrieval","task_name":"Retrieval"},{"task_slug":null,"task_name":"Triplet"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/audio-classification-on-audioset","task":"Audio Classification","dataset":"AudioSet","model":"Triplet","rank_in_archive_order":51,"of":51,"metrics":{"Test mAP":"0.244"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1711.02209","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}