{"url":"/dataset/cal500","name":"CAL500","full_name":"Computer Audition Lab 500","description_markdown":"**CAL500** (**Computer Audition Lab 500**) is a dataset aimed for evaluation of music information retrieval systems. It consists of 502 songs picked from western popular music. The audio is represented as a time series of the first 13 Mel-frequency cepstral coefficients (and their first and second derivatives) extracted by sliding a 12 ms half-overlapping short-time window over the waveform of each song. Each song has been annotated by at least 3 people with 135 musically-relevant concepts spanning six semantic categories:\r\n\r\n* 29 instruments were annotated as present in the song or not,\r\n* 22 vocal characteristics were annotated as relevant to the singer or not,\r\n* 36 genres,\r\n* 18 emotions were rated on a scale from one to three (e.g., ``not happy\", ``neutral\", ``happy\"),\r\n* 15 song concepts describing the acoustic qualities of the song, artist and recording (e.g., tempo, energy, sound quality),\r\n* 15 usage terms (e.g., \"I would listen to this song while driving, sleeping, etc.\").\r\n\r\nSource: [http://calab1.ucsd.edu/~datasets/cal500/details_cal500.txt](http://calab1.ucsd.edu/~datasets/cal500/details_cal500.txt)\r\nAudio Source: [http://calab1.ucsd.edu/~datasets/cal500/cal500data/](http://calab1.ucsd.edu/~datasets/cal500/cal500data/)","description_withheld":null,"homepage":"http://calab1.ucsd.edu/~datasets/","introduced_date":"2008-01-01","introduced_date_note":null,"introduced_by":{"paper":null,"title":"Semantic Annotation and Retrieval of Music and Sound Effects","first_author":null,"url":"https://doi.org/10.1109/TASL.2007.913750"},"license":{"name":"Custom","url":"http://calab1.ucsd.edu/~datasets/cal500/use_cal500.txt"},"modalities":[{"name":"Audio","url":"/datasets/modality/audio"},{"name":"Tabular","url":"/datasets/modality/tabular"}],"tasks":[{"name":"Multi-Label Classification","url":"/task/multi-label-classification","datasets_with_task":"/datasets/task/multi-label-classification"},{"name":"Multi-Task Learning","url":"/task/multi-task-learning","datasets_with_task":"/datasets/task/multi-task-learning"},{"name":"Matrix Completion","url":"/task/matrix-completion","datasets_with_task":"/datasets/task/matrix-completion"}],"languages":[],"variants":["CAL500"],"data_loaders":[],"num_papers_in_archive":21,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}