Datasets › CAL500

CAL500 (Computer Audition Lab 500)

Introduced in Semantic Annotation and Retrieval of Music and Sound Effects1 Jan 2008 archive 2025-07-28

CAL500 (Computer Audition Lab 500) is a dataset aimed for evaluation of music information retrieval systems. It consists of 502 songs picked from western popular music. The audio is represented as a time series of the first 13 Mel-frequency cepstral coefficients (and their first and second derivatives) extracted by sliding a 12 ms half-overlapping short-time window over the waveform of each song. Each song has been annotated by at least 3 people with 135 musically-relevant concepts spanning six semantic categories:

  • 29 instruments were annotated as present in the song or not,
  • 22 vocal characteristics were annotated as relevant to the singer or not,
  • 36 genres,
  • 18 emotions were rated on a scale from one to three (e.g., not happy",neutral", ``happy"),
  • 15 song concepts describing the acoustic qualities of the song, artist and recording (e.g., tempo, energy, sound quality),
  • 15 usage terms (e.g., "I would listen to this song while driving, sleeping, etc.").

Source: http://calab1.ucsd.edu/~datasets/cal500/details_cal500.txt Audio Source: http://calab1.ucsd.edu/~datasets/cal500/cal500data/

Benchmarks archive 2025-07-28

No leaderboard in the archive resolves to this dataset.

Papers archive 2025-07-28

No paper in the archive has a leaderboard row on this dataset; the archive counts 21 papers for it but never published that list.

Dataset loaders archive 2025-07-28

No loader listed in the archive.

Tasks archive 2025-07-28

License archive 2025-07-28

Custom

Modalities archive 2025-07-28

Languages archive 2025-07-28

No language tagged.

Variants archive 2025-07-28

  • CAL500

1 variant name, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections