{"url":"/method/jukebox","slug":"jukebox","name":"Jukebox","full_name":"Jukebox","full_name_withheld":false,"description_markdown":"**Jukebox** is a model that generates music with singing in the raw audio domain. It tackles the long context of raw audio using a multi-scale [VQ-VAE](https://paperswithcode.com/method/vq-vae) to compress it to discrete codes, and modeling those using [autoregressive Transformers](https://paperswithcode.com/methods/category/autoregressive-transformers). It can condition on artist and genre to steer the musical and vocal style, and on unaligned lyrics to make the singing more controllable.\r\n\r\nThree separate VQ-VAE models are trained with different temporal resolutions. At each level, the input audio is segmented and encoded into latent vectors $\\mathbf{h}\\_{t}$, which are then quantized to the closest codebook vectors $\\mathbf{e}\\_{z\\_{t}}$. The code $z\\_{t}$ is a discrete representation of the audio that we later train our prior on. The decoder takes the sequence of codebook vectors and reconstructs the audio. The top level learns the highest degree of abstraction, since it is encoding longer audio per token while keeping the codebook size the same. Audio can be reconstructed using the codes at any one of the abstraction levels, where the least abstract bottom-level codes result in the highest-quality audio.","description_state":"present","introduced_year":null,"introduced_by":{"title":null,"paper":null,"first_author":null,"n_authors":0,"url_abs":null,"archive_paper_url":null},"source":{"url":"https://arxiv.org/abs/2005.00341v1","title":"Jukebox: A Generative Model for Music","url_on_a_paper_host":true},"code_snippet_url":null,"code_snippet_url_on_a_code_host":false,"categories":[{"area":"Audio","area_id":"audio","collection":"Generative Audio Models","url":"/methods/category/generative-audio-models","pwc_aliases":[]}],"n_papers_tagged":13,"archive_num_papers":null,"papers_newest_first":[{"paper":"/paper/comparative-analysis-of-pretrained-audio","title":"Comparative Analysis of Pretrained Audio Representations in Music Recommender Systems","date":"2024-09-13","arxiv_id":"2409.08987","n_code_links":1,"syntology":null},{"paper":null,"title":"An End-to-End Approach for Chord-Conditioned Song Generation","date":"2024-09-10","arxiv_id":"2409.06307","n_code_links":0,"syntology":null},{"paper":null,"title":"From Audio Encoders to Piano Judges: Benchmarking Performance Understanding for Solo Piano","date":"2024-07-05","arxiv_id":"2407.04518","n_code_links":0,"syntology":null},{"paper":null,"title":"A Novel Audio Representation for Music Genre Identification in MIR","date":"2024-04-01","arxiv_id":"2404.01058","n_code_links":0,"syntology":null},{"paper":null,"title":"Advancements in Generative AI: A Comprehensive Review of GANs, GPT, Autoencoders, Diffusion Model, and Transformers","date":"2023-11-17","arxiv_id":"2311.10242","n_code_links":0,"syntology":null},{"paper":null,"title":"MAP-Music2Vec: A Simple and Effective Baseline for Self-Supervised Music Audio Representation Learning","date":"2022-12-05","arxiv_id":"2212.02508","n_code_links":0,"syntology":null},{"paper":"/paper/melody-transcription-via-generative-pre","title":"Melody transcription via generative pre-training","date":"2022-12-04","arxiv_id":"2212.01884","n_code_links":1,"syntology":null},{"paper":"/paper/edge-editable-dance-generation-from-music","title":"EDGE: Editable Dance Generation From Music","date":"2022-11-19","arxiv_id":"2211.10658","n_code_links":1,"syntology":{"ran":7,"of":11,"unverified":4,"pointer_only":0}},{"paper":null,"title":"Evaluating Deep Music Generation Methods Using Data Augmentation","date":"2021-12-31","arxiv_id":"2201.00052","n_code_links":0,"syntology":null},{"paper":"/paper/transfer-learning-with-jukebox-for-music","title":"Transfer Learning with Jukebox for Music Source Separation","date":"2021-11-28","arxiv_id":"2111.14200","n_code_links":1,"syntology":null},{"paper":"/paper/unsupervised-source-separation-by-steering","title":"Unsupervised Source Separation By Steering Pretrained Music Models","date":"2021-10-25","arxiv_id":"2110.13071","n_code_links":1,"syntology":null},{"paper":"/paper/codified-audio-language-modeling-learns","title":"Codified audio language modeling learns useful representations for music information retrieval","date":"2021-07-12","arxiv_id":"2107.05677","n_code_links":1,"syntology":{"ran":3,"of":3,"unverified":0,"pointer_only":0}},{"paper":"/paper/jukebox-a-generative-model-for-music-1","title":"Jukebox: A Generative Model for Music","date":"2020-04-30","arxiv_id":"2005.00341","n_code_links":12,"syntology":null}],"papers_shown":13,"tasks":[{"task":"/task/information-retrieval","name":"Information Retrieval","papers":4},{"task":"/task/music-information-retrieval","name":"Music Information Retrieval","papers":4},{"task":"/task/genre-classification","name":"Genre classification","papers":3},{"task":"/task/music-generation","name":"Music Generation","papers":3},{"task":"/task/retrieval","name":"Retrieval","papers":3},{"task":"/task/audio-source-separation","name":"Audio Source Separation","papers":2},{"task":"/task/music-genre-classification","name":"Music Genre Classification","papers":2},{"task":"/task/music-tagging","name":"Music Tagging","papers":2},{"task":"/task/transfer-learning","name":"Transfer Learning","papers":2},{"task":"/task/attribute","name":"Attribute","papers":1},{"task":"/task/audio-generation","name":"Audio Generation","papers":1},{"task":"/task/benchmarking","name":"Benchmarking","papers":1},{"task":"/task/chord-recognition","name":"Chord Recognition","papers":1},{"task":"/task/code-generation","name":"Code Generation","papers":1},{"task":"/task/data-augmentation","name":"Data Augmentation","papers":1},{"task":"/task/diversity","name":"Diversity","papers":1},{"task":"/task/emotion-recognition","name":"Emotion Recognition","papers":1},{"task":"/task/key-detection","name":"Key Detection","papers":1},{"task":"/task/language-modeling","name":"Language Modeling","papers":1},{"task":"/task/language-modelling","name":"Language Modelling","papers":1}],"tasks_shown":20,"n_tasks":30,"usage_by_year":[{"year":"2020","papers":1},{"year":"2021","papers":4},{"year":"2022","papers":3},{"year":"2023","papers":1},{"year":"2024","papers":4}],"row_source":"embedded","archive":{"source":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","archive_url":"https://paperswithcode.com/method/jukebox"},"syntology_read_at":"2026-09-24T18:15:14+00:00"}