Methods › Computer Vision › Likelihood-Based Generative Models › VQ-VAE
VQ-VAE
Introduced by Aaron van den Oord et al. in Neural Discrete Representation Learning
archive 2025-07-28 Description, source and code snippet are the archive's method entry.
VQ-VAE is a type of variational autoencoder that uses vector quantisation to obtain a discrete latent representation. It differs from VAEs in two key ways: the encoder network outputs discrete, rather than continuous, codes; and the prior is learnt rather than static. In order to learn a discrete latent representation, ideas from vector quantisation (VQ) are incorporated. Using the VQ method allows the model to circumvent issues of posterior collapse - where the latents are ignored when they are paired with a powerful autoregressive decoder - typically observed in the VAE framework. Pairing these representations with an autoregressive prior, the model can generate high quality images, videos, and speech as well as doing high quality speaker conversion and unsupervised learning of phonemes.
Papers archive 2025-07-28
30 shown of 197, newest first. Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code; it is per sample and not a correctness claim.
-
DuetGen: Music Driven Two-Person Dance Generation via Hierarchical Masked Modeling 23 Jun 2025 · 1 repository · arXiv:2506.18680
-
Policy-Based Trajectory Clustering in Offline Reinforcement Learning 10 Jun 2025 · 0 repositories · arXiv:2506.09202
-
STAR: Learning Diverse Robot Skill Abstractions through Rotation-Augmented Vector Quantization 4 Jun 2025 · 1 repository · arXiv:2506.03863
-
VesselGPT: Autoregressive Modeling of Vascular Geometry 19 May 2025 · 1 repository · arXiv:2505.13318
-
M3G: Multi-Granular Gesture Generator for Audio-Driven Full-Body Human Motion Synthesis 13 May 2025 · 0 repositories · arXiv:2505.08293
-
Towards Foundation Models for Experimental Readout Systems Combining Discrete and Continuous Data 13 May 2025 · 1 repository · arXiv:2505.08736
-
Ego4o: Egocentric Human Motion Capture and Understanding from Multi-Modal Input 11 Apr 2025 · 0 repositories · arXiv:2504.08449
-
Instruction-Guided Autoregressive Neural Network Parameter Generation 2 Apr 2025 · 0 repositories · arXiv:2504.02012
-
Make Some Noise: Towards LLM audio reasoning and generation using sound tokens 28 Mar 2025 · 0 repositories · arXiv:2503.22275
-
AGIR: Assessing 3D Gait Impairment with Reasoning based on LLMs 23 Mar 2025 · 0 repositories · arXiv:2503.18141
-
Improving Autoregressive Image Generation through Coarse-to-Fine Token Prediction 20 Mar 2025 · 0 repositories · arXiv:2503.16194
-
A Foundation Model for Patient Behavior Monitoring and Suicide Detection 19 Mar 2025 · 0 repositories · arXiv:2503.15221
-
GenM³: Generative Pretrained Multi-path Motion Model for Text Conditional Human Motion Generation 19 Mar 2025 · 0 repositories · arXiv:2503.14919
-
BioSerenity-E1: a self-supervised EEG model for medical applications 13 Mar 2025 · 0 repositories · arXiv:2503.10362
-
UniGenX: Unified Generation of Sequence and Structure with Autoregressive Diffusion 9 Mar 2025 · 0 repositories · arXiv:2503.06687
-
Enhancing Spoken Discourse Modeling in Language Models Using Gestural Cues 5 Mar 2025 · 0 repositories · arXiv:2503.03474
-
CAPS: Context-Aware Priority Sampling for Enhanced Imitation Learning in Autonomous Driving 3 Mar 2025 · 0 repositories · arXiv:2503.01650
-
Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement 11 Feb 2025 · 0 repositories · arXiv:2502.07243
-
Token Assorted: Mixing Latent and Text Tokens for Improved Language Model Reasoning 5 Feb 2025 · 0 repositories · arXiv:2502.03275
-
Patch-aware Vector Quantized Codebook Learning for Unsupervised Visual Defect Detection 15 Jan 2025 · 0 repositories · arXiv:2501.09187
-
HOIGPT: Learning Long-Sequence Hand-Object Interaction with Language Models 1 Jan 2025 · 0 repositories
-
LLM-driven Multimodal and Multi-Identity Listening Head Generation 1 Jan 2025 · 0 repositories
-
TAR3D: Creating High-Quality 3D Assets via Next-Part Prediction 22 Dec 2024 · 0 repositories · arXiv:2412.16919
-
CSL-L2M: Controllable Song-Level Lyric-to-Melody Generation Based on Conditional Transformer with Fine-Grained Lyric and Musical Controls 13 Dec 2024 · 0 repositories · arXiv:2412.09887
-
Adaptive²: Adaptive Domain Mining for Fine-grained Domain Adaptation Modeling 11 Dec 2024 · 0 repositories · arXiv:2412.08198
-
SweetTokenizer: Semantic-Aware Spatial-Temporal Tokenizer for Compact Visual Discretization 11 Dec 2024 · 0 repositories · arXiv:2412.10443
-
VQalAttent: a Transparent Speech Generation Pipeline based on Transformer-learned VQ-VAE Latent Space 22 Nov 2024 · 0 repositories · arXiv:2411.14642
-
Augmenting Training Data with Vector-Quantized Variational Autoencoder for Classifying RF Signals 23 Oct 2024 · 0 repositories · arXiv:2410.18283
-
Gaussian Mixture Vector Quantization with Aggregated Categorical Posterior 14 Oct 2024 · 0 repositories · arXiv:2410.10180
-
InterMask: 3D Human Interaction Generation via Collaborative Masked Modelling 13 Oct 2024 · 1 repository · arXiv:2410.10010
Tasks archive 2025-07-28
20 shown of 171 tasks the archive attaches to papers tagged with this method, by distinct papers. A task without a page in the catalog is plain text.
Usage over time archive 2025-07-28
Components: the archive holds no method-to-method composition, so PwC's Components table cannot be rebuilt; the Papers list carries no Results column for the same reason (the archive does not join its leaderboard rows to method tags).
Categories archive 2025-07-28
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections