Methods › Computer Vision › Likelihood-Based Generative Models › VQ-VAE › Papers, page 2
VQ-VAE
Papers archive 2025-07-28
archive papers tagged: 197 · with a code link: 82 · where Syntology ran a sample: 35 (32 with a run with no instrument failure, 3 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (35 of 197 tagged: 32 with a run with no instrument failure, 3 where every run was a failure of Syntology's instrument)
Page 2 of 2: papers 101 to 197 of 197, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
LayoutDiffuse: Adapting Foundational Diffusion Models for Layout-to-Image Generation 16 Feb 2023 · 0 repositories · arXiv:2302.08908
-
Vector Quantized Wasserstein Auto-Encoder 12 Feb 2023 · 0 repositories · arXiv:2302.05917
-
Language Quantized AutoEncoders: Towards Unsupervised Text-Image Alignment 2 Feb 2023 · 1 repository · arXiv:2302.00902
-
Improving Statistical Fidelity for Neural Image Compression with Implicit Local Likelihood Models 26 Jan 2023 · 1 repository · arXiv:2301.11189
-
T2M-GPT: Generating Human Motion from Textual Descriptions with Discrete Representations 15 Jan 2023 · 1 repository · arXiv:2301.06052Syntology official (archive's flag): 5 ran · 5 ran (of which 4 constructed an object rather than computing a result; 5 with no instrument failure: 1 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 7 harvested samples)
-
Latent Autoregressive Source Separation 9 Jan 2023 · 1 repository · arXiv:2301.08562
-
CodeTalker: Speech-Driven 3D Facial Animation with Discrete Motion Prior 6 Jan 2023 · 1 repository · arXiv:2301.02379Syntology official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 8 with no instrument failure: 0 honoured, 1 violated, 7 with no contract checked; 0 where Syntology's instrument failed) · 5 unverified (of 13 harvested samples)
-
All in Tokens: Unifying Output Space of Visual Tasks via Soft Token 5 Jan 2023 · 1 repository · arXiv:2301.02229
-
Generating Human Motion From Textual Descriptions With Discrete Representations 1 Jan 2023 · 0 repositories
-
Image Compression with Product Quantized Masked Image Modeling 14 Dec 2022 · 0 repositories · arXiv:2212.07372
-
Generating Holistic 3D Human Motion from Speech 8 Dec 2022 · 3 repositories · arXiv:2212.04420
-
MAP-Music2Vec: A Simple and Effective Baseline for Self-Supervised Music Audio Representation Learning 5 Dec 2022 · 0 repositories · arXiv:2212.02508
-
Melody transcription via generative pre-training 4 Dec 2022 · 1 repository · arXiv:2212.01884
-
Accelerating Antimicrobial Peptide Discovery with Latent Structure 28 Nov 2022 · 1 repository · arXiv:2212.09450
-
Homology-constrained vector quantization entropy regularizer 25 Nov 2022 · 1 repository · arXiv:2211.14363Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
Disentangled Feature Learning for Real-Time Neural Speech Coding 22 Nov 2022 · 0 repositories · arXiv:2211.11960
-
EDGE: Editable Dance Generation From Music 19 Nov 2022 · 1 repository · arXiv:2211.10658Syntology official (archive's flag): 10 ran · 10 ran (of which 0 constructed an object rather than computing a result; 8 with no instrument failure: 3 honoured, 1 violated, 4 with no contract checked; 2 where Syntology's instrument failed) · 1 unverified (of 11 harvested samples) · 1 pointer-only (licence)
-
SSGVS: Semantic Scene Graph-to-Video Synthesis 11 Nov 2022 · 0 repositories · arXiv:2211.06119
-
Learning Control by Iterative Inversion 3 Nov 2022 · 0 repositories · arXiv:2211.01724
-
Character-Centric Story Visualization via Visual Planning and Token Alignment 16 Oct 2022 · 2 repositories · arXiv:2210.08465Syntology official: no sample here; runs from other or unrecorded repositories · 12 ran (of which 7 constructed an object rather than computing a result; 9 with no instrument failure: 1 honoured, 0 violated, 8 with no contract checked; 3 where Syntology's instrument failed) · 9 unverified (of 21 harvested samples)
-
JukeDrummer: Conditional Beat-aware Audio-domain Drum Accompaniment Generation via Transformer VQ-VAE 12 Oct 2022 · 1 repository · arXiv:2210.06007
-
Race Bias Analysis of Bona Fide Errors in face anti-spoofing 11 Oct 2022 · 0 repositories · arXiv:2210.05366
-
LMQFormer: A Laplace-Prior-Guided Mask Query Transformer for Lightweight Snow Removal 10 Oct 2022 · 1 repository · arXiv:2210.04787
-
A deep learning approach for detection and localization of leaf anomalies 7 Oct 2022 · 0 repositories · arXiv:2210.03558
-
Efficient Planning in a Compact Latent Action Space 22 Aug 2022 · 1 repository · arXiv:2208.10291Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 2 harvested samples)
-
Hierarchical Residual Learning Based Vector Quantized Variational Autoencoder for Image Reconstruction and Generation 9 Aug 2022 · 1 repository · arXiv:2208.04554
-
Augmenting Vision Language Pretraining by Learning Codebook with Visual Semantics 31 Jul 2022 · 0 repositories · arXiv:2208.00475
-
Cancer Subtyping by Improved Transcriptomic Features Using Vector Quantized Variational Autoencoder 20 Jul 2022 · 0 repositories · arXiv:2207.09783
-
Diffsound: Discrete Diffusion Model for Text-to-sound Generation 20 Jul 2022 · 1 repository · arXiv:2207.09983
-
Latent-Domain Predictive Neural Speech Coding 18 Jul 2022 · 0 repositories · arXiv:2207.08363
-
PILC: Practical Image Lossless Compression with an End-to-end GPU Oriented Neural Framework 10 Jun 2022 · 0 repositories · arXiv:2206.05279
-
Draft-and-Revise: Effective Image Generation with Contextual RQ-Transformer 9 Jun 2022 · 0 repositories · arXiv:2206.04452
-
DiVAE: Photorealistic Images Synthesis with Denoising Diffusion Decoder 1 Jun 2022 · 0 repositories · arXiv:2206.00386
-
SQ-VAE: Variational Bayes on Discrete Representation with Self-annealed Stochastic Quantization 16 May 2022 · 1 repository · arXiv:2205.07547Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 3 where Syntology's instrument failed) · 1 unverified (of 5 harvested samples)
-
Learning to Listen: Modeling Non-Deterministic Dyadic Facial Motion 18 Apr 2022 · 0 repositories · arXiv:2204.08451
-
Unconditional Image-Text Pair Generation with Multimodal Cross Quantizer 15 Apr 2022 · 1 repository · arXiv:2204.07537
-
Analysis of Voice Conversion and Code-Switching Synthesis Using VQ-VAE 28 Mar 2022 · 0 repositories · arXiv:2203.14640
-
Pixel VQ-VAEs for Improved Pixel Art Representation 23 Mar 2022 · 1 repository · arXiv:2203.12130
-
Diffusion bridges vector quantized Variational AutoEncoders 10 Feb 2022 · 1 repository · arXiv:2202.04895Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
Robust Vector Quantized-Variational Autoencoder 4 Feb 2022 · 0 repositories · arXiv:2202.01987
-
Evaluating Deep Music Generation Methods Using Data Augmentation 31 Dec 2021 · 0 repositories · arXiv:2201.00052
-
Global Context with Discrete Diffusion in Vector Quantised Modelling for Image Generation 3 Dec 2021 · 0 repositories · arXiv:2112.01799
-
Transfer Learning with Jukebox for Music Source Separation 28 Nov 2021 · 1 repository · arXiv:2111.14200
-
A model of semantic completion in generative episodic memory 26 Nov 2021 · 0 repositories · arXiv:2111.13537
-
Learning source-aware representations of music in a discrete latent space 26 Nov 2021 · 0 repositories · arXiv:2111.13321
-
Layered Controllable Video Generation 24 Nov 2021 · 0 repositories · arXiv:2111.12747
-
L-Verse: Bidirectional Generation Between Image and Text 22 Nov 2021 · 1 repository · arXiv:2111.11133
-
Unsupervised Source Separation By Steering Pretrained Music Models 25 Oct 2021 · 1 repository · arXiv:2110.13071
-
Discrete Acoustic Space for an Efficient Sampling in Neural Text-To-Speech 24 Oct 2021 · 0 repositories · arXiv:2110.12539
-
Taming Visually Guided Sound Generation 17 Oct 2021 · 3 repositories · arXiv:2110.08791Syntology official (archive's flag): 1 ran · 7 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 7 harvested samples)
-
KaraSinger: Score-Free Singing Voice Synthesis with VQ-VAE using Mel-spectrograms 8 Oct 2021 · 0 repositories · arXiv:2110.04005
-
An Unsupervised Video Game Playstyle Metric via State Discretization 3 Oct 2021 · 1 repository · arXiv:2110.00950
-
Generating Antimicrobial Peptides from Latent Secondary Structure Space 29 Sep 2021 · 0 repositories
-
Applying the Information Bottleneck Principle to Prosodic Representation Learning 5 Aug 2021 · 0 repositories · arXiv:2108.02821
-
RockGPT: Reconstructing three-dimensional digital rocks from single two-dimensional slice from the perspective of video generation 5 Aug 2021 · 0 repositories · arXiv:2108.03132
-
Codified audio language modeling learns useful representations for music information retrieval 12 Jul 2021 · 1 repository · arXiv:2107.05677Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
VIMPAC: Video Pre-Training via Masked Token Prediction and Contrastive Learning 21 Jun 2021 · 1 repository · arXiv:2106.11250Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Neural Distributed Source Coding 5 Jun 2021 · 1 repository · arXiv:2106.02797Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified (of 7 harvested samples)
-
Factorising Meaning and Form for Intent-Preserving Paraphrasing 31 May 2021 · 1 repository · arXiv:2105.15053
-
CogView: Mastering Text-to-Image Generation via Transformers 26 May 2021 · 4 repositories · arXiv:2105.13290Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 1 where Syntology's instrument failed) · 2 unverified (of 6 harvested samples)
-
Exploring Disentanglement with Multilingual and Monolingual VQ-VAE 4 May 2021 · 1 repository · arXiv:2105.01573
-
VideoGPT: Video Generation using VQ-VAE and Transformers 20 Apr 2021 · 3 repositories · arXiv:2104.10157
-
Generating Diverse Structure for Image Inpainting With Hierarchical VQ-VAE 18 Mar 2021 · 2 repositories · arXiv:2103.10022Syntology community repositories only · 4 ran (of which 3 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 5 harvested samples) · 5 pointer-only (licence)
-
Wav2vec-C: A Self-supervised Model for Speech Representation Learning 9 Mar 2021 · 0 repositories · arXiv:2103.08393
-
Generating Images with Sparse Representations 5 Mar 2021 · 2 repositories · arXiv:2103.03841Syntology 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 2 honoured, 0 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
Predicting Video with VQVAE 2 Mar 2021 · 1 repository · arXiv:2103.01950
-
Enhancing into the codec: Noise Robust Speech Coding with Vector-Quantized Autoencoders 12 Feb 2021 · 0 repositories · arXiv:2102.06610
-
Self-Supervised VQ-VAE for One-Shot Music Style Transfer 10 Feb 2021 · 1 repository · arXiv:2102.05749
-
VideoGen: Generative Modeling of Videos using VQ-VAE and Transformers 1 Jan 2021 · 0 repositories
-
Anomaly detection through latent space restoration using vector-quantized variational autoencoders 12 Dec 2020 · 0 repositories · arXiv:2012.06765
-
Multi-Instrumentalist Net: Unsupervised Generation of Music from Body Movements 7 Dec 2020 · 0 repositories · arXiv:2012.03478
-
Learning Vector Quantized Shape Code for Amodal Blastomere Instance Segmentation 2 Dec 2020 · 0 repositories · arXiv:2012.00985
-
A Comparison of Discrete Latent Variable Models for Speech Representation Learning 24 Oct 2020 · 0 repositories · arXiv:2010.14230
-
Learning Disentangled Phone and Speaker Representations in a Semi-Supervised VQ-VAE Paradigm 21 Oct 2020 · 1 repository · arXiv:2010.10727Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 2 harvested samples)
-
End-to-End Text-to-Speech using Latent Duration based on VQ-VAE 19 Oct 2020 · 0 repositories · arXiv:2010.09602
-
The Utility of Decorrelating Colour Spaces in Vector Quantised Variational Autoencoders 30 Sep 2020 · 1 repository · arXiv:2009.14487
-
Incorporating Reinforced Adversarial Learning in Autoregressive Image Generation 20 Jul 2020 · 0 repositories · arXiv:2007.09923
-
Generating Annotated High-Fidelity Images Containing Multiple Coherent Objects 22 Jun 2020 · 1 repository · arXiv:2006.12150
-
UWSpeech: Speech to Speech Translation for Unwritten Languages 14 Jun 2020 · 0 repositories · arXiv:2006.07926
-
Transformer VQ-VAE for Unsupervised Unit Discovery and Speech Synthesis: ZeroSpeech 2020 Challenge 24 May 2020 · 0 repositories · arXiv:2005.11676
-
Vector-quantized neural networks for acoustic unit discovery in the ZeroSpeech 2020 challenge 19 May 2020 · 2 repositories · arXiv:2005.09409
-
Robust Training of Vector Quantized Bottleneck Models 18 May 2020 · 1 repository · arXiv:2005.08520Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 7 harvested samples)
-
DiscreTalk: Text-to-Speech as a Machine Translation Problem 12 May 2020 · 0 repositories · arXiv:2005.05525
-
Jukebox: A Generative Model for Music 30 Apr 2020 · 12 repositories · arXiv:2005.00341
-
Adversarial Feature Learning and Unsupervised Clustering based Speech Synthesis for Found Data with Acoustic and Textual Noise 28 Apr 2020 · 0 repositories · arXiv:2004.13595
-
Hierarchical Quantized Autoencoders 19 Feb 2020 · 1 repository · arXiv:2002.08111
-
Neuromorphologicaly-preserving Volumetric data encoding using VQ-VAE 13 Feb 2020 · 0 repositories · arXiv:2002.05692
-
Semi-supervised Grasp Detection by Representation Learning in a Vector Quantized Latent Space 23 Jan 2020 · 0 repositories · arXiv:2001.08477
-
Low Bit-Rate Speech Coding with VQ-VAE and a WaveNet Decoder 14 Oct 2019 · 0 repositories · arXiv:1910.06464
-
MelGAN: Generative Adversarial Networks for Conditional Waveform Synthesis 8 Oct 2019 · 21 repositories · arXiv:1910.06711Syntology official (archive's flag): 1 ran · 6 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 1 violated, 1 with no contract checked; 4 where Syntology's instrument failed) · 1 unverified (of 7 harvested samples) · 1 pointer-only (licence)
-
VQVAE Unsupervised Unit Discovery and Multi-scale Code2Spec Inverter for Zerospeech Challenge 2019 27 May 2019 · 0 repositories · arXiv:1905.11449
-
Towards a better understanding of Vector Quantized Autoencoders 1 May 2019 · 0 repositories
-
Unsupervised acoustic unit discovery for speech synthesis using discrete latent-variable neural networks 16 Apr 2019 · 0 repositories · arXiv:1904.07556
-
Unsupervised speech representation learning using WaveNet autoencoders 25 Jan 2019 · 5 repositories · arXiv:1901.08810
-
Variational Information Bottleneck on Vector Quantized Autoencoders 2 Aug 2018 · 0 repositories · arXiv:1808.01048
-
Learning Product Codebooks using Vector Quantized Autoencoders for Image Retrieval 12 Jul 2018 · 0 repositories · arXiv:1807.04629
-
Neural Discrete Representation Learning 2 Nov 2017 · 50 repositories · arXiv:1711.00937Syntology community repositories only · 64 ran (of which 35 constructed an object rather than computing a result; 60 with no instrument failure: 0 honoured, 0 violated, 60 with no contract checked; 4 where Syntology's instrument failed) · 16 unverified (of 80 harvested samples) · 32 pointer-only (licence)