{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/theory-and-experiments-on-vector-quantized","title":"Theory and Experiments on Vector Quantized Autoencoders","arxiv_id":"1805.11063","date":"2018-05-28","proceeding":null,"authors":["Aurko Roy","Ashish Vaswani","Arvind Neelakantan","Niki Parmar"],"abstract":"Deep neural networks with discrete latent variables offer the promise of\nbetter symbolic reasoning, and learning abstractions that are more useful to\nnew tasks. There has been a surge in interest in discrete latent variable\nmodels, however, despite several recent improvements, the training of discrete\nlatent variable models has remained challenging and their performance has\nmostly failed to match their continuous counterparts. Recent work on vector\nquantized autoencoders (VQ-VAE) has made substantial progress in this\ndirection, with its perplexity almost matching that of a VAE on datasets such\nas CIFAR-10. In this work, we investigate an alternate training technique for\nVQ-VAE, inspired by its connection to the Expectation Maximization (EM)\nalgorithm. Training the discrete bottleneck with EM helps us achieve better\nimage generation results on CIFAR-10, and together with knowledge distillation,\nallows us to develop a non-autoregressive machine translation model whose\naccuracy almost matches a strong greedy autoregressive baseline Transformer,\nwhile being 3.3 times faster at inference.","url_abs":"http://arxiv.org/abs/1805.11063v2","url_pdf":"http://arxiv.org/pdf/1805.11063v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"theory-and-experiments-on-vector-quantized","repo_url":"https://github.com/jaywalnut310/Vector-Quantized-Autoencoders","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok"}},{"paper_slug":"theory-and-experiments-on-vector-quantized","repo_url":"https://github.com/swasun/VQ-VAE-images","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"image-generation","task_name":"Image Generation"},{"task_slug":"knowledge-distillation","task_name":"Knowledge Distillation"},{"task_slug":"machine-translation","task_name":"Machine Translation"},{"task_slug":"translation","task_name":"Translation"}],"methods":[{"method_slug":"absolute-position-encodings","method_name":"Absolute Position Encodings"},{"method_slug":"adam","method_name":"Adam"},{"method_slug":"attention","method_name":"Attention"},{"method_slug":"bpe","method_name":"BPE"},{"method_slug":"dense-connections","method_name":"Dense Connections"},{"method_slug":"dropout","method_name":"Dropout"},{"method_slug":"label-smoothing","method_name":"Label Smoothing"},{"method_slug":"layer-normalization","method_name":"Layer Normalization"},{"method_slug":"linear-layer","method_name":"Linear Layer"},{"method_slug":"multi-head-attention","method_name":"Multi-Head Attention"},{"method_slug":"position-wise-feed-forward-layer","method_name":"Position-Wise Feed-Forward Layer"},{"method_slug":"relu","method_name":"ReLU"},{"method_slug":"residual-connection","method_name":"Residual Connection"},{"method_slug":"softmax","method_name":"Softmax"},{"method_slug":"transformer","method_name":"Transformer"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1805.11063","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}