{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/b-multivariational-autoencoder-for-entangled","title":"$β$-Multivariational Autoencoder for Entangled Representation Learning in Video Frames","arxiv_id":"2211.12627","date":"2022-11-22","proceeding":null,"authors":["Fatemeh Nouri","Robert Bergevin"],"abstract":"It is crucial to choose actions from an appropriate distribution while learning a sequential decision-making process in which a set of actions is expected given the states and previous reward. Yet, if there are more than two latent variables and every two variables have a covariance value, learning a known prior from data becomes challenging. Because when the data are big and diverse, many posterior estimate methods experience posterior collapse. In this paper, we propose the $\\beta$-Multivariational Autoencoder ($\\beta$MVAE) to learn a Multivariate Gaussian prior from video frames for use as part of a single object-tracking in form of a decision-making process. We present a novel formulation for object motion in videos with a set of dependent parameters to address a single object-tracking task. The true values of the motion parameters are obtained through data analysis on the training set. The parameters population is then assumed to have a Multivariate Gaussian distribution. The $\\beta$MVAE is developed to learn this entangled prior $p = N(\\mu, \\Sigma)$ directly from frame patches where the output is the object masks of the frame patches. We devise a bottleneck to estimate the posterior's parameters, i.e. $\\mu', \\Sigma'$. Via a new reparameterization trick, we learn the likelihood $p(\\hat{x}|z)$ as the object mask of the input. Furthermore, we alter the neural network of $\\beta$MVAE with the U-Net architecture and name the new network $\\beta$Multivariational U-Net ($\\beta$MVUnet). Our networks are trained from scratch via over 85k video frames for 24 ($\\beta$MVUnet) and 78 ($\\beta$MVAE) million steps. We show that $\\beta$MVUnet enhances both posterior estimation and segmentation functioning over the test set. Our code and the trained networks are publicly released.","url_abs":"https://arxiv.org/abs/2211.12627v1","url_pdf":"https://arxiv.org/pdf/2211.12627v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"b-multivariational-autoencoder-for-entangled","repo_url":"https://github.com/fatemehN/entangled_representation","is_official":1,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"decision-making","task_name":"Decision Making"},{"task_slug":"object","task_name":"Object"},{"task_slug":"object-tracking","task_name":"Object Tracking"},{"task_slug":"representation-learning","task_name":"Representation Learning"},{"task_slug":"sequential-decision-making","task_name":"Sequential Decision Making"}],"methods":[{"method_slug":"concatenated-skip-connection","method_name":"Concatenated Skip Connection"},{"method_slug":"convolution","method_name":"Convolution"},{"method_slug":"max-pooling","method_name":"Max Pooling"},{"method_slug":"relu","method_name":"ReLU"},{"method_slug":"test","method_name":"Test"},{"method_slug":"u-net","method_name":"U-Net"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}