{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/disentangling-motion-foreground-and","title":"Disentangling Motion, Foreground and Background Features in Videos","arxiv_id":"1707.04092","date":"2017-07-13","proceeding":null,"authors":["Xunyu Lin","Victor Campos","Xavier Giro-i-Nieto","Jordi Torres","Cristian Canton Ferrer"],"abstract":"This paper introduces an unsupervised framework to extract semantically rich\nfeatures for video representation. Inspired by how the human visual system\ngroups objects based on motion cues, we propose a deep convolutional neural\nnetwork that disentangles motion, foreground and background information. The\nproposed architecture consists of a 3D convolutional feature encoder for blocks\nof 16 frames, which is trained for reconstruction tasks over the first and last\nframes of the sequence. A preliminary supervised experiment was conducted to\nverify the feasibility of proposed method by training the model with a fraction\nof videos from the UCF-101 dataset taking as ground truth the bounding boxes\naround the activity regions. Qualitative results indicate that the network can\nsuccessfully segment foreground and background in videos as well as update the\nforeground appearance based on disentangled motion features. The benefits of\nthese learned features are shown in a discriminative classification task, where\ninitializing the network with the proposed pretraining method outperforms both\nrandom initialization and autoencoder pretraining. Our model and source code\nare publicly available at https://imatge-upc.github.io/unsupervised-2017-cvprw/ .","url_abs":"http://arxiv.org/abs/1707.04092v2","url_pdf":"http://arxiv.org/pdf/1707.04092v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"disentangling-motion-foreground-and","repo_url":"https://github.com/imatge-upc/unsupervised-2017-cvprw","is_official":1,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":null}],"tasks":[],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1707.04092","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}