{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/spatio-temporal-video-autoencoder-with","title":"Spatio-temporal video autoencoder with differentiable memory","arxiv_id":"1511.06309","date":"2015-11-19","proceeding":null,"authors":["Viorica Patraucean","Ankur Handa","Roberto Cipolla"],"abstract":"We describe a new spatio-temporal video autoencoder, based on a classic\nspatial image autoencoder and a novel nested temporal autoencoder. The temporal\nencoder is represented by a differentiable visual memory composed of\nconvolutional long short-term memory (LSTM) cells that integrate changes over\ntime. Here we target motion changes and use as temporal decoder a robust\noptical flow prediction module together with an image sampler serving as\nbuilt-in feedback loop. The architecture is end-to-end differentiable. At each\ntime step, the system receives as input a video frame, predicts the optical\nflow based on the current observation and the LSTM memory state as a dense\ntransformation map, and applies it to the current frame to generate the next\nframe. By minimising the reconstruction error between the predicted next frame\nand the corresponding ground truth next frame, we train the whole system to\nextract features useful for motion estimation without any supervision effort.\nWe present one direct application of the proposed framework in\nweakly-supervised semantic segmentation of videos through label propagation\nusing optical flow.","url_abs":"http://arxiv.org/abs/1511.06309v5","url_pdf":"http://arxiv.org/pdf/1511.06309v5.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"spatio-temporal-video-autoencoder-with","repo_url":"https://github.com/viorik/ConvLSTM","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok"}}],"tasks":[{"task_slug":"decoder","task_name":"Decoder"},{"task_slug":"motion-estimation","task_name":"Motion Estimation"},{"task_slug":"optical-flow-estimation","task_name":"Optical Flow Estimation"},{"task_slug":"semantic-segmentation","task_name":"Semantic Segmentation"},{"task_slug":"weakly-supervised-semantic-segmentation-1","task_name":"Weakly supervised Semantic Segmentation"},{"task_slug":"weakly-supervised-semantic-segmentation","task_name":"Weakly-Supervised Semantic Segmentation"}],"methods":[{"method_slug":"lstm","method_name":"LSTM"},{"method_slug":"sigmoid-activation","method_name":"Sigmoid Activation"},{"method_slug":"tanh-activation","method_name":"Tanh Activation"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":"https://app.syntology.ai/?focus=1511.06309","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}