{"url":"/method/movinet","slug":"movinet","name":"MoViNet","full_name":"MoViNet","full_name_withheld":false,"description_markdown":"**Mobile Video Network**, or **MoViNet**, is a type of computation and memory efficient video network that can operate on streaming video for online inference. Three techniques are used to improve efficiency while reducing the peak memory usage of 3D CNNs. First, a video network search space is designed and [neural architecture search](https://paperswithcode.com/method/neural-architecture-search) employed to generate efficient and diverse 3D CNN architectures. Second, a Stream Buffer technique is introduced that decouples memory from video clip duration, allowing 3D CNNs to embed arbitrary-length streaming video sequences for both training and inference with a small constant memory footprint. Third, a simple ensembling technique is used to improve accuracy further without sacrificing efficiency.","description_state":"present","introduced_year":null,"introduced_by":{"title":null,"paper":null,"first_author":null,"n_authors":0,"url_abs":null,"archive_paper_url":null},"source":{"url":"https://arxiv.org/abs/2103.11511v2","title":"MoViNets: Mobile Video Networks for Efficient Video Recognition","url_on_a_paper_host":true},"code_snippet_url":null,"code_snippet_url_on_a_code_host":false,"categories":[{"area":"Computer Vision","area_id":"computer-vision","collection":"Video Recognition Models","url":"/methods/category/video-recognition-models","pwc_aliases":[]}],"n_papers_tagged":2,"archive_num_papers":null,"papers_newest_first":[{"paper":"/paper/real-time-streaming-video-denoising-with","title":"Real-time Streaming Video Denoising with Bidirectional Buffers","date":"2022-07-14","arxiv_id":"2207.06937","n_code_links":1,"syntology":{"ran":3,"of":3,"unverified":0,"pointer_only":0}},{"paper":"/paper/movinets-mobile-video-networks-for-efficient","title":"MoViNets: Mobile Video Networks for Efficient Video Recognition","date":"2021-03-21","arxiv_id":"2103.11511","n_code_links":3,"syntology":{"ran":8,"of":13,"unverified":5,"pointer_only":0}}],"papers_shown":2,"tasks":[{"task":"/task/action-classification","name":"Action Classification","papers":1},{"task":"/task/action-recognition-in-videos","name":"Action Recognition","papers":1},{"task":"/task/computational-efficiency","name":"Computational Efficiency","papers":1},{"task":"/task/denoising","name":"Denoising","papers":1},{"task":"/task/architecture-search","name":"Neural Architecture Search","papers":1},{"task":"/task/action-recognition","name":"Temporal Action Localization","papers":1},{"task":"/task/video-denoising","name":"Video Denoising","papers":1},{"task":"/task/video-recognition","name":"Video Recognition","papers":1}],"tasks_shown":8,"n_tasks":8,"usage_by_year":[{"year":"2021","papers":1},{"year":"2022","papers":1}],"row_source":"embedded","archive":{"source":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","archive_url":"https://paperswithcode.com/method/movinet"},"syntology_read_at":"2026-09-24T18:15:14+00:00"}