{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/eco-efficient-convolutional-network-for","title":"ECO: Efficient Convolutional Network for Online Video Understanding","arxiv_id":"1804.09066","date":"2018-04-24","proceeding":"ECCV 2018 9","authors":["Mohammadreza Zolfaghari","Kamaljeet Singh","Thomas Brox"],"abstract":"The state of the art in video understanding suffers from two problems: (1)\nThe major part of reasoning is performed locally in the video, therefore, it\nmisses important relationships within actions that span several seconds. (2)\nWhile there are local methods with fast per-frame processing, the processing of\nthe whole video is not efficient and hampers fast video retrieval or online\nclassification of long-term activities. In this paper, we introduce a network\narchitecture that takes long-term content into account and enables fast\nper-video processing at the same time. The architecture is based on merging\nlong-term content already in the network rather than in a post-hoc fusion.\nTogether with a sampling strategy, which exploits that neighboring frames are\nlargely redundant, this yields high-quality action classification and video\ncaptioning at up to 230 videos per second, where each video can consist of a\nfew hundred frames. The approach achieves competitive performance across all\ndatasets while being 10x to 80x faster than state-of-the-art methods.","url_abs":"http://arxiv.org/abs/1804.09066v2","url_pdf":"http://arxiv.org/pdf/1804.09066v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"eco-efficient-convolutional-network-for","repo_url":"https://github.com/mzolfaghari/ECO-efficient-video-understanding","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"MIT"}},{"paper_slug":"eco-efficient-convolutional-network-for","repo_url":"https://github.com/MindSpore-paper-code-3/code8/tree/main/ecolite","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"mindspore","reach":null},{"paper_slug":"eco-efficient-convolutional-network-for","repo_url":"https://github.com/code-implementation1/Code1/tree/main/ecolite","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"mindspore","reach":null},{"paper_slug":"eco-efficient-convolutional-network-for","repo_url":"https://github.com/code-implementation1/Code2/tree/main/AdderNGD","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"mindspore","reach":null},{"paper_slug":"eco-efficient-convolutional-network-for","repo_url":"https://github.com/kingcong/ecolite","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"mindspore","reach":{"status":"gone","observed_at":"2026-09-18","how":"tree_404+repo_404"}},{"paper_slug":"eco-efficient-convolutional-network-for","repo_url":"https://github.com/mindspore-ai/models/tree/master/research/cv/ecolite","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"mindspore","reach":null}],"tasks":[{"task_slug":"action-classification","task_name":"Action Classification"},{"task_slug":"action-recognition-in-videos","task_name":"Action Recognition"},{"task_slug":"classification","task_name":"General Classification"},{"task_slug":"retrieval","task_name":"Retrieval"},{"task_slug":"video-captioning","task_name":"Video Captioning"},{"task_slug":"video-classification","task_name":"Video Classification"},{"task_slug":"video-retrieval","task_name":"Video Retrieval"},{"task_slug":"video-understanding","task_name":"Video Understanding"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/action-recognition-in-videos-on-something-1","task":"Action Recognition","dataset":"Something-Something V1","model":"ECO-Net (ImageNet pretrained)","rank_in_archive_order":66,"of":74,"metrics":{"Top 1 Accuracy":"46.4"},"uses_additional_data":true},{"leaderboard":"/sota/action-recognition-in-videos-on-something-1","task":"Action Recognition","dataset":"Something-Something V1","model":"ECO-Net","rank_in_archive_order":67,"of":74,"metrics":{"Top 1 Accuracy":"46.4"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1804.09066","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"1804.09066"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/code-implementation1/Code1/tree/main/ecolite","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/kingcong/ecolite","reach":{"status":"gone","observed_at":"2026-09-18","how":"tree_404+repo_404"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/mzolfaghari/ECO-efficient-video-understanding","reach":{"status":"ok","spdx":"MIT"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/MindSpore-paper-code-3/code8/tree/main/ecolite","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/mindspore-ai/models/tree/master/research/cv/ecolite","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/code-implementation1/Code2/tree/main/AdderNGD","reach":null}],"summary":{"unverified":2},"by_repo_kind":{"official":{"samples":2,"ran":0,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"d0250bc6fedca800","entry":"VideoSpatialPrediction","repo":"mzolfaghari/ECO-efficient-video-understanding","repo_kind":"official","path":"caffe_3d/action_python/VideoSpatialPrediction.py","file_url":"https://github.com/mzolfaghari/ECO-efficient-video-understanding/blob/HEAD/caffe_3d/action_python/VideoSpatialPrediction.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"d0250bc6fedca800"}},{"code_sha256_prefix":"4549877dcbf99fda","entry":"VideoTemporalPrediction","repo":"mzolfaghari/ECO-efficient-video-understanding","repo_kind":"official","path":"caffe_3d/action_python/VideoTemporalPrediction.py","file_url":"https://github.com/mzolfaghari/ECO-efficient-video-understanding/blob/HEAD/caffe_3d/action_python/VideoTemporalPrediction.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"4549877dcbf99fda"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}