{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/multi-fiber-networks-for-video-recognition","title":"Multi-Fiber Networks for Video Recognition","arxiv_id":"1807.11195","date":"2018-07-30","proceeding":"ECCV 2018 9","authors":["Yunpeng Chen","Yannis Kalantidis","Jianshu Li","Shuicheng Yan","Jiashi Feng"],"abstract":"In this paper, we aim to reduce the computational cost of spatio-temporal\ndeep neural networks, making them run as fast as their 2D counterparts while\npreserving state-of-the-art accuracy on video recognition benchmarks. To this\nend, we present the novel Multi-Fiber architecture that slices a complex neural\nnetwork into an ensemble of lightweight networks or fibers that run through the\nnetwork. To facilitate information flow between fibers we further incorporate\nmultiplexer modules and end up with an architecture that reduces the\ncomputational cost of 3D networks by an order of magnitude, while increasing\nrecognition performance at the same time. Extensive experimental results show\nthat our multi-fiber architecture significantly boosts the efficiency of\nexisting convolution networks for both image and video recognition tasks,\nachieving state-of-the-art performance on UCF-101, HMDB-51 and Kinetics\ndatasets. Our proposed model requires over 9x and 13x less computations than\nthe I3D and R(2+1)D models, respectively, yet providing higher accuracy.","url_abs":"http://arxiv.org/abs/1807.11195v3","url_pdf":"http://arxiv.org/pdf/1807.11195v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"action-classification","task_name":"Action Classification"},{"task_slug":"action-recognition-in-videos","task_name":"Action Recognition"},{"task_slug":"video-recognition","task_name":"Video Recognition"}],"methods":[{"method_slug":"2-1-d-convolution","method_name":"(2+1)D Convolution"},{"method_slug":"average-pooling","method_name":"Average Pooling"},{"method_slug":"batch-normalization","method_name":"Batch Normalization"},{"method_slug":"convolution","method_name":"Convolution"},{"method_slug":"dense-connections","method_name":"Dense Connections"},{"method_slug":"global-average-pooling","method_name":"Global Average Pooling"},{"method_slug":"r-2-1-d","method_name":"R(2+1)D"},{"method_slug":"relu","method_name":"ReLU"},{"method_slug":"residual-connection","method_name":"Residual Connection"}],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/action-classification-on-kinetics-400","task":"Action Classification","dataset":"Kinetics-400","model":"MFNet","rank_in_archive_order":171,"of":207,"metrics":{"Acc@1":"72.8","Acc@5":"90.4"},"uses_additional_data":false},{"leaderboard":"/sota/action-recognition-in-videos-on-ucf101","task":"Action Recognition","dataset":"UCF101","model":"MF-Net, RGB only (ImageNet+Kinetics pretrained)","rank_in_archive_order":39,"of":91,"metrics":{"3-fold Accuracy":"96.0"},"uses_additional_data":true}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1807.11195","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}