{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/deep-temporal-linear-encoding-networks","title":"Deep Temporal Linear Encoding Networks","arxiv_id":"1611.06678","date":"2016-11-21","proceeding":"CVPR 2017 7","authors":["Ali Diba","Vivek Sharma","Luc van Gool"],"abstract":"The CNN-encoding of features from entire videos for the representation of\nhuman actions has rarely been addressed. Instead, CNN work has focused on\napproaches to fuse spatial and temporal networks, but these were typically\nlimited to processing shorter sequences. We present a new video representation,\ncalled temporal linear encoding (TLE) and embedded inside of CNNs as a new\nlayer, which captures the appearance and motion throughout entire videos. It\nencodes this aggregated information into a robust video feature representation,\nvia end-to-end learning. Advantages of TLEs are: (a) they encode the entire\nvideo into a compact feature representation, learning the semantics and a\ndiscriminative feature space; (b) they are applicable to all kinds of networks\nlike 2D and 3D CNNs for video classification; and (c) they model feature\ninteractions in a more expressive way and without loss of information. We\nconduct experiments on two challenging human action datasets: HMDB51 and\nUCF101. The experiments show that TLE outperforms current state-of-the-art\nmethods on both datasets.","url_abs":"http://arxiv.org/abs/1611.06678v1","url_pdf":"http://arxiv.org/pdf/1611.06678v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"deep-temporal-linear-encoding-networks","repo_url":"https://github.com/AbdalaDiasse/Video-classification-for-oil-quality-estimation","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"unanswered"}},{"paper_slug":"deep-temporal-linear-encoding-networks","repo_url":"https://github.com/bryanyzhu/two-stream-pytorch","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"unanswered"}}],"tasks":[{"task_slug":"representation-learning","task_name":"Representation Learning"},{"task_slug":"video-classification","task_name":"Video Classification"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1611.06678","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}