{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/action-recognition-with-trajectory-pooled","title":"Action Recognition with Trajectory-Pooled Deep-Convolutional Descriptors","arxiv_id":"1505.04868","date":"2015-05-19","proceeding":"CVPR 2015 6","authors":["Limin Wang","Yu Qiao","Xiaoou Tang"],"abstract":"Visual features are of vital importance for human action understanding in\nvideos. This paper presents a new video representation, called\ntrajectory-pooled deep-convolutional descriptor (TDD), which shares the merits\nof both hand-crafted features and deep-learned features. Specifically, we\nutilize deep architectures to learn discriminative convolutional feature maps,\nand conduct trajectory-constrained pooling to aggregate these convolutional\nfeatures into effective descriptors. To enhance the robustness of TDDs, we\ndesign two normalization methods to transform convolutional feature maps,\nnamely spatiotemporal normalization and channel normalization. The advantages\nof our features come from (i) TDDs are automatically learned and contain high\ndiscriminative capacity compared with those hand-crafted features; (ii) TDDs\ntake account of the intrinsic characteristics of temporal dimension and\nintroduce the strategies of trajectory-constrained sampling and pooling for\naggregating deep-learned features. We conduct experiments on two challenging\ndatasets: HMDB51 and UCF101. Experimental results show that TDDs outperform\nprevious hand-crafted features and deep-learned features. Our method also\nachieves superior performance to the state of the art on these datasets (HMDB51\n65.9%, UCF101 91.5%).","url_abs":"http://arxiv.org/abs/1505.04868v1","url_pdf":"http://arxiv.org/pdf/1505.04868v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"action-recognition-with-trajectory-pooled","repo_url":"https://github.com/damien911224/theWorldInSafety","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok","spdx":"GPL-3.0"}}],"tasks":[{"task_slug":"action-recognition-in-videos","task_name":"Action Recognition"},{"task_slug":"action-understanding","task_name":"Action Understanding"},{"task_slug":"activity-recognition-in-videos","task_name":"Activity Recognition In Videos"},{"task_slug":"action-recognition","task_name":"Temporal Action Localization"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/action-recognition-in-videos-on-hmdb-51","task":"Action Recognition","dataset":"HMDB-51","model":"TDD + IDT","rank_in_archive_order":60,"of":77,"metrics":{"Average accuracy of 3 splits":"65.9"},"uses_additional_data":false},{"leaderboard":"/sota/action-recognition-in-videos-on-ucf101","task":"Action Recognition","dataset":"UCF101","model":"TDD + IDT","rank_in_archive_order":68,"of":91,"metrics":{"3-fold Accuracy":"91.5"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1505.04868","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}