{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/improved-dense-trajectory-with-cross-streams","title":"Improved Dense Trajectory with Cross Streams","arxiv_id":"1604.08826","date":"2016-04-29","proceeding":null,"authors":["Katsunori Ohnishi","Masatoshi Hidaka","Tatsuya Harada"],"abstract":"Improved dense trajectories (iDT) have shown great performance in action\nrecognition, and their combination with the two-stream approach has achieved\nstate-of-the-art performance. It is, however, difficult for iDT to completely\nremove background trajectories from video with camera shaking. Trajectories in\nless discriminative regions should be given modest weights in order to create\nmore discriminative local descriptors for action recognition. In addition, the\ntwo-stream approach, which learns appearance and motion information separately,\ncannot focus on motion in important regions when extracting features from\nspatial convolutional layers of the appearance network, and vice versa. In\norder to address the above mentioned problems, we propose a new local\ndescriptor that pools a new convolutional layer obtained from crossing two\nnetworks along iDT. This new descriptor is calculated by applying\ndiscriminative weights learned from one network to a convolutional layer of the\nother network. Our method has achieved state-of-the-art performance on ordinal\naction recognition datasets, 92.3% on UCF101, and 66.2% on HMDB51.","url_abs":"http://arxiv.org/abs/1604.08826v1","url_pdf":"http://arxiv.org/pdf/1604.08826v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"action-classification","task_name":"Action Classification"},{"task_slug":"action-recognition-in-videos","task_name":"Action Recognition"},{"task_slug":"action-recognition","task_name":"Temporal Action Localization"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/action-classification-on-toyota-smarthome","task":"Action Classification","dataset":"Toyota Smarthome dataset","model":"Dense Trajectories","rank_in_archive_order":11,"of":13,"metrics":{"CS":"41.9","CV1":"20.9","CV2":"23.7"},"uses_additional_data":false}],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}