{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/ts-lstm-and-temporal-inception-exploiting","title":"TS-LSTM and Temporal-Inception: Exploiting Spatiotemporal Dynamics for Activity Recognition","arxiv_id":"1703.10667","date":"2017-03-30","proceeding":null,"authors":["Chih-Yao Ma","Min-Hung Chen","Zsolt Kira","Ghassan AlRegib"],"abstract":"Recent two-stream deep Convolutional Neural Networks (ConvNets) have made\nsignificant progress in recognizing human actions in videos. Despite their\nsuccess, methods extending the basic two-stream ConvNet have not systematically\nexplored possible network architectures to further exploit spatiotemporal\ndynamics within video sequences. Further, such networks often use different\nbaseline two-stream networks. Therefore, the differences and the distinguishing\nfactors between various methods using Recurrent Neural Networks (RNN) or\nconvolutional networks on temporally-constructed feature vectors\n(Temporal-ConvNet) are unclear. In this work, we first demonstrate a strong\nbaseline two-stream ConvNet using ResNet-101. We use this baseline to\nthoroughly examine the use of both RNNs and Temporal-ConvNets for extracting\nspatiotemporal information. Building upon our experimental results, we then\npropose and investigate two different networks to further integrate\nspatiotemporal information: 1) temporal segment RNN and 2) Inception-style\nTemporal-ConvNet. We demonstrate that using both RNNs (using LSTMs) and\nTemporal-ConvNets on spatiotemporal feature matrices are able to exploit\nspatiotemporal dynamics to improve the overall performance. However, each of\nthese methods require proper care to achieve state-of-the-art performance; for\nexample, LSTMs require pre-segmented data or else they cannot fully exploit\ntemporal information. Our analysis identifies specific limitations for each\nmethod that could form the basis of future work. Our experimental results on\nUCF101 and HMDB51 datasets achieve state-of-the-art performances, 94.1% and\n69.0%, respectively, without requiring extensive temporal augmentation.","url_abs":"http://arxiv.org/abs/1703.10667v1","url_pdf":"http://arxiv.org/pdf/1703.10667v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"ts-lstm-and-temporal-inception-exploiting","repo_url":"https://github.com/chihyaoma/Activity-Recognition-with-CNN-and-RNN","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"torch","reach":{"status":"unanswered"}},{"paper_slug":"ts-lstm-and-temporal-inception-exploiting","repo_url":"https://github.com/ChetanTayal138/HAR-Web","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"unanswered"}},{"paper_slug":"ts-lstm-and-temporal-inception-exploiting","repo_url":"https://github.com/jeffreyhuang1/two-stream-action-recognition","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"unanswered"}},{"paper_slug":"ts-lstm-and-temporal-inception-exploiting","repo_url":"https://github.com/jeffreyyihuang/two-stream-action-recognition","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"unanswered"}}],"tasks":[{"task_slug":"action-classification","task_name":"Action Classification"},{"task_slug":"action-recognition-in-videos","task_name":"Action Recognition"},{"task_slug":"activity-recognition","task_name":"Activity Recognition"},{"task_slug":"action-recognition","task_name":"Temporal Action Localization"},{"task_slug":"video-classification","task_name":"Video Classification"},{"task_slug":"video-understanding","task_name":"Video Understanding"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/action-recognition-in-videos-on-hmdb-51","task":"Action Recognition","dataset":"HMDB-51","model":"TS-LSTM","rank_in_archive_order":57,"of":77,"metrics":{"Average accuracy of 3 splits":"69"},"uses_additional_data":false},{"leaderboard":"/sota/action-recognition-in-videos-on-ucf101","task":"Action Recognition","dataset":"UCF101","model":"TS-LSTM","rank_in_archive_order":57,"of":91,"metrics":{"3-fold Accuracy":"94.1"},"uses_additional_data":false}],"syntology":{"syntology_url":"https://syntology.ai/paper/1703.10667","atlas_url":"https://app.syntology.ai/?focus=1703.10667","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}