{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/temporal-relational-reasoning-in-videos","title":"Temporal Relational Reasoning in Videos","arxiv_id":"1711.08496","date":"2017-11-22","proceeding":"ECCV 2018 9","authors":["Bolei Zhou","Alex Andonian","Aude Oliva","Antonio Torralba"],"abstract":"Temporal relational reasoning, the ability to link meaningful transformations\nof objects or entities over time, is a fundamental property of intelligent\nspecies. In this paper, we introduce an effective and interpretable network\nmodule, the Temporal Relation Network (TRN), designed to learn and reason about\ntemporal dependencies between video frames at multiple time scales. We evaluate\nTRN-equipped networks on activity recognition tasks using three recent video\ndatasets - Something-Something, Jester, and Charades - which fundamentally\ndepend on temporal relational reasoning. Our results demonstrate that the\nproposed TRN gives convolutional neural networks a remarkable capacity to\ndiscover temporal relations in videos. Through only sparsely sampled video\nframes, TRN-equipped networks can accurately predict human-object interactions\nin the Something-Something dataset and identify various human gestures on the\nJester dataset with very competitive performance. TRN-equipped networks also\noutperform two-stream networks and 3D convolution networks in recognizing daily\nactivities in the Charades dataset. Further analyses show that the models learn\nintuitive and interpretable visual common sense knowledge in videos.","url_abs":"http://arxiv.org/abs/1711.08496v2","url_pdf":"http://arxiv.org/pdf/1711.08496v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"temporal-relational-reasoning-in-videos","repo_url":"https://github.com/metalbubble/TRN-pytorch","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"unanswered"}},{"paper_slug":"temporal-relational-reasoning-in-videos","repo_url":"https://github.com/okankop/MFF-pytorch","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"unanswered"}},{"paper_slug":"temporal-relational-reasoning-in-videos","repo_url":"https://github.com/zhoubolei/TRN-pytorch","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null},{"paper_slug":"temporal-relational-reasoning-in-videos","repo_url":"https://github.com/2023-MindSpore-1/ms-code-7/tree/main/trn","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"mindspore","reach":null},{"paper_slug":"temporal-relational-reasoning-in-videos","repo_url":"https://github.com/MindSpore-paper-code-2/code3/tree/main/trn","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"mindspore","reach":null}],"tasks":[{"task_slug":"action-classification","task_name":"Action Classification"},{"task_slug":"action-recognition-in-videos","task_name":"Action Recognition"},{"task_slug":"action-recognition-in-videos-2","task_name":"Action Recognition In Videos"},{"task_slug":"activity-recognition","task_name":"Activity Recognition"},{"task_slug":"common-sense-reasoning","task_name":"Common Sense Reasoning"},{"task_slug":"human-object-interaction-detection","task_name":"Human-Object Interaction Detection"},{"task_slug":"relation-network","task_name":"Relation Network"},{"task_slug":"relational-reasoning","task_name":"Relational Reasoning"}],"methods":[{"method_slug":"3d-convolution","method_name":"3D Convolution"},{"method_slug":"convolution","method_name":"Convolution"}],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/action-classification-on-charades","task":"Action Classification","dataset":"Charades","model":"MultiScale TRN","rank_in_archive_order":42,"of":49,"metrics":{"MAP":"25.2"},"uses_additional_data":true},{"leaderboard":"/sota/action-classification-on-moments-in-time","task":"Action Classification","dataset":"MiT","model":"TRN-Multiscale","rank_in_archive_order":26,"of":29,"metrics":{"Top 1 Accuracy":"28.27","Top 5 Accuracy":"53.87"},"uses_additional_data":false},{"leaderboard":"/sota/action-recognition-in-videos-on-something-1","task":"Action Recognition","dataset":"Something-Something V1","model":"2-Stream TRN","rank_in_archive_order":71,"of":74,"metrics":{"Top 1 Accuracy":"42.01"},"uses_additional_data":false},{"leaderboard":"/sota/action-recognition-in-videos-on-something-1","task":"Action Recognition","dataset":"Something-Something V1","model":"M-TRN","rank_in_archive_order":74,"of":74,"metrics":{"Top 1 Accuracy":"34.4"},"uses_additional_data":false},{"leaderboard":"/sota/action-recognition-in-videos-on-jester-1","task":"Action Recognition In Videos","dataset":"Jester (Gesture Recognition)","model":"MultiScale TRN","rank_in_archive_order":4,"of":9,"metrics":{"Val":"95.31"},"uses_additional_data":false},{"leaderboard":"/sota/action-recognition-in-videos-on-something-2","task":"Action Recognition In Videos","dataset":"Something-Something V1","model":"2-Stream TRN","rank_in_archive_order":3,"of":3,"metrics":{"Top 1 Accuracy":"42.01"},"uses_additional_data":false},{"leaderboard":"/sota/action-recognition-in-videos-on-something-3","task":"Action Recognition In Videos","dataset":"Something-Something V2","model":"2-Stream TRN","rank_in_archive_order":3,"of":4,"metrics":{"Top-1 Accuracy":"55.52","Top-5 Accuracy":"83.06"},"uses_additional_data":false},{"leaderboard":"/sota/hand-gesture-recognition-on-jester-test","task":"Hand Gesture Recognition","dataset":"Jester test","model":"Multiscale TRN","rank_in_archive_order":2,"of":2,"metrics":{"Top 1 Accuracy":"94.78"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1711.08496","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"1711.08496"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/MindSpore-paper-code-2/code3/tree/main/trn","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/zhoubolei/TRN-pytorch","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/okankop/MFF-pytorch","reach":{"status":"unanswered"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/2023-MindSpore-1/ms-code-7/tree/main/trn","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/metalbubble/TRN-pytorch","reach":{"status":"unanswered"}}],"summary":{"ran_honours":1,"ran_draft_wrong":1,"unverified":1},"by_repo_kind":{"listed":{"samples":3,"ran":2,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":3,"samples":[{"code_sha256_prefix":"8949b39f8daef7b5","entry":"load_frames","repo":"zhoubolei/TRN-pytorch","repo_kind":"listed","path":"test_video.py","file_url":"https://github.com/zhoubolei/TRN-pytorch/blob/HEAD/test_video.py","link_basis":"first_harvest_node","language":"python","status":"ran_honours","verification_level":1,"contract_check":"HONOURS","metamorphic_tier":"well_formed","behaviour_fingerprint":false,"licence":"NOASSERTION","inline_ok":false,"mcp_get_code":{"code_sha256":"8949b39f8daef7b5"}},{"code_sha256_prefix":"9d6cb70a03f4ceb4","entry":"return_TRN","repo":"zhoubolei/TRN-pytorch","repo_kind":"listed","path":"TRNmodule.py","file_url":"https://github.com/zhoubolei/TRN-pytorch/blob/HEAD/TRNmodule.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"OUTPUT_MISDECLARED","metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"NOASSERTION","inline_ok":false,"mcp_get_code":{"code_sha256":"9d6cb70a03f4ceb4"}},{"code_sha256_prefix":"c89e8de7543d2de8","entry":"extract_frames","repo":"zhoubolei/TRN-pytorch","repo_kind":"listed","path":"test_video.py","file_url":"https://github.com/zhoubolei/TRN-pytorch/blob/HEAD/test_video.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NOASSERTION","inline_ok":false,"mcp_get_code":{"code_sha256":"c89e8de7543d2de8"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}