{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/tdn-temporal-difference-networks-for","title":"TDN: Temporal Difference Networks for Efficient Action Recognition","arxiv_id":"2012.10071","date":"2020-12-18","proceeding":"CVPR 2021 1","authors":["LiMin Wang","Zhan Tong","Bin Ji","Gangshan Wu"],"abstract":"Temporal modeling still remains challenging for action recognition in videos. To mitigate this issue, this paper presents a new video architecture, termed as Temporal Difference Network (TDN), with a focus on capturing multi-scale temporal information for efficient action recognition. The core of our TDN is to devise an efficient temporal module (TDM) by explicitly leveraging a temporal difference operator, and systematically assess its effect on short-term and long-term motion modeling. To fully capture temporal information over the entire video, our TDN is established with a two-level difference modeling paradigm. Specifically, for local motion modeling, temporal difference over consecutive frames is used to supply 2D CNNs with finer motion pattern, while for global motion modeling, temporal difference across segments is incorporated to capture long-range structure for motion feature excitation. TDN provides a simple and principled temporal modeling framework and could be instantiated with the existing CNNs at a small extra computational cost. Our TDN presents a new state of the art on the Something-Something V1 & V2 datasets and is on par with the best performance on the Kinetics-400 dataset. In addition, we conduct in-depth ablation studies and plot the visualization results of our TDN, hopefully providing insightful analysis on temporal difference modeling. We release the code at https://github.com/MCG-NJU/TDN.","url_abs":"https://arxiv.org/abs/2012.10071v2","url_pdf":"https://arxiv.org/pdf/2012.10071v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"tdn-temporal-difference-networks-for","repo_url":"https://github.com/MCG-NJU/TDN","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"action-classification","task_name":"Action Classification"},{"task_slug":"action-recognition-in-videos","task_name":"Action Recognition"},{"task_slug":"action-recognition-in-videos-2","task_name":"Action Recognition In Videos"}],"methods":[{"method_slug":"tdn","method_name":"TDN"}],"datasets_introduced":[],"methods_introduced":[{"slug":"tdn","name":"TDN","full_name":"Temporaral Difference Network"}],"results":[{"leaderboard":"/sota/action-classification-on-kinetics-400","task":"Action Classification","dataset":"Kinetics-400","model":"TDN-ResNet101 (ensemble, ImageNet pretrained, RGB only)","rank_in_archive_order":111,"of":207,"metrics":{"Acc@1":"79.4","Acc@5":"94.4"},"uses_additional_data":false},{"leaderboard":"/sota/action-recognition-in-videos-on-something-1","task":"Action Recognition","dataset":"Something-Something V1","model":"TDN ResNet101 (one clip, center crop, 8+16 ensemble, ImageNet pretrained, RGB only)","rank_in_archive_order":18,"of":74,"metrics":{"Top 1 Accuracy":"56.8","Top 5 Accuracy":"84.1"},"uses_additional_data":false},{"leaderboard":"/sota/action-recognition-in-videos-on-something","task":"Action Recognition","dataset":"Something-Something V2","model":"TDN ResNet101 (one clip, three crop, 8+16 ensemble, ImageNet pretrained, RGB only)","rank_in_archive_order":44,"of":123,"metrics":{"GFLOPs":"198x3","Top-1 Accuracy":"69.6","Top-5 Accuracy":"92.2"},"uses_additional_data":true},{"leaderboard":"/sota/action-recognition-in-videos-on-something","task":"Action Recognition","dataset":"Something-Something V2","model":"TDN ResNet101 (one clip, center crop, 8+16 ensemble, ImageNet pretrained, RGB only)","rank_in_archive_order":52,"of":123,"metrics":{"GFLOPs":"198x1","Top-1 Accuracy":"68.2","Top-5 Accuracy":"91.6"},"uses_additional_data":true}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2012.10071","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2012.10071"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/MCG-NJU/TDN","reach":null}],"summary":{"ran":1,"ran_fixture":1,"unverified":2},"by_repo_kind":{"official":{"samples":4,"ran":2,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"0492518f68c77840","entry":"TDN_Net","repo":"MCG-NJU/TDN","repo_kind":"official","path":"ops/tdn_net.py","file_url":"https://github.com/MCG-NJU/TDN/blob/HEAD/ops/tdn_net.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"0492518f68c77840"}},{"code_sha256_prefix":"8843a54c072af40d","entry":"accuracy","repo":"MCG-NJU/TDN","repo_kind":"official","path":"test_models_center_crop.py","file_url":"https://github.com/MCG-NJU/TDN/blob/HEAD/test_models_center_crop.py","link_basis":"first_harvest_node","language":"python","status":"ran_fixture","verification_level":1,"contract_check":"RAISES","metamorphic_tier":"invariant","behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"8843a54c072af40d"}},{"code_sha256_prefix":"753b07af0ef51461","entry":"eval_video","repo":"MCG-NJU/TDN","repo_kind":"official","path":"test_models_center_crop.py","file_url":"https://github.com/MCG-NJU/TDN/blob/HEAD/test_models_center_crop.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"753b07af0ef51461"}},{"code_sha256_prefix":"97b9705aaede3cbc","entry":"tdn_net","repo":"MCG-NJU/TDN","repo_kind":"official","path":"ops/tdn_net.py","file_url":"https://github.com/MCG-NJU/TDN/blob/HEAD/ops/tdn_net.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"97b9705aaede3cbc"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}