{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/orthogonal-temporal-interpolation-for-zero","title":"Orthogonal Temporal Interpolation for Zero-Shot Video Recognition","arxiv_id":"2308.06897","date":"2023-08-14","proceeding":null,"authors":["Yan Zhu","Junbao Zhuo","Bin Ma","Jiajia Geng","Xiaoming Wei","Xiaolin Wei","Shuhui Wang"],"abstract":"Zero-shot video recognition (ZSVR) is a task that aims to recognize video categories that have not been seen during the model training process. Recently, vision-language models (VLMs) pre-trained on large-scale image-text pairs have demonstrated impressive transferability for ZSVR. To make VLMs applicable to the video domain, existing methods often use an additional temporal learning module after the image-level encoder to learn the temporal relationships among video frames. Unfortunately, for video from unseen categories, we observe an abnormal phenomenon where the model that uses spatial-temporal feature performs much worse than the model that removes temporal learning module and uses only spatial feature. We conjecture that improper temporal modeling on video disrupts the spatial feature of the video. To verify our hypothesis, we propose Feature Factorization to retain the orthogonal temporal feature of the video and use interpolation to construct refined spatial-temporal feature. The model using appropriately refined spatial-temporal feature performs better than the one using only spatial feature, which verifies the effectiveness of the orthogonal temporal feature for the ZSVR task. Therefore, an Orthogonal Temporal Interpolation module is designed to learn a better refined spatial-temporal video feature during training. Additionally, a Matching Loss is introduced to improve the quality of the orthogonal temporal feature. We propose a model called OTI for ZSVR by employing orthogonal temporal interpolation and the matching loss based on VLMs. The ZSVR accuracies on popular video datasets (i.e., Kinetics-600, UCF101 and HMDB51) show that OTI outperforms the previous state-of-the-art method by a clear margin.","url_abs":"https://arxiv.org/abs/2308.06897v1","url_pdf":"https://arxiv.org/pdf/2308.06897v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"orthogonal-temporal-interpolation-for-zero","repo_url":"https://github.com/sweetorangezhuyan/mm2023_oti","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"video-recognition","task_name":"Video Recognition"},{"task_slug":"zero-shot-action-recognition","task_name":"Zero-Shot Action Recognition"},{"task_slug":null,"task_name":"Zero-Shot Action Recognition on HMDB51"},{"task_slug":null,"task_name":"Zero-Shot Action Recognition on UCF101"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/zero-shot-action-recognition-on-hmdb51","task":"Zero-Shot Action Recognition","dataset":"HMDB51","model":"OTI(ViT-L/14)","rank_in_archive_order":2,"of":29,"metrics":{"Top-1 Accuracy":"64"},"uses_additional_data":false},{"leaderboard":"/sota/zero-shot-action-recognition-on-kinetics","task":"Zero-Shot Action Recognition","dataset":"Kinetics","model":"OTI（ViT-L/14）","rank_in_archive_order":5,"of":20,"metrics":{"Top-1 Accuracy":"70.6"},"uses_additional_data":false},{"leaderboard":"/sota/zero-shot-action-recognition-on-ucf101","task":"Zero-Shot Action Recognition","dataset":"UCF101","model":"OTI(ViT-L/14)","rank_in_archive_order":1,"of":35,"metrics":{"Top-1 Accuracy":"92.8"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2308.06897","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2308.06897"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/sweetorangezhuyan/mm2023_oti","reach":null}],"summary":{"ran_honours":2,"ran_draft_wrong":1},"by_repo_kind":{"official":{"samples":3,"ran":3,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":3,"samples":[{"code_sha256_prefix":"f550a530634eb2be","entry":"get_cswinmodel_pa","repo":"sweetorangezhuyan/mm2023_oti","repo_kind":"official","path":"train_oti.py","file_url":"https://github.com/sweetorangezhuyan/mm2023_oti/blob/HEAD/train_oti.py","link_basis":"first_harvest_node","language":"python","status":"ran_honours","verification_level":1,"contract_check":"HONOURS","metamorphic_tier":"well_formed","behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"f550a530634eb2be"}},{"code_sha256_prefix":"4c19dee3f3367478","entry":"get_orthogonal_num_vec","repo":"sweetorangezhuyan/mm2023_oti","repo_kind":"official","path":"modules/oti_zsvr.py","file_url":"https://github.com/sweetorangezhuyan/mm2023_oti/blob/HEAD/modules/oti_zsvr.py","link_basis":"first_harvest_node","language":"python","status":"ran_honours","verification_level":1,"contract_check":"HONOURS","metamorphic_tier":"invariant","behaviour_fingerprint":true,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"4c19dee3f3367478"}},{"code_sha256_prefix":"d452c540ad7ecbbc","entry":"update_dict","repo":"sweetorangezhuyan/mm2023_oti","repo_kind":"official","path":"train_oti.py","file_url":"https://github.com/sweetorangezhuyan/mm2023_oti/blob/HEAD/train_oti.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"OUTPUT_MISDECLARED","metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"d452c540ad7ecbbc"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}