{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/few-shot-video-classification-via-temporal","title":"Few-Shot Video Classification via Temporal Alignment","arxiv_id":"1906.11415","date":"2019-06-27","proceeding":"CVPR 2020 6","authors":["Kaidi Cao","Jingwei Ji","Zhangjie Cao","Chien-Yi Chang","Juan Carlos Niebles"],"abstract":"There is a growing interest in learning a model which could recognize novel classes with only a few labeled examples. In this paper, we propose Temporal Alignment Module (TAM), a novel few-shot learning framework that can learn to classify a previous unseen video. While most previous works neglect long-term temporal ordering information, our proposed model explicitly leverages the temporal ordering information in video data through temporal alignment. This leads to strong data-efficiency for few-shot learning. In concrete, TAM calculates the distance value of query video with respect to novel class proxies by averaging the per frame distances along its alignment path. We introduce continuous relaxation to TAM so the model can be learned in an end-to-end fashion to directly optimize the few-shot learning objective. We evaluate TAM on two challenging real-world datasets, Kinetics and Something-Something-V2, and show that our model leads to significant improvement of few-shot video classification over a wide range of competitive baselines.","url_abs":"https://arxiv.org/abs/1906.11415v1","url_pdf":"https://arxiv.org/pdf/1906.11415v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"action-recognition-in-videos","task_name":"Action Recognition"},{"task_slug":"classification-1","task_name":"Classification"},{"task_slug":"few-shot-action-recognition","task_name":"Few Shot Action Recognition"},{"task_slug":"few-shot-learning","task_name":"Few-Shot Learning"},{"task_slug":"classification","task_name":"General Classification"},{"task_slug":"video-classification","task_name":"Video Classification"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/action-recognition-in-videos-on-something","task":"Action Recognition","dataset":"Something-Something V2","model":"TAM (5-shot)","rank_in_archive_order":115,"of":123,"metrics":{"Top-1 Accuracy":"52.3"},"uses_additional_data":false},{"leaderboard":"/sota/few-shot-action-recognition-on-kinetics-100","task":"Few Shot Action Recognition","dataset":"Kinetics-100","model":"OTAM","rank_in_archive_order":6,"of":8,"metrics":{"Accuracy":"85.8"},"uses_additional_data":false},{"leaderboard":"/sota/few-shot-action-recognition-on-something","task":"Few Shot Action Recognition","dataset":"Something-Something-100","model":"OTAM","rank_in_archive_order":5,"of":5,"metrics":{"1:1 Accuracy":"52.3"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1906.11415","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}