{"url":"/task/action-localization","name":"Action Localization","slug":"action-localization","description_markdown":"Action Localization is finding the spatial and temporal co ordinates for an action in a video. An action localization model will identify which frame an action start and ends in video and return the x,y coordinates of an action. Further the co ordinates will change  when the object performing action  undergoes a displacement.","categories":[{"name":"Computer Vision","url":"/area/computer-vision"}],"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","slug_source":"archive_url"},"counts":{"papers_tagged":369,"papers_with_code":169,"benchmarks":0,"benchmark_tables_in_archive":0,"benchmark_tables_shown":0,"benchmark_tables_withheld_as_spam":0,"benchmark_definition":"a leaderboard table with at least one row; benchmark_tables_shown also counts the zero-row tables; benchmark_tables_in_archive adds the tables withheld as spam","datasets":4,"subtasks":4,"parent_tasks":0},"benchmarks":[],"datasets":[{"url":"/dataset/coin","name":"COIN","full_name":"","num_papers_in_archive":105},{"url":"/dataset/hacs","name":"HACS","full_name":"Human Action Clips and Segments","num_papers_in_archive":75},{"url":"/dataset/grasp","name":"GraSP","full_name":"Holistic and Multi-Granular Surgical Scene Understanding of Prostatectomies","num_papers_in_archive":3},{"url":"/dataset/cvb-a-video-dataset-of-cattle-visual","name":"CVB","full_name":"Video Dataset of Cattle Visual Behaviors","num_papers_in_archive":2}],"subtasks":[{"url":"/task/action-recognition","name":"Temporal Action Localization"},{"url":"/task/action-segmentation","name":"Action Segmentation"},{"url":"/task/spatio-temporal-action-localization","name":"Spatio-Temporal Action Localization"},{"url":"/task/unusual-activity-localization","name":"Unusual Activity Localization"}],"parent_tasks":[],"papers":{"order":"repositories listed in the archive (desc), then date (desc); the archive holds no stars","population":"papers tagged with this task that list at least one repository in the archive","shown":30,"of":169,"tagged_in_all":369,"items":[{"url":"/paper/ava-a-video-dataset-of-spatio-temporally","title":"AVA: A Video Dataset of Spatio-temporally Localized Atomic Visual Actions","date":"2017-05-23","arxiv_id":"1705.08421","repositories_listed":9,"syntology":null},{"url":"/paper/acgnet-action-complement-graph-network-for","title":"ACGNet: Action Complement Graph Network for Weakly-supervised Temporal Action Localization","date":"2021-12-21","arxiv_id":"2112.10977","repositories_listed":5,"syntology":null},{"url":"/paper/you-only-watch-once-a-unified-cnn","title":"You Only Watch Once: A Unified CNN Architecture for Real-Time Spatiotemporal Action Localization","date":"2019-11-15","arxiv_id":"1911.06644","repositories_listed":5,"syntology":{"n":12,"n_ran":3,"n_unverified":9,"n_pointer_only":2}},{"url":"/paper/recognition-of-instrument-tissue-interactions","title":"Recognition of Instrument-Tissue Interactions in Endoscopic Videos via Action Triplets","date":"2020-07-10","arxiv_id":"2007.05405","repositories_listed":4,"syntology":null},{"url":"/paper/end-to-end-learning-of-visual-representations","title":"End-to-End Learning of Visual Representations from Uncurated Instructional Videos","date":"2019-12-13","arxiv_id":"1912.06430","repositories_listed":4,"syntology":{"n":5,"n_ran":3,"n_unverified":2,"n_pointer_only":0}},{"url":"/paper/howto100m-learning-a-text-video-embedding-by","title":"HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video Clips","date":"2019-06-07","arxiv_id":"1906.03327","repositories_listed":4,"syntology":{"n":1,"n_ran":1,"n_unverified":0,"n_pointer_only":1}},{"url":"/paper/temporal-action-localization-with-enhanced","title":"Temporal Action Localization with Enhanced Instant Discriminability","date":"2023-09-11","arxiv_id":"2309.05590","repositories_listed":3,"syntology":null},{"url":"/paper/videomix-rethinking-data-augmentation-for","title":"VideoMix: Rethinking Data Augmentation for Video Classification","date":"2020-12-07","arxiv_id":"2012.03457","repositories_listed":3,"syntology":{"n":3,"n_ran":2,"n_unverified":1,"n_pointer_only":0}},{"url":"/paper/1st-place-solution-for-ava-kinetics-crossover","title":"1st place solution for AVA-Kinetics Crossover in AcitivityNet Challenge 2020","date":"2020-06-16","arxiv_id":"2006.09116","repositories_listed":3,"syntology":null},{"url":"/paper/actor-context-actor-relation-network-for","title":"Actor-Context-Actor Relation Network for Spatio-Temporal Action Localization","date":"2020-06-14","arxiv_id":"2006.07976","repositories_listed":3,"syntology":{"n":27,"n_ran":1,"n_unverified":26,"n_pointer_only":0}},{"url":"/paper/weakly-supervised-action-localization-by","title":"Weakly Supervised Action Localization by Sparse Temporal Pooling Network","date":"2017-12-14","arxiv_id":"1712.05080","repositories_listed":3,"syntology":null},{"url":"/paper/hide-and-seek-forcing-a-network-to-be","title":"Hide-and-Seek: Forcing a Network to be Meticulous for Weakly-supervised Object and Action Localization","date":"2017-04-13","arxiv_id":"1704.04232","repositories_listed":3,"syntology":null},{"url":"/paper/gpt-4v-in-wonderland-large-multimodal-models","title":"GPT-4V in Wonderland: Large Multimodal Models for Zero-Shot Smartphone GUI Navigation","date":"2023-11-13","arxiv_id":"2311.07562","repositories_listed":2,"syntology":{"n":3,"n_ran":2,"n_unverified":1,"n_pointer_only":3}},{"url":"/paper/fha-kitchens-a-novel-dataset-for-fine-grained","title":"Multi-Granularity Hand Action Detection","date":"2023-06-19","arxiv_id":"2306.10858","repositories_listed":2,"syntology":null},{"url":"/paper/where-a-strong-backbone-meets-strong-features","title":"Where a Strong Backbone Meets Strong Features -- ActionFormer for Ego4D Moment Queries Challenge","date":"2022-11-16","arxiv_id":"2211.09074","repositories_listed":2,"syntology":null},{"url":"/paper/weakly-supervised-action-localization-via","title":"Weakly-Supervised Temporal Action Localization by Progressive Complementary Learning","date":"2022-06-22","arxiv_id":"2206.11011","repositories_listed":2,"syntology":null},{"url":"/paper/structured-attention-composition-for-temporal","title":"Structured Attention Composition for Temporal Action Localization","date":"2022-05-20","arxiv_id":"2205.09956","repositories_listed":2,"syntology":null},{"url":"/paper/contextualized-spatio-temporal-contrastive","title":"Contextualized Spatio-Temporal Contrastive Learning with Self-Supervision","date":"2021-12-09","arxiv_id":"2112.05181","repositories_listed":2,"syntology":null},{"url":"/paper/videoclip-contrastive-pre-training-for-zero","title":"VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding","date":"2021-09-28","arxiv_id":"2109.14084","repositories_listed":2,"syntology":null},{"url":"/paper/cross-modal-consensus-network-for-weakly","title":"Cross-modal Consensus Network for Weakly Supervised Temporal Action Localization","date":"2021-07-27","arxiv_id":"2107.12589","repositories_listed":2,"syntology":null},{"url":"/paper/acm-net-action-context-modeling-network-for","title":"ACM-Net: Action Context Modeling Network for Weakly-Supervised Temporal Action Localization","date":"2021-04-07","arxiv_id":"2104.02967","repositories_listed":2,"syntology":null},{"url":"/paper/background-modeling-via-uncertainty","title":"Weakly-supervised Temporal Action Localization by Uncertainty Modeling","date":"2020-06-12","arxiv_id":"2006.07006","repositories_listed":2,"syntology":null},{"url":"/paper/learning-sparse-2d-temporal-adjacent-networks","title":"Learning Sparse 2D Temporal Adjacent Networks for Temporal Action Localization","date":"2019-12-08","arxiv_id":"1912.03612","repositories_listed":2,"syntology":null},{"url":"/paper/background-suppression-network-for-weakly","title":"Background Suppression Network for Weakly-supervised Temporal Action Localization","date":"2019-11-22","arxiv_id":"1911.09963","repositories_listed":2,"syntology":null},{"url":"/paper/deep-concept-wise-temporal-convolutional","title":"Deep Concept-wise Temporal Convolutional Networks for Action Localization","date":"2019-08-26","arxiv_id":"1908.09442","repositories_listed":2,"syntology":null},{"url":"/paper/hide-and-seek-a-data-augmentation-technique","title":"Hide-and-Seek: A Data Augmentation Technique for Weakly-Supervised Localization and Beyond","date":"2018-11-06","arxiv_id":"1811.02545","repositories_listed":2,"syntology":null},{"url":"/paper/guess-where-actor-supervision-for","title":"Guess Where? Actor-Supervision for Spatiotemporal Action Localization","date":"2018-04-05","arxiv_id":"1804.01824","repositories_listed":2,"syntology":null},{"url":"/paper/hacs-human-action-clips-and-segments-dataset","title":"HACS: Human Action Clips and Segments Dataset for Recognition and Temporal Localization","date":"2017-12-26","arxiv_id":"1712.09374","repositories_listed":2,"syntology":{"n":1,"n_ran":0,"n_unverified":1,"n_pointer_only":0}},{"url":"/paper/action-tubelet-detector-for-spatio-temporal","title":"Action Tubelet Detector for Spatio-Temporal Action Localization","date":"2017-05-04","arxiv_id":"1705.01861","repositories_listed":2,"syntology":null},{"url":"/paper/zero-shot-temporal-interaction-localization","title":"Zero-Shot Temporal Interaction Localization for Egocentric Videos","date":"2025-06-04","arxiv_id":"2506.03662","repositories_listed":1,"syntology":null}],"syntology_records":7,"syntology_note":"a paper without a record is not a recorded non-run: it may lack an arXiv id or simply be absent from the graph layer"},"description_links":{"kept":0,"unwrapped_to_text":0,"bare_urls_linked":0,"relative_images_dropped":0,"rule":"internal links are kept only when the target slug exists in the catalog"},"syntology":{"read_at":"2026-09-24T18:15:14+00:00","claim":"Per-sample execution status on synthesized fixtures ('ran N of M samples'); not a correctness claim and not a ranking signal.","status_vocabulary":{"ran_honours":"ran, honoured the contract we drafted","ran_violates":"ran, violated the contract we drafted","ran_draft_wrong":"ran; our contract draft was wrong, not the code","ran_fixture":"ran; our fixture could not drive it","ran":"ran on a synthesized input","unverified":"unverified (harvested, no recorded run)"}},"not_shown":{"libraries":"the archive has no per-task library table","trend_sparklines":"the Trend column of the benchmarks table was a rendered image; it is not in the archive","social_and_latest_sorts":"stars and social signals are not in the archive"}}