{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/every-moment-counts-dense-detailed-labeling","title":"Every Moment Counts: Dense Detailed Labeling of Actions in Complex Videos","arxiv_id":"1507.05738","date":"2015-07-21","proceeding":null,"authors":["Serena Yeung","Olga Russakovsky","Ning Jin","Mykhaylo Andriluka","Greg Mori","Li Fei-Fei"],"abstract":"Every moment counts in action recognition. A comprehensive understanding of\nhuman activity in video requires labeling every frame according to the actions\noccurring, placing multiple labels densely over a video sequence. To study this\nproblem we extend the existing THUMOS dataset and introduce MultiTHUMOS, a new\ndataset of dense labels over unconstrained internet videos. Modeling multiple,\ndense labels benefits from temporal relations within and across classes. We\ndefine a novel variant of long short-term memory (LSTM) deep networks for\nmodeling these temporal relations via multiple input and output connections. We\nshow that this model improves action labeling accuracy and further enables\ndeeper understanding tasks ranging from structured retrieval to action\nprediction.","url_abs":"http://arxiv.org/abs/1507.05738v3","url_pdf":"http://arxiv.org/pdf/1507.05738v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"every-moment-counts-dense-detailed-labeling","repo_url":"https://github.com/lauradhatt/Interesting-Reads","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":{"status":"unanswered"}}],"tasks":[{"task_slug":"action-recognition-in-videos","task_name":"Action Recognition"},{"task_slug":"retrieval","task_name":"Retrieval"},{"task_slug":"action-recognition","task_name":"Temporal Action Localization"}],"methods":[],"datasets_introduced":[{"slug":"multithumos","name":"MultiTHUMOS","full_name":""}],"methods_introduced":[],"results":[{"leaderboard":"/sota/action-detection-on-multi-thumos","task":"Action Detection","dataset":"Multi-THUMOS","model":"Two-stream + LSTM","rank_in_archive_order":7,"of":8,"metrics":{"mAP":"28.1"},"uses_additional_data":false},{"leaderboard":"/sota/action-detection-on-multi-thumos","task":"Action Detection","dataset":"Multi-THUMOS","model":"Two-stream","rank_in_archive_order":8,"of":8,"metrics":{"mAP":"27.6"},"uses_additional_data":false}],"syntology":{"syntology_url":"https://syntology.ai/paper/1507.05738","atlas_url":"https://app.syntology.ai/?focus=1507.05738","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}