{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/the-something-something-video-database-for","title":"The \"something something\" video database for learning and evaluating visual common sense","arxiv_id":"1706.04261","date":"2017-06-13","proceeding":"ICCV 2017 10","authors":["Raghav Goyal","Samira Ebrahimi Kahou","Vincent Michalski","Joanna Materzyńska","Susanne Westphal","Heuna Kim","Valentin Haenel","Ingo Fruend","Peter Yianilos","Moritz Mueller-Freitag","Florian Hoppe","Christian Thurau","Ingo Bax","Roland Memisevic"],"abstract":"Neural networks trained on datasets such as ImageNet have led to major\nadvances in visual object classification. One obstacle that prevents networks\nfrom reasoning more deeply about complex scenes and situations, and from\nintegrating visual knowledge with natural language, like humans do, is their\nlack of common sense knowledge about the physical world. Videos, unlike still\nimages, contain a wealth of detailed information about the physical world.\nHowever, most labelled video datasets represent high-level concepts rather than\ndetailed physical aspects about actions and scenes. In this work, we describe\nour ongoing collection of the \"something-something\" database of video\nprediction tasks whose solutions require a common sense understanding of the\ndepicted situation. The database currently contains more than 100,000 videos\nacross 174 classes, which are defined as caption-templates. We also describe\nthe challenges in crowd-sourcing this data at scale.","url_abs":"http://arxiv.org/abs/1706.04261v2","url_pdf":"http://arxiv.org/pdf/1706.04261v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"the-something-something-video-database-for","repo_url":"https://github.com/akshyta/Human-Activity-Recognition","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"unanswered"}},{"paper_slug":"the-something-something-video-database-for","repo_url":"https://github.com/bit-ml/dyreg-gnn","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null},{"paper_slug":"the-something-something-video-database-for","repo_url":"https://github.com/caspillaga/Conv3DSelfAttention","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"unanswered"}},{"paper_slug":"the-something-something-video-database-for","repo_url":"https://github.com/jayleicn/singularity","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null},{"paper_slug":"the-something-something-video-database-for","repo_url":"https://github.com/latte488/smth-smth-v2","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"unanswered"}}],"tasks":[{"task_slug":"action-recognition-in-videos","task_name":"Action Recognition"},{"task_slug":"common-sense-reasoning","task_name":"Common Sense Reasoning"},{"task_slug":"classification","task_name":"General Classification"},{"task_slug":"video-prediction","task_name":"Video Prediction"}],"methods":[],"datasets_introduced":[{"slug":"something-something-v1","name":"Something-Something V1","full_name":""},{"slug":"something-something-v2","name":"Something-Something V2","full_name":""}],"methods_introduced":[],"results":[{"leaderboard":"/sota/action-recognition-in-videos-on-something","task":"Action Recognition","dataset":"Something-Something V2","model":"model3D_1 with left-right augmentation and fps jitter","rank_in_archive_order":116,"of":123,"metrics":{"Top-1 Accuracy":"51.33","Top-5 Accuracy":"80.46"},"uses_additional_data":false}],"syntology":{"syntology_url":"https://syntology.ai/paper/1706.04261","atlas_url":"https://app.syntology.ai/?focus=1706.04261","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}