{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/uncovering-what-why-and-how-a-comprehensive","title":"Uncovering What, Why and How: A Comprehensive Benchmark for Causation Understanding of Video Anomaly","arxiv_id":"2405.00181","date":"2024-04-30","proceeding":null,"authors":["Hang Du","Sicheng Zhang","Binzhu Xie","Guoshun Nan","Jiayang Zhang","Junrui Xu","Hangyu Liu","Sicong Leng","Jiangming Liu","Hehe Fan","Dajiu Huang","Jing Feng","Linli Chen","Can Zhang","Xuhuan Li","Hao Zhang","Jianhang Chen","Qimei Cui","Xiaofeng Tao"],"abstract":"Video anomaly understanding (VAU) aims to automatically comprehend unusual occurrences in videos, thereby enabling various applications such as traffic surveillance and industrial manufacturing. While existing VAU benchmarks primarily concentrate on anomaly detection and localization, our focus is on more practicality, prompting us to raise the following crucial questions: \"what anomaly occurred?\", \"why did it happen?\", and \"how severe is this abnormal event?\". In pursuit of these answers, we present a comprehensive benchmark for Causation Understanding of Video Anomaly (CUVA). Specifically, each instance of the proposed benchmark involves three sets of human annotations to indicate the \"what\", \"why\" and \"how\" of an anomaly, including 1) anomaly type, start and end times, and event descriptions, 2) natural language explanations for the cause of an anomaly, and 3) free text reflecting the effect of the abnormality. In addition, we also introduce MMEval, a novel evaluation metric designed to better align with human preferences for CUVA, facilitating the measurement of existing LLMs in comprehending the underlying cause and corresponding effect of video anomalies. Finally, we propose a novel prompt-based method that can serve as a baseline approach for the challenging CUVA. We conduct extensive experiments to show the superiority of our evaluation metric and the prompt-based approach. Our code and dataset are available at https://github.com/fesvhtr/CUVA.","url_abs":"https://arxiv.org/abs/2405.00181v2","url_pdf":"https://arxiv.org/pdf/2405.00181v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"uncovering-what-why-and-how-a-comprehensive","repo_url":"https://github.com/fesvhtr/cuva","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":null},{"paper_slug":"uncovering-what-why-and-how-a-comprehensive","repo_url":"https://github.com/dulpy/ecva","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok"}}],"tasks":[{"task_slug":"anomaly-detection","task_name":"Anomaly Detection"}],"methods":[{"method_slug":"align","method_name":"ALIGN"},{"method_slug":"focus","method_name":"Focus"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/2405.00181","atlas_url":"https://app.syntology.ai/?focus=2405.00181","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2405.00181"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-25T09:33:49+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/dulpy/ecva","reach":{"status":"ok"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/fesvhtr/cuva","reach":null}],"summary":{"ran_draft_wrong":2},"by_repo_kind":{"official":{"samples":2,"ran":2,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":2,"samples":[{"code_sha256_prefix":"d473ef9bac3be71a","entry":"upload_video","repo":"fesvhtr/cuva","repo_kind":"official","path":"Models/Video-ChatGPT/video_chatgpt/CUVA/CUVA.py","file_url":"https://github.com/fesvhtr/cuva/blob/HEAD/Models/Video-ChatGPT/video_chatgpt/CUVA/CUVA.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"OUTPUT_MISDECLARED","metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"d473ef9bac3be71a"}},{"code_sha256_prefix":"7c32a225488d3d7d","entry":"upload_video","repo":"fesvhtr/cuva","repo_kind":"official","path":"Models/Video-ChatGPT/video_chatgpt/CUVA/mmEval.py","file_url":"https://github.com/fesvhtr/cuva/blob/HEAD/Models/Video-ChatGPT/video_chatgpt/CUVA/mmEval.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"OUTPUT_MISDECLARED","metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"7c32a225488d3d7d"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}