{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/a-temporal-fusion-approach-for-video","title":"A Temporal Fusion Approach for Video Classification with Convolutional and LSTM Neural Networks Applied to Violence Detection","arxiv_id":null,"date":"2021-02-20","proceeding":"Inteligencia Artificial 2021 2","authors":["Jean Phelipe de Oliveira Lima","Carlos Maur´ıcio Ser´odio Figueiredo"],"abstract":"In modern smart cities, there is a quest for the highest level of integration and automation service. In\r\nthe surveillance sector, one of the main challenges is to automate the analysis of videos in real-time to identify\r\ncritical situations. This paper presents intelligent models based on Convolutional Neural Networks (in which the\r\nMobileNet, InceptionV3 and VGG16 networks had used), LSTM networks and feedforward networks for the task\r\nof classifying videos under the classes \"Violence\" and \"Non-Violence\", using for this the RLVS database. Different\r\ndata representations held used according to the Temporal Fusion techniques. The best outcome achieved was 0.91\r\nand 0.90 of Accuracy and F1-Score, respectively, a higher result compared to those found in similar researches\r\nfor works conducted on the same database.","url_abs":"https://journal.iberamia.org/index.php/intartif/article/view/573","url_pdf":"https://journal.iberamia.org/index.php/intartif/article/view/573/138","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"action-recognition-in-videos","task_name":"Action Recognition"},{"task_slug":"video-classification","task_name":"Video Classification"}],"methods":[{"method_slug":null,"method_name":null},{"method_slug":"lstm","method_name":"LSTM"},{"method_slug":"sigmoid-activation","method_name":"Sigmoid Activation"},{"method_slug":"tanh-activation","method_name":"Tanh Activation"}],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/action-recognition-on-real-life-violence","task":"Action Recognition","dataset":"Real Life Violence Situations Dataset","model":"Temporal Fusion cnn+lstm","rank_in_archive_order":2,"of":3,"metrics":{"accuracy":"91%"},"uses_additional_data":false}],"syntology":{"syntology_url":null,"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}