{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/sparse-adversarial-perturbations-for-videos","title":"Sparse Adversarial Perturbations for Videos","arxiv_id":"1803.02536","date":"2018-03-07","proceeding":null,"authors":["Xingxing Wei","Jun Zhu","Hang Su"],"abstract":"Although adversarial samples of deep neural networks (DNNs) have been\nintensively studied on static images, their extensions in videos are never\nexplored. Compared with images, attacking a video needs to consider not only\nspatial cues but also temporal cues. Moreover, to improve the imperceptibility\nas well as reduce the computation cost, perturbations should be added on as\nfewer frames as possible, i.e., adversarial perturbations are temporally\nsparse. This further motivates the propagation of perturbations, which denotes\nthat perturbations added on the current frame can transfer to the next frames\nvia their temporal interactions. Thus, no (or few) extra perturbations are\nneeded for these frames to misclassify them. To this end, we propose an\nl2,1-norm based optimization algorithm to compute the sparse adversarial\nperturbations for videos. We choose the action recognition as the targeted\ntask, and networks with a CNN+RNN architecture as threat models to verify our\nmethod. Thanks to the propagation, we can compute perturbations on a shortened\nversion video, and then adapt them to the long version video to fool DNNs.\nExperimental results on the UCF101 dataset demonstrate that even only one frame\nin a video is perturbed, the fooling rate can still reach 59.7%.","url_abs":"http://arxiv.org/abs/1803.02536v1","url_pdf":"http://arxiv.org/pdf/1803.02536v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"sparse-adversarial-perturbations-for-videos","repo_url":"https://github.com/yanhui002/video_adv","is_official":1,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"tf","reach":{"status":"ok"}},{"paper_slug":"sparse-adversarial-perturbations-for-videos","repo_url":"https://github.com/anonymous-p/Flickering_Adversarial_Video","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok"}},{"paper_slug":"sparse-adversarial-perturbations-for-videos","repo_url":"https://github.com/thakurnupur/Sparse-Adversarial-Perturbations-PyTorch","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok"}}],"tasks":[{"task_slug":"action-recognition-in-videos","task_name":"Action Recognition"},{"task_slug":"action-recognition","task_name":"Temporal Action Localization"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":"https://app.syntology.ai/?focus=1803.02536","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}