{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/rate-accuracy-trade-off-in-video","title":"Rate-Accuracy Trade-Off In Video Classification With Deep Convolutional Neural Networks","arxiv_id":"1810.03964","date":"2018-09-27","proceeding":null,"authors":["Mohammad Jubran","Alhabib Abbas","Aaron Chadha","Yiannis Andreopoulos"],"abstract":"Advanced video classification systems decode video frames to derive the\nnecessary texture and motion representations for ingestion and analysis by\nspatio-temporal deep convolutional neural networks (CNNs). However, when\nconsidering visual Internet-of-Things applications, surveillance systems and\nsemantic crawlers of large video repositories, the video capture and the\nCNN-based semantic analysis parts do not tend to be co-located. This\nnecessitates the transport of compressed video over networks and incurs\nsignificant overhead in bandwidth and energy consumption, thereby significantly\nundermining the deployment potential of such systems. In this paper, we\ninvestigate the trade-off between the encoding bitrate and the achievable\naccuracy of CNN-based video classification models that directly ingest\nAVC/H.264 and HEVC encoded videos. Instead of retaining entire compressed video\nbitstreams and applying complex optical flow calculations prior to CNN\nprocessing, we only retain motion vector and select texture information at\nsignificantly-reduced bitrates and apply no additional processing prior to CNN\ningestion. Based on three CNN architectures and two action recognition\ndatasets, we achieve 11%-94% saving in bitrate with marginal effect on\nclassification accuracy. A model-based selection between multiple CNNs\nincreases these savings further, to the point where, if up to 7% loss of\naccuracy can be tolerated, video classification can take place with as little\nas 3 kbps for the transport of the required compressed video information to the\nsystem implementing the CNN models.","url_abs":"http://arxiv.org/abs/1810.03964v2","url_pdf":"http://arxiv.org/pdf/1810.03964v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"rate-accuracy-trade-off-in-video","repo_url":"https://github.com/rate-accuracy-mvcnn/main","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"none","reach":null}],"tasks":[{"task_slug":"action-recognition-in-videos","task_name":"Action Recognition"},{"task_slug":"classification-1","task_name":"Classification"},{"task_slug":"classification","task_name":"General Classification"},{"task_slug":"optical-flow-estimation","task_name":"Optical Flow Estimation"},{"task_slug":"action-recognition","task_name":"Temporal Action Localization"},{"task_slug":"video-classification","task_name":"Video Classification"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}