{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/a-mid-level-video-representation-based-on","title":"A Mid-level Video Representation based on Binary Descriptors: A Case Study for Pornography Detection","arxiv_id":"1605.03804","date":"2016-05-12","proceeding":null,"authors":["Carlos Caetano","Sandra Avila","William Robson Schwartz","Silvio Jamil F. Guimarães","Arnaldo de A. Araújo"],"abstract":"With the growing amount of inappropriate content on the Internet, such as\npornography, arises the need to detect and filter such material. The reason for\nthis is given by the fact that such content is often prohibited in certain\nenvironments (e.g., schools and workplaces) or for certain publics (e.g.,\nchildren). In recent years, many works have been mainly focused on detecting\npornographic images and videos based on visual content, particularly on the\ndetection of skin color. Although these approaches provide good results, they\ngenerally have the disadvantage of a high false positive rate since not all\nimages with large areas of skin exposure are necessarily pornographic images,\nsuch as people wearing swimsuits or images related to sports. Local feature\nbased approaches with Bag-of-Words models (BoW) have been successfully applied\nto visual recognition tasks in the context of pornography detection. Even\nthough existing methods provide promising results, they use local feature\ndescriptors that require a high computational processing time yielding\nhigh-dimensional vectors. In this work, we propose an approach for pornography\ndetection based on local binary feature extraction and BossaNova image\nrepresentation, a BoW model extension that preserves more richly the visual\ninformation. Moreover, we propose two approaches for video description based on\nthe combination of mid-level representations namely BossaNova Video Descriptor\n(BNVD) and BoW Video Descriptor (BoW-VD). The proposed techniques are\npromising, achieving an accuracy of 92.40%, thus reducing the classification\nerror by 16% over the current state-of-the-art local features approach on the\nPornography dataset.","url_abs":"http://arxiv.org/abs/1605.03804v1","url_pdf":"http://arxiv.org/pdf/1605.03804v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"a-mid-level-video-representation-based-on","repo_url":"https://github.com/jackaduma/nude-detect","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"none","reach":null}],"tasks":[{"task_slug":"pornography-detection","task_name":"Pornography Detection"},{"task_slug":"video-description","task_name":"Video Description"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}