{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/mining-mid-level-visual-patterns-with-deep","title":"Mining Mid-level Visual Patterns with Deep CNN Activations","arxiv_id":"1506.06343","date":"2015-06-21","proceeding":null,"authors":["Yao Li","Lingqiao Liu","Chunhua Shen","Anton Van Den Hengel"],"abstract":"The purpose of mid-level visual element discovery is to find clusters of\nimage patches that are both representative and discriminative. Here we study\nthis problem from the prospective of pattern mining while relying on the\nrecently popularized Convolutional Neural Networks (CNNs). We observe that a\nfully-connected CNN activation extracted from an image patch typically\npossesses two appealing properties that enable its seamless integration with\npattern mining techniques. The marriage between CNN activations and association\nrule mining, a well-known pattern mining technique in the literature, leads to\nfast and effective discovery of representative and discriminative patterns from\na huge number of image patches. When we retrieve and visualize image patches\nwith the same pattern, surprisingly, they are not only visually similar but\nalso semantically consistent, and thus give rise to a mid-level visual element\nin our work. Given the patterns and retrieved mid-level visual elements, we\npropose two methods to generate image feature representations for each. The\nfirst method is to use the patterns as codewords in a dictionary, similar to\nthe Bag-of-Visual-Words model, we compute a Bag-of-Patterns representation. The\nsecond one relies on the retrieved mid-level visual elements to construct a\nBag-of-Elements representation. We evaluate the two encoding methods on scene\nand object classification tasks, and demonstrate that our approach outperforms\nor matches recent works using CNN activations for these tasks.","url_abs":"http://arxiv.org/abs/1506.06343v3","url_pdf":"http://arxiv.org/pdf/1506.06343v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"mining-mid-level-visual-patterns-with-deep","repo_url":"https://github.com/yaoliUoA/MDPM","is_official":1,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":null}],"tasks":[],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}