{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/effective-aesthetics-prediction-with-multi","title":"Effective Aesthetics Prediction with Multi-level Spatially Pooled Features","arxiv_id":"1904.01382","date":"2019-04-02","proceeding":"CVPR 2019 6","authors":["Vlad Hosu","Bastian Goldlucke","Dietmar Saupe"],"abstract":"We propose an effective deep learning approach to aesthetics quality\nassessment that relies on a new type of pre-trained features, and apply it to\nthe AVA data set, the currently largest aesthetics database. While previous\napproaches miss some of the information in the original images, due to taking\nsmall crops, down-scaling or warping the originals during training, we propose\nthe first method that efficiently supports full resolution images as an input,\nand can be trained on variable input sizes. This allows us to significantly\nimprove upon the state of the art, increasing the Spearman rank-order\ncorrelation coefficient (SRCC) of ground-truth mean opinion scores (MOS) from\nthe existing best reported of 0.612 to 0.756. To achieve this performance, we\nextract multi-level spatially pooled (MLSP) features from all convolutional\nblocks of a pre-trained InceptionResNet-v2 network, and train a custom shallow\nConvolutional Neural Network (CNN) architecture on these new features.","url_abs":"http://arxiv.org/abs/1904.01382v1","url_pdf":"http://arxiv.org/pdf/1904.01382v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"effective-aesthetics-prediction-with-multi","repo_url":"https://github.com/subpic/ava-mlsp","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"tf","reach":{"status":"ok","spdx":"MIT"}}],"tasks":[{"task_slug":"aesthetics-quality-assessment","task_name":"Aesthetics Quality Assessment"},{"task_slug":"image-quality-assessment","task_name":"Image Quality Assessment"},{"task_slug":"prediction","task_name":"Prediction"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/aesthetics-quality-assessment-on-ava","task":"Aesthetics Quality Assessment","dataset":"AVA","model":"Pool-3FC","rank_in_archive_order":3,"of":9,"metrics":{"Accuracy":"81.7%"},"uses_additional_data":false}],"syntology":{"syntology_url":"https://syntology.ai/paper/1904.01382","atlas_url":"https://app.syntology.ai/?focus=1904.01382","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}