{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/generating-high-quality-crowd-density-maps","title":"Generating High-Quality Crowd Density Maps using Contextual Pyramid CNNs","arxiv_id":"1708.00953","date":"2017-08-02","proceeding":"ICCV 2017 10","authors":["Vishwanath A. Sindagi","Vishal M. Patel"],"abstract":"We present a novel method called Contextual Pyramid CNN (CP-CNN) for\ngenerating high-quality crowd density and count estimation by explicitly\nincorporating global and local contextual information of crowd images. The\nproposed CP-CNN consists of four modules: Global Context Estimator (GCE), Local\nContext Estimator (LCE), Density Map Estimator (DME) and a Fusion-CNN (F-CNN).\nGCE is a VGG-16 based CNN that encodes global context and it is trained to\nclassify input images into different density classes, whereas LCE is another\nCNN that encodes local context information and it is trained to perform\npatch-wise classification of input images into different density classes. DME\nis a multi-column architecture-based CNN that aims to generate high-dimensional\nfeature maps from the input image which are fused with the contextual\ninformation estimated by GCE and LCE using F-CNN. To generate high resolution\nand high-quality density maps, F-CNN uses a set of convolutional and\nfractionally-strided convolutional layers and it is trained along with the DME\nin an end-to-end fashion using a combination of adversarial loss and\npixel-level Euclidean loss. Extensive experiments on highly challenging\ndatasets show that the proposed method achieves significant improvements over\nthe state-of-the-art methods.","url_abs":"http://arxiv.org/abs/1708.00953v1","url_pdf":"http://arxiv.org/pdf/1708.00953v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"crowd-counting","task_name":"Crowd Counting"},{"task_slug":"high","task_name":"Vocal Bursts Intensity Prediction"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/crowd-counting-on-shanghaitech-a","task":"Crowd Counting","dataset":"ShanghaiTech A","model":"CP-CNN","rank_in_archive_order":28,"of":35,"metrics":{"MAE":"73.6"},"uses_additional_data":false},{"leaderboard":"/sota/crowd-counting-on-shanghaitech-b","task":"Crowd Counting","dataset":"ShanghaiTech B","model":"CP-CNN","rank_in_archive_order":29,"of":32,"metrics":{"MAE":"20.1"},"uses_additional_data":false},{"leaderboard":"/sota/crowd-counting-on-ucf-cc-50","task":"Crowd Counting","dataset":"UCF CC 50","model":"CP-CNN","rank_in_archive_order":16,"of":22,"metrics":{"MAE":"295.8"},"uses_additional_data":false},{"leaderboard":"/sota/crowd-counting-on-worldexpo10","task":"Crowd Counting","dataset":"WorldExpo’10","model":"CP-CNN","rank_in_archive_order":8,"of":15,"metrics":{"Average MAE":"8.9"},"uses_additional_data":false}],"syntology":{"syntology_url":null,"atlas_url":"https://app.syntology.ai/?focus=1708.00953","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}