{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/divide-and-grow-capturing-huge-diversity-in","title":"Divide and Grow: Capturing Huge Diversity in Crowd Images with Incrementally Growing CNN","arxiv_id":"1807.09993","date":"2018-07-26","proceeding":"CVPR 2018 6","authors":["Deepak Babu Sam","Neeraj N Sajjan","R. Venkatesh Babu"],"abstract":"Automated counting of people in crowd images is a challenging task. The major\ndifficulty stems from the large diversity in the way people appear in crowds.\nIn fact, features available for crowd discrimination largely depend on the\ncrowd density to the extent that people are only seen as blobs in a highly\ndense scene. We tackle this problem with a growing CNN which can progressively\nincrease its capacity to account for the wide variability seen in crowd scenes.\nOur model starts from a base CNN density regressor, which is trained in\nequivalence on all types of crowd images. In order to adapt with the huge\ndiversity, we create two child regressors which are exact copies of the base\nCNN. A differential training procedure divides the dataset into two clusters\nand fine-tunes the child networks on their respective specialties.\nConsequently, without any hand-crafted criteria for forming specialties, the\nchild regressors become experts on certain types of crowds. The child networks\nare again split recursively, creating two experts at every division. This\nhierarchical training leads to a CNN tree, where the child regressors are more\nfine experts than any of their parents. The leaf nodes are taken as the final\nexperts and a classifier network is then trained to predict the correct\nspecialty for a given test image patch. The proposed model achieves higher\ncount accuracy on major crowd datasets. Further, we analyse the characteristics\nof specialties mined automatically by our method.","url_abs":"http://arxiv.org/abs/1807.09993v1","url_pdf":"http://arxiv.org/pdf/1807.09993v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"crowd-counting","task_name":"Crowd Counting"},{"task_slug":"diversity","task_name":"Diversity"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/crowd-counting-on-shanghaitech-a","task":"Crowd Counting","dataset":"ShanghaiTech A","model":"IG-CNN","rank_in_archive_order":26,"of":35,"metrics":{"MAE":"72.5"},"uses_additional_data":false},{"leaderboard":"/sota/crowd-counting-on-shanghaitech-b","task":"Crowd Counting","dataset":"ShanghaiTech B","model":"IG-CNN","rank_in_archive_order":24,"of":32,"metrics":{"MAE":"13.6"},"uses_additional_data":false},{"leaderboard":"/sota/crowd-counting-on-ucf-cc-50","task":"Crowd Counting","dataset":"UCF CC 50","model":"IG-CNN","rank_in_archive_order":15,"of":22,"metrics":{"MAE":"291.4"},"uses_additional_data":false},{"leaderboard":"/sota/crowd-counting-on-worldexpo10","task":"Crowd Counting","dataset":"WorldExpo’10","model":"IG-CNN","rank_in_archive_order":13,"of":15,"metrics":{"Average MAE":"11.3"},"uses_additional_data":false}],"syntology":{"syntology_url":"https://syntology.ai/paper/1807.09993","atlas_url":"https://app.syntology.ai/?focus=1807.09993","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}