{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/learning-from-less-data-a-unified-data-subset","title":"Learning From Less Data: A Unified Data Subset Selection and Active Learning Framework for Computer Vision","arxiv_id":"1901.01151","date":"2019-01-03","proceeding":null,"authors":["Vishal Kaushal","Rishabh Iyer","Suraj Kothawade","Rohan Mahadev","Khoshrav Doctor","Ganesh Ramakrishnan"],"abstract":"Supervised machine learning based state-of-the-art computer vision techniques\nare in general data hungry. Their data curation poses the challenges of\nexpensive human labeling, inadequate computing resources and larger experiment\nturn around times. Training data subset selection and active learning\ntechniques have been proposed as possible solutions to these challenges. A\nspecial class of subset selection functions naturally model notions of\ndiversity, coverage and representation and can be used to eliminate redundancy\nthus lending themselves well for training data subset selection. They can also\nhelp improve the efficiency of active learning in further reducing human\nlabeling efforts by selecting a subset of the examples obtained using the\nconventional uncertainty sampling based techniques. In this work, we\nempirically demonstrate the effectiveness of two diversity models, namely the\nFacility-Location and Dispersion models for training-data subset selection and\nreducing labeling effort. We demonstrate this across the board for a variety of\ncomputer vision tasks including Gender Recognition, Face Recognition, Scene\nRecognition, Object Detection and Object Recognition. Our results show that\ndiversity based subset selection done in the right way can increase the\naccuracy by upto 5 - 10% over existing baselines, particularly in settings in\nwhich less training data is available. This allows the training of complex\nmachine learning models like Convolutional Neural Networks with much less\ntraining data and labeling costs while incurring minimal performance loss.","url_abs":"http://arxiv.org/abs/1901.01151v1","url_pdf":"http://arxiv.org/pdf/1901.01151v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"learning-from-less-data-a-unified-data-subset","repo_url":"https://github.com/caihuaiguang/CHG-Shapley-for-Data-Selection","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null},{"paper_slug":"learning-from-less-data-a-unified-data-subset","repo_url":"https://github.com/decile-team/cords","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"MIT"}}],"tasks":[{"task_slug":"active-learning","task_name":"Active Learning"},{"task_slug":"machine-learning","task_name":"BIG-bench Machine Learning"},{"task_slug":"diversity","task_name":"Diversity"},{"task_slug":"face-recognition","task_name":"Face Recognition"},{"task_slug":"object-detection","task_name":"Object Detection"},{"task_slug":"object-recognition","task_name":"Object Recognition"},{"task_slug":"scene-recognition","task_name":"Scene Recognition"},{"task_slug":"object-detection-1","task_name":"object-detection"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1901.01151","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}