{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/sampling-bias-in-deep-active-classification","title":"Sampling Bias in Deep Active Classification: An Empirical Study","arxiv_id":"1909.09389","date":"2019-09-20","proceeding":"IJCNLP 2019 11","authors":["Ameya Prabhu","Charles Dognin","Maneesh Singh"],"abstract":"The exploding cost and time needed for data labeling and model training are bottlenecks for training DNN models on large datasets. Identifying smaller representative data samples with strategies like active learning can help mitigate such bottlenecks. Previous works on active learning in NLP identify the problem of sampling bias in the samples acquired by uncertainty-based querying and develop costly approaches to address it. Using a large empirical study, we demonstrate that active set selection using the posterior entropy of deep models like FastText.zip (FTZ) is robust to sampling biases and to various algorithmic choices (query size and strategies) unlike that suggested by traditional literature. We also show that FTZ based query strategy produces sample sets similar to those from more sophisticated approaches (e.g ensemble networks). Finally, we show the effectiveness of the selected samples by creating tiny high-quality datasets, and utilizing them for fast and cheap training of large models. Based on the above, we propose a simple baseline for deep active text classification that outperforms the state-of-the-art. We expect the presented work to be useful and informative for dataset compression and for problems involving active, semi-supervised or online learning scenarios. Code and models are available at: https://github.com/drimpossible/Sampling-Bias-Active-Learning","url_abs":"https://arxiv.org/abs/1909.09389v1","url_pdf":"https://arxiv.org/pdf/1909.09389v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"sampling-bias-in-deep-active-classification","repo_url":"https://github.com/drimpossible/Sampling-Bias-Active-Learning","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok"}},{"paper_slug":"sampling-bias-in-deep-active-classification","repo_url":"https://github.com/Xtra-Computing/thundersvm","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"none","reach":{"status":"ok","spdx":"Apache-2.0"}}],"tasks":[{"task_slug":"active-learning","task_name":"Active Learning"},{"task_slug":"classification-1","task_name":"Classification"},{"task_slug":"classification","task_name":"General Classification"},{"task_slug":"text-classification","task_name":"Text Classification"},{"task_slug":"text-classification-1","task_name":"text-classification"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/text-classification-on-ag-news","task":"Text Classification","dataset":"AG News","model":"ULMFiT (Small data)","rank_in_archive_order":7,"of":24,"metrics":{"Error":"6.3"},"uses_additional_data":false},{"leaderboard":"/sota/text-classification-on-amazon-2","task":"Text Classification","dataset":"Amazon-2","model":"ULMFiT (Small data)","rank_in_archive_order":3,"of":4,"metrics":{"Error":"3.9"},"uses_additional_data":false},{"leaderboard":"/sota/text-classification-on-amazon-5","task":"Text Classification","dataset":"Amazon-5","model":"ULMFiT (Small data)","rank_in_archive_order":2,"of":3,"metrics":{"Error":"35.9"},"uses_additional_data":false},{"leaderboard":"/sota/text-classification-on-dbpedia","task":"Text Classification","dataset":"DBpedia","model":"ULMFiT (Small data)","rank_in_archive_order":7,"of":21,"metrics":{"Error":"0.8"},"uses_additional_data":false},{"leaderboard":"/sota/text-classification-on-sogou-news","task":"Text Classification","dataset":"Sogou News","model":"ULMFiT (Small data)","rank_in_archive_order":3,"of":3,"metrics":{"Accuracy":"97"},"uses_additional_data":false},{"leaderboard":"/sota/text-classification-on-yahoo-answers","task":"Text Classification","dataset":"Yahoo! Answers","model":"ULMFiT (Small data)","rank_in_archive_order":6,"of":10,"metrics":{"Accuracy":"74.3"},"uses_additional_data":false},{"leaderboard":"/sota/text-classification-on-yelp-2","task":"Text Classification","dataset":"Yelp-2","model":"ULMFiT (Small data)","rank_in_archive_order":4,"of":5,"metrics":{"Accuracy":"97.1%"},"uses_additional_data":false},{"leaderboard":"/sota/text-classification-on-yelp-5","task":"Text Classification","dataset":"Yelp-5","model":"ULMFiT (Small data)","rank_in_archive_order":7,"of":7,"metrics":{"Accuracy":"67.6%"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1909.09389","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"1909.09389"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/drimpossible/Sampling-Bias-Active-Learning","reach":{"status":"ok"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/Xtra-Computing/thundersvm","reach":{"status":"ok","spdx":"Apache-2.0"}}],"summary":{"unverified":1},"by_repo_kind":{"official":{"samples":1,"ran":0,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"80f8630d0b01b42b","entry":"svm_read_problem","repo":"Xtra-Computing/thundersvm","repo_kind":"official","path":"python/svm.py","file_url":"https://github.com/Xtra-Computing/thundersvm/blob/HEAD/python/svm.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"80f8630d0b01b42b"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}