{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/semi-supervised-active-clustering-with-weak","title":"Semi-Supervised Active Clustering with Weak Oracles","arxiv_id":"1709.03202","date":"2017-09-11","proceeding":null,"authors":["Taewan Kim","Joydeep Ghosh"],"abstract":"Semi-supervised active clustering (SSAC) utilizes the knowledge of a domain\nexpert to cluster data points by interactively making pairwise \"same-cluster\"\nqueries. However, it is impractical to ask human oracles to answer every\npairwise query. In this paper, we study the influence of allowing \"not-sure\"\nanswers from a weak oracle and propose algorithms to efficiently handle\nuncertainties. Different types of model assumptions are analyzed to cover\nrealistic scenarios of oracle abstraction. In the first model, random-weak\noracle, an oracle randomly abstains with a certain probability. We also\nproposed two distance-weak oracle models which simulate the case of getting\nconfused based on the distance between two points in a pairwise query. For each\nweak oracle model, we show that a small query complexity is adequate for the\neffective $k$ means clustering with high probability. Sufficient conditions for\nthe guarantee include a $\\gamma$-margin property of the data, and an existence\nof a point close to each cluster center. Furthermore, we provide a sample\ncomplexity with a reduced effect of the cluster's margin and only a logarithmic\ndependency on the data dimension. Our results allow significantly less number\nof same-cluster queries if the margin of the clusters is tight, i.e. $\\gamma\n\\approx 1$. Experimental results on synthetic data show the effective\nperformance of our approach in overcoming uncertainties.","url_abs":"http://arxiv.org/abs/1709.03202v1","url_pdf":"http://arxiv.org/pdf/1709.03202v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"semi-supervised-active-clustering-with-weak","repo_url":"https://github.com/twankim/weaksemi","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"none","reach":null}],"tasks":[{"task_slug":"clustering","task_name":"Clustering"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}