{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/a-clustering-based-combinatorial-approach-to","title":"A Clustering-Based Combinatorial Approach to Unsupervised Matching of Product Titles","arxiv_id":"1903.04276","date":"2019-03-07","proceeding":null,"authors":["Leonidas Akritidis","Athanasios Fevgas","Panayiotis Bozanis","Christos Makris"],"abstract":"The constant growth of the e-commerce industry has rendered the problem of\nproduct retrieval particularly important. As more enterprises move their\nactivities on the Web, the volume and the diversity of the product-related\ninformation increase quickly. These factors make it difficult for the users to\nidentify and compare the features of their desired products. Recent studies\nproved that the standard similarity metrics cannot effectively identify\nidentical products, since similar titles often refer to different products and\nvice-versa. Other studies employed external data sources (search engines) to\nenrich the titles; these solutions are rather impractical mainly because the\nexternal data fetching is slow. In this paper we introduce UPM, an unsupervised\nalgorithm for matching products by their titles. UPM is independent of any\nexternal sources, since it analyzes the titles and extracts combinations of\nwords out of them. These combinations are evaluated according to several\ncriteria, and the most appropriate of them constitutes the cluster where a\nproduct is classified into. UPM is also parameter-free, it avoids product\npairwise comparisons, and includes a post-processing verification stage which\ncorrects the erroneous matches. The experimental evaluation of UPM demonstrated\nits superiority against the state-of-the-art approaches in terms of both\nefficiency and effectiveness.","url_abs":"http://arxiv.org/abs/1903.04276v1","url_pdf":"http://arxiv.org/pdf/1903.04276v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"a-clustering-based-combinatorial-approach-to","repo_url":"https://github.com/BinaryWiz/UPM","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":null},{"paper_slug":"a-clustering-based-combinatorial-approach-to","repo_url":"https://github.com/BinaryWiz/Unsupervised-Product-Matching-Using-Combinations-and-Permutations","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":null}],"tasks":[{"task_slug":"clustering","task_name":"Clustering"},{"task_slug":"retrieval","task_name":"Retrieval"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}