{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/multilingual-clustering-of-streaming-news","title":"Multilingual Clustering of Streaming News","arxiv_id":"1809.00540","date":"2018-09-03","proceeding":"EMNLP 2018 10","authors":["Sebastião Miranda","Artūrs Znotiņš","Shay B. Cohen","Guntis Barzdins"],"abstract":"Clustering news across languages enables efficient media monitoring by\naggregating articles from multilingual sources into coherent stories. Doing so\nin an online setting allows scalable processing of massive news streams. To\nthis end, we describe a novel method for clustering an incoming stream of\nmultilingual documents into monolingual and crosslingual story clusters. Unlike\ntypical clustering approaches that consider a small and known number of labels,\nwe tackle the problem of discovering an ever growing number of cluster labels\nin an online fashion, using real news datasets in multiple languages. Our\nmethod is simple to implement, computationally efficient and produces\nstate-of-the-art results on datasets in German, English and Spanish.","url_abs":"http://arxiv.org/abs/1809.00540v1","url_pdf":"http://arxiv.org/pdf/1809.00540v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"multilingual-clustering-of-streaming-news","repo_url":"https://github.com/priberam/news-clustering","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"none","reach":null},{"paper_slug":"multilingual-clustering-of-streaming-news","repo_url":"https://github.com/bourrel/French-News-Clustering","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok"}}],"tasks":[{"task_slug":"articles","task_name":"Articles"},{"task_slug":"clustering","task_name":"Clustering"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}