{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/growing-story-forest-online-from-massive","title":"Growing Story Forest Online from Massive Breaking News","arxiv_id":"1803.00189","date":"2018-03-01","proceeding":null,"authors":["Bang Liu","Di Niu","Kunfeng Lai","Linglong Kong","Yu Xu"],"abstract":"We describe our experience of implementing a news content organization system\nat Tencent that discovers events from vast streams of breaking news and evolves\nnews story structures in an online fashion. Our real-world system has distinct\nrequirements in contrast to previous studies on topic detection and tracking\n(TDT) and event timeline or graph generation, in that we 1) need to accurately\nand quickly extract distinguishable events from massive streams of long text\ndocuments that cover diverse topics and contain highly redundant information,\nand 2) must develop the structures of event stories in an online manner,\nwithout repeatedly restructuring previously formed stories, in order to\nguarantee a consistent user viewing experience. In solving these challenges, we\npropose Story Forest, a set of online schemes that automatically clusters\nstreaming documents into events, while connecting related events in growing\ntrees to tell evolving stories. We conducted extensive evaluation based on 60\nGB of real-world Chinese news data, although our ideas are not\nlanguage-dependent and can easily be extended to other languages, through\ndetailed pilot user experience studies. The results demonstrate the superior\ncapability of Story Forest to accurately identify events and organize news text\ninto a logical structure that is appealing to human readers, compared to\nmultiple existing algorithm frameworks.","url_abs":"http://arxiv.org/abs/1803.00189v1","url_pdf":"http://arxiv.org/pdf/1803.00189v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"growing-story-forest-online-from-massive","repo_url":"https://github.com/BangLiu/StoryForest","is_official":1,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"none","reach":null}],"tasks":[{"task_slug":"graph-generation","task_name":"Graph Generation"},{"task_slug":"information-threading","task_name":"Information Threading"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/information-threading-on-newshead","task":"Information Threading","dataset":"NewSHead","model":"EventX","rank_in_archive_order":3,"of":4,"metrics":{"NMI":"0.2405"},"uses_additional_data":false}],"syntology":{"syntology_url":"https://syntology.ai/paper/1803.00189","atlas_url":"https://app.syntology.ai/?focus=1803.00189","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}