{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/neural-article-pair-modeling-for-wikipedia","title":"Neural Article Pair Modeling for Wikipedia Sub-article Matching","arxiv_id":"1807.11689","date":"2018-07-31","proceeding":null,"authors":["Muhao Chen","Changping Meng","Gang Huang","Carlo Zaniolo"],"abstract":"Nowadays, editors tend to separate different subtopics of a long Wiki-pedia\narticle into multiple sub-articles. This separation seeks to improve human\nreadability. However, it also has a deleterious effect on many Wikipedia-based\ntasks that rely on the article-as-concept assumption, which requires each\nentity (or concept) to be described solely by one article. This underlying\nassumption significantly simplifies knowledge representation and extraction,\nand it is vital to many existing technologies such as automated knowledge base\nconstruction, cross-lingual knowledge alignment, semantic search and data\nlineage of Wikipedia entities. In this paper we provide an approach to match\nthe scattered sub-articles back to their corresponding main-articles, with the\nintent of facilitating automated Wikipedia curation and processing. The\nproposed model adopts a hierarchical learning structure that combines multiple\nvariants of neural document pair encoders with a comprehensive set of explicit\nfeatures. A large crowdsourced dataset is created to support the evaluation and\nfeature extraction for the task. Based on the large dataset, the proposed model\nachieves promising results of cross-validation and significantly outperforms\nprevious approaches. Large-scale serving on the entire English Wikipedia also\nproves the practicability and scalability of the proposed model by effectively\nextracting a vast collection of newly paired main and sub-articles.","url_abs":"http://arxiv.org/abs/1807.11689v2","url_pdf":"http://arxiv.org/pdf/1807.11689v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"neural-article-pair-modeling-for-wikipedia","repo_url":"https://github.com/muhaochen/subarticle","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"tf","reach":null}],"tasks":[{"task_slug":"articles","task_name":"Articles"},{"task_slug":"knowledge-base-construction","task_name":"Knowledge Base Construction"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1807.11689","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}