{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/web2text-deep-structured-boilerplate-removal","title":"Web2Text: Deep Structured Boilerplate Removal","arxiv_id":"1801.02607","date":"2018-03-27","proceeding":null,"authors":["Vogels Thijs","Ganea Octavian-Eugen","Eickhoff Carsten"],"abstract":"Web pages are a valuable source of information for many natural language\nprocessing and information retrieval tasks. Extracting the main content from\nthose documents is essential for the performance of derived applications. To\naddress this issue, we introduce a novel model that performs sequence labeling\nto collectively classify all text blocks in an HTML page as either boilerplate\nor main content. Our method uses a hidden Markov model on top of potentials\nderived from DOM tree features using convolutional neural networks. The\nproposed method sets a new state-of-the-art performance for boilerplate removal\non the CleanEval benchmark. As a component of information retrieval pipelines,\nit improves retrieval performance on the ClueWeb12 collection.","url_abs":"http://arxiv.org/abs/1801.02607v3","url_pdf":"http://arxiv.org/pdf/1801.02607v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"web2text-deep-structured-boilerplate-removal","repo_url":"https://github.com/dalab/web2text","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"tf","reach":null}],"tasks":[{"task_slug":"information-retrieval","task_name":"Information Retrieval"},{"task_slug":"retrieval","task_name":"Retrieval"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}