{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/extracting-scientific-figures-with-distantly","title":"Extracting Scientific Figures with Distantly Supervised Neural Networks","arxiv_id":"1804.02445","date":"2018-04-06","proceeding":null,"authors":["Noah Siegel","Nicholas Lourie","Russell Power","Waleed Ammar"],"abstract":"Non-textual components such as charts, diagrams and tables provide key\ninformation in many scientific documents, but the lack of large labeled\ndatasets has impeded the development of data-driven methods for scientific\nfigure extraction. In this paper, we induce high-quality training labels for\nthe task of figure extraction in a large number of scientific documents, with\nno human intervention. To accomplish this we leverage the auxiliary data\nprovided in two large web collections of scientific documents (arXiv and\nPubMed) to locate figures and their associated captions in the rasterized PDF.\nWe share the resulting dataset of over 5.5 million induced labels---4,000 times\nlarger than the previous largest figure extraction dataset---with an average\nprecision of 96.8%, to enable the development of modern data-driven methods for\nthis task. We use this dataset to train a deep neural network for end-to-end\nfigure detection, yielding a model that can be more easily extended to new\ndomains compared to previous work. The model was successfully deployed in\nSemantic Scholar, a large-scale academic search engine, and used to extract\nfigures in 13 million scientific documents.","url_abs":"http://arxiv.org/abs/1804.02445v2","url_pdf":"http://arxiv.org/pdf/1804.02445v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"extracting-scientific-figures-with-distantly","repo_url":"https://github.com/allenai/deepfigures-open","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"tf","reach":null}],"tasks":[],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":"https://app.syntology.ai/?focus=1804.02445","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}