{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/improving-the-representation-and-conversion","title":"Improving the Representation and Conversion of Mathematical Formulae by Considering their Textual Context","arxiv_id":"1804.04956","date":"2018-04-13","proceeding":null,"authors":["Schubotz Moritz","Greiner-Petter Andre","Scharpf Philipp","Meuschke Norman","Cohl Howard","Gipp Bela"],"abstract":"Mathematical formulae represent complex semantic information in a concise\nform. Especially in Science, Technology, Engineering, and Mathematics,\nmathematical formulae are crucial to communicate information, e.g., in\nscientific papers, and to perform computations using computer algebra systems.\nEnabling computers to access the information encoded in mathematical formulae\nrequires machine-readable formats that can represent both the presentation and\ncontent, i.e., the semantics, of formulae. Exchanging such information between\nsystems additionally requires conversion methods for mathematical\nrepresentation formats. We analyze how the semantic enrichment of formulae\nimproves the format conversion process and show that considering the textual\ncontext of formulae reduces the error rate of such conversions. Our main\ncontributions are: (1) providing an openly available benchmark dataset for the\nmathematical format conversion task consisting of a newly created test\ncollection, an extensive, manually curated gold standard and task-specific\nevaluation metrics; (2) performing a quantitative evaluation of\nstate-of-the-art tools for mathematical format conversions; (3) presenting a\nnew approach that considers the textual context of formulae to reduce the error\nrate for mathematical format conversions. Our benchmark dataset facilitates\nfuture research on mathematical format conversions as well as research on many\nproblems in mathematical information retrieval. Because we annotated and linked\nall components of formulae, e.g., identifiers, operators and other entities, to\nWikidata entries, the gold standard can, for instance, be used to train methods\nfor formula concept discovery and recognition. Such methods can then be applied\nto improve mathematical information retrieval systems, e.g., for semantic\nformula search, recommendation of mathematical content, or detection of\nmathematical plagiarism.","url_abs":"http://arxiv.org/abs/1804.04956v1","url_pdf":"http://arxiv.org/pdf/1804.04956v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"improving-the-representation-and-conversion","repo_url":"https://github.com/ag-gipp/MathMLben","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":null}],"tasks":[{"task_slug":"information-retrieval","task_name":"Information Retrieval"},{"task_slug":"retrieval","task_name":"Retrieval"}],"methods":[],"datasets_introduced":[{"slug":"mathmlben","name":"MathMLben","full_name":"Formula semantics benchmark"}],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1804.04956","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}