{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/improved-infilling-of-missing-metadata-from","title":"Improved Infilling of Missing Metadata from Expendable Bathythermographs (XBTs) Using Multiple Machine Learning Methods","arxiv_id":null,"date":"2022-09-01","proceeding":"Journal of Atmospheric and Oceanic Technology 2022 9","authors":["Stephen Haddad","Rachel E. Killick","Matthew D. Palmer","Mark J. Webb","Rachel Prudden","Francesco Capponi","and Samantha V. Adams"],"abstract":"Historical in situ ocean temperature profile measurements are important for a wide range of ocean and\r\n\r\nclimate research activities. A large proportion of the profile observations have been recorded using expendable bathyther-\r\nmographs (XBTs), and required bias corrections for use in climate change studies. It is generally accepted that the bias,\r\n\r\nand therefore bias correction, depends on the type of XBT used. However, poor historical metadata collection practices\r\n\r\nmean the XBT probe type information is often missing, for 59% of profiles between 1967 and 2000, limiting the develop-\r\nment of reliable bias corrections. We develop a process of estimating missing instrument type metadata (the combination\r\n\r\nof both model and manufacturer) systematically, constructing a machine learning pipeline based on thorough data explo-\r\nration to inform these choices. The predicted instrument type, where missing, will facilitate improved XBT bias correc-\r\ntions. The new approach improves the accuracy of the XBT type classification compared to previous approaches from a\r\n\r\nrecall value of 0.75–0.94. We also develop an approach to account for the uncertainty associated with metadata assign-\r\nments using ensembles of decision trees, which could feed into an ensemble approach to creating ocean temperature data-\r\nsets. We describe the challenges arising from the nature of the dataset in applying standard machine learning techniques\r\n\r\nto the problem. We have implemented this in a portable, reproducible way using standard data science tools, with a view\r\nto these techniques being applied to other similar problems in climate science.","url_abs":"https://journals.ametsoc.org/view/journals/atot/39/9/JTECH-D-21-0117.1.xml?rskey=GVhzHc&result=5","url_pdf":"https://journals.ametsoc.org/view/journals/atot/39/9/JTECH-D-21-0117.1.xml?rskey=GVhzHc&result=5","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"improved-infilling-of-missing-metadata-from","repo_url":"https://github.com/MetOffice/XBTs_classification","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"none","reach":null}],"tasks":[{"task_slug":"type","task_name":"Vocal Bursts Type Prediction"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}