{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/data-readiness-levels","title":"Data Readiness Levels","arxiv_id":"1705.02245","date":"2017-05-05","proceeding":null,"authors":["Neil D. Lawrence"],"abstract":"Application of models to data is fraught. Data-generating collaborators often\nonly have a very basic understanding of the complications of collating,\nprocessing and curating data. Challenges include: poor data collection\npractices, missing values, inconvenient storage mechanisms, intellectual\nproperty, security and privacy. All these aspects obstruct the sharing and\ninterconnection of data, and the eventual interpretation of data through\nmachine learning or other approaches. In project reporting, a major challenge\nis in encapsulating these problems and enabling goals to be built around the\nprocessing of data. Project overruns can occur due to failure to account for\nthe amount of time required to curate and collate. But to understand these\nfailures we need to have a common language for assessing the readiness of a\nparticular data set. This position paper proposes the use of data readiness\nlevels: it gives a rough outline of three stages of data preparedness and\nspeculates on how formalisation of these levels into a common language for data\nreadiness could facilitate project management.","url_abs":"http://arxiv.org/abs/1705.02245v1","url_pdf":"http://arxiv.org/pdf/1705.02245v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"data-readiness-levels","repo_url":"https://github.com/fredriko/draviz","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":{"status":"unanswered"}}],"tasks":[{"task_slug":"management","task_name":"Management"},{"task_slug":"missing-values","task_name":"Missing Values"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1705.02245","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}