{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/replication-study-development-and-validation","title":"Replication study: Development and validation of deep learning algorithm for detection of diabetic retinopathy in retinal fundus photographs","arxiv_id":"1803.04337","date":"2018-03-12","proceeding":null,"authors":["Mike Voets","Kajsa Møllersen","Lars Ailo Bongo"],"abstract":"Replication studies are essential for validation of new methods, and are\ncrucial to maintain the high standards of scientific publications, and to use\nthe results in practice. We have attempted to replicate the main method in\n'Development and validation of a deep learning algorithm for detection of\ndiabetic retinopathy in retinal fundus photographs' published in JAMA 2016;\n316(22). We re-implemented the method since the source code is not available,\nand we used publicly available data sets. The original study used non-public\nfundus images from EyePACS and three hospitals in India for training. We used a\ndifferent EyePACS data set from Kaggle. The original study used the benchmark\ndata set Messidor-2 to evaluate the algorithm's performance. We used the same\ndata set. In the original study, ophthalmologists re-graded all images for\ndiabetic retinopathy, macular edema, and image gradability. There was one\ndiabetic retinopathy grade per image for our data sets, and we assessed image\ngradability ourselves. Hyper-parameter settings were not described in the\noriginal study. But some of these were later published. We were not able to\nreplicate the original study. Our algorithm's area under the receiver operating\ncurve (AUC) of 0.94 on the Kaggle EyePACS test set and 0.80 on Messidor-2 did\nnot come close to the reported AUC of 0.99 in the original study. This may be\ncaused by the use of a single grade per image, different data, or different not\ndescribed hyper-parameter settings. This study shows the challenges of\nreplicating deep learning, and the need for more replication studies to\nvalidate deep learning methods, especially for medical image analysis.\n  Our source code and instructions are available at:\nhttps://github.com/mikevoets/jama16-retina-replication","url_abs":"http://arxiv.org/abs/1803.04337v3","url_pdf":"http://arxiv.org/pdf/1803.04337v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"replication-study-development-and-validation","repo_url":"https://github.com/mikevoets/jama16-retina-replication","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"tf","reach":null}],"tasks":[{"task_slug":"diabetic-retinopathy-detection","task_name":"Diabetic Retinopathy Detection"},{"task_slug":"medical-image-analysis","task_name":"Medical Image Analysis"},{"task_slug":"medical-image-segmentation","task_name":"Medical Image Segmentation"},{"task_slug":"mitosis-detection","task_name":"Mitosis Detection"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}