{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/reverse-knowledge-distillation-training-a","title":"Reverse Knowledge Distillation: Training a Large Model using a Small One for Retinal Image Matching on Limited Data","arxiv_id":"2307.10698","date":"2023-07-20","proceeding":null,"authors":["Sahar Almahfouz Nasser","Nihar Gupte","Amit Sethi"],"abstract":"Retinal image matching plays a crucial role in monitoring disease progression and treatment response. However, datasets with matched keypoints between temporally separated pairs of images are not available in abundance to train transformer-based model. We propose a novel approach based on reverse knowledge distillation to train large models with limited data while preventing overfitting. Firstly, we propose architectural modifications to a CNN-based semi-supervised method called SuperRetina that help us improve its results on a publicly available dataset. Then, we train a computationally heavier model based on a vision transformer encoder using the lighter CNN-based model, which is counter-intuitive in the field knowledge-distillation research where training lighter models based on heavier ones is the norm. Surprisingly, such reverse knowledge distillation improves generalization even further. Our experiments suggest that high-dimensional fitting in representation space may prevent overfitting unlike training directly to match the final output. We also provide a public dataset with annotations for retinal image keypoint detection and matching to help the research community develop algorithms for retinal image applications.","url_abs":"https://arxiv.org/abs/2307.10698v2","url_pdf":"https://arxiv.org/pdf/2307.10698v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"reverse-knowledge-distillation-training-a","repo_url":"https://github.com/SaharAlmahfouzNasser/MeDAL-Retina","is_official":1,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":null}],"tasks":[{"task_slug":"image-registration","task_name":"Image Registration"},{"task_slug":"keypoint-detection","task_name":"Keypoint Detection"},{"task_slug":"keypoint-detection-and-image-matching","task_name":"Keypoint detection and image matching"},{"task_slug":"knowledge-distillation","task_name":"Knowledge Distillation"}],"methods":[{"method_slug":"attention","method_name":"Attention"},{"method_slug":"dense-connections","method_name":"Dense Connections"},{"method_slug":"knowledge-distillation","method_name":"Knowledge Distillation"},{"method_slug":"layer-normalization","method_name":"Layer Normalization"},{"method_slug":"linear-layer","method_name":"Linear Layer"},{"method_slug":"multi-head-attention","method_name":"Multi-Head Attention"},{"method_slug":"residual-connection","method_name":"Residual Connection"},{"method_slug":"softmax","method_name":"Softmax"},{"method_slug":"vision-transformer","method_name":"Vision Transformer"}],"datasets_introduced":[{"slug":"medal-retina-dataset","name":"MeDAL Retina Dataset","full_name":"MeDAL  Retina Dataset"}],"methods_introduced":[],"results":[{"leaderboard":"/sota/image-registration-on-fire","task":"Image Registration","dataset":"FIRE","model":"LKRetina","rank_in_archive_order":1,"of":6,"metrics":{"mAUC":"0.761"},"uses_additional_data":false}],"syntology":{"syntology_url":null,"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}