{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/cast-contrastive-adaptation-and-distillation","title":"CAST: Contrastive Adaptation and Distillation for Semi-Supervised Instance Segmentation","arxiv_id":"2505.21904","date":"2025-05-28","proceeding":null,"authors":["Pardis Taghavi","Tian Liu","Renjie Li","Reza Langari","Zhengzhong Tu"],"abstract":"Instance segmentation demands costly per-pixel annotations and large models. We introduce CAST, a semi-supervised knowledge distillation (SSKD) framework that compresses pretrained vision foundation models (VFM) into compact experts using limited labeled and abundant unlabeled data. CAST unfolds in three stages: (1) domain adaptation of the VFM teacher(s) via self-training with contrastive pixel calibration, (2) distillation into a compact student via a unified multi-objective loss that couples standard supervision and pseudo-labels with our instance-aware pixel-wise contrastive term, and (3) fine-tuning on labeled data to remove residual pseudo-label bias. Central to CAST is an \\emph{instance-aware pixel-wise contrastive loss} that fuses mask and class scores to mine informative negatives and enforce clear inter-instance margins. By maintaining this contrastive signal across both adaptation and distillation, we align teacher and student embeddings and fully leverage unlabeled images. On Cityscapes and ADE20K, our ~11X smaller student surpasses its adapted VFM teacher(s) by +3.4 AP (33.9 vs. 30.5) and +1.5 AP (16.7 vs. 15.2) and outperforms state-of-the-art semi-supervised approaches.","url_abs":"https://arxiv.org/abs/2505.21904v2","url_pdf":"https://arxiv.org/pdf/2505.21904v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"domain-adaptation","task_name":"Domain Adaptation"},{"task_slug":"instance-segmentation","task_name":"Instance Segmentation"},{"task_slug":"knowledge-distillation","task_name":"Knowledge Distillation"},{"task_slug":"pseudo-label","task_name":"Pseudo Label"},{"task_slug":"semantic-segmentation","task_name":"Semantic Segmentation"},{"task_slug":"semi-supervised-instance-segmentation","task_name":"Semi-Supervised Instance Segmentation"}],"methods":[{"method_slug":"align","method_name":"ALIGN"},{"method_slug":"knowledge-distillation","method_name":"Knowledge Distillation"}],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/instance-segmentation-on-cityscapes-1","task":"Instance Segmentation","dataset":"Cityscapes","model":"CAST","rank_in_archive_order":1,"of":1,"metrics":{"AP":"33.9"},"uses_additional_data":false},{"leaderboard":"/sota/knowledge-distillation-on-cityscapes","task":"Knowledge Distillation","dataset":"Cityscapes","model":"CAST","rank_in_archive_order":1,"of":1,"metrics":{"AP":"33.9"},"uses_additional_data":false},{"leaderboard":"/sota/semi-supervised-instance-segmentation-on-1","task":"Semi-Supervised Instance Segmentation","dataset":"ADE20K","model":"CAST","rank_in_archive_order":1,"of":1,"metrics":{"AP":"16.7"},"uses_additional_data":false},{"leaderboard":"/sota/semi-supervised-instance-segmentation-on","task":"Semi-Supervised Instance Segmentation","dataset":"Cityscapes","model":"CAST","rank_in_archive_order":1,"of":1,"metrics":{"AP":"33.9"},"uses_additional_data":false}],"syntology":{"syntology_url":null,"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}