{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/self-supervised-learning-of-echocardiographic","title":"Self-supervised Learning of Echocardiographic Video Representations via Online Cluster Distillation","arxiv_id":"2506.11777","date":"2025-06-13","proceeding":null,"authors":["Divyanshu Mishra","Mohammadreza Salehi","Pramit Saha","Olga Patey","Aris T. Papageorghiou","Yuki M. Asano","J. Alison Noble"],"abstract":"Self-supervised learning (SSL) has achieved major advances in natural images and video understanding, but challenges remain in domains like echocardiography (heart ultrasound) due to subtle anatomical structures, complex temporal dynamics, and the current lack of domain-specific pre-trained models. Existing SSL approaches such as contrastive, masked modeling, and clustering-based methods struggle with high intersample similarity, sensitivity to low PSNR inputs common in ultrasound, or aggressive augmentations that distort clinically relevant features. We present DISCOVR (Distilled Image Supervision for Cross Modal Video Representation), a self-supervised dual branch framework for cardiac ultrasound video representation learning. DISCOVR combines a clustering-based video encoder that models temporal dynamics with an online image encoder that extracts fine-grained spatial semantics. These branches are connected through a semantic cluster distillation loss that transfers anatomical knowledge from the evolving image encoder to the video encoder, enabling temporally coherent representations enriched with fine-grained semantic understanding. Evaluated on six echocardiography datasets spanning fetal, pediatric, and adult populations, DISCOVR outperforms both specialized video anomaly detection methods and state-of-the-art video-SSL baselines in zero-shot and linear probing setups, and achieves superior segmentation transfer.","url_abs":"https://arxiv.org/abs/2506.11777v1","url_pdf":"https://arxiv.org/pdf/2506.11777v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"self-supervised-learning-of-echocardiographic","repo_url":"https://github.com/mdivyanshu97/discovr","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"anomaly-detection","task_name":"Anomaly Detection"},{"task_slug":"clustering","task_name":"Clustering"},{"task_slug":"representation-learning","task_name":"Representation Learning"},{"task_slug":"self-supervised-learning","task_name":"Self-Supervised Learning"},{"task_slug":"video-anomaly-detection","task_name":"Video Anomaly Detection"},{"task_slug":"video-understanding","task_name":"Video Understanding"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2506.11777","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2506.11777"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/mdivyanshu97/discovr","reach":null}],"summary":{"ran_draft_wrong":2,"ran":3,"ran_fixture":3,"unverified":2},"by_repo_kind":{"official":{"samples":10,"ran":8,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":10,"samples":[{"code_sha256_prefix":"4619da37e019258a","entry":"clean_state_dict","repo":"mdivyanshu97/discovr","repo_kind":"official","path":"run_discovr_encoder.py","file_url":"https://github.com/mdivyanshu97/discovr/blob/HEAD/run_discovr_encoder.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":2,"contract_check":"OUTPUT_MISDECLARED","metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"4619da37e019258a"}},{"code_sha256_prefix":"9cb7158ad2ea3981","entry":"DINOHead","repo":"mdivyanshu97/discovr","repo_kind":"official","path":"models/modeling_pretrain.py","file_url":"https://github.com/mdivyanshu97/discovr/blob/HEAD/models/modeling_pretrain.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"9cb7158ad2ea3981"}},{"code_sha256_prefix":"955f481a84b0e8d0","entry":"PretrainVisionTransformerDecoder","repo":"mdivyanshu97/discovr","repo_kind":"official","path":"models/modeling_pretrain.py","file_url":"https://github.com/mdivyanshu97/discovr/blob/HEAD/models/modeling_pretrain.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"955f481a84b0e8d0"}},{"code_sha256_prefix":"df29747ca407c6d4","entry":"SparseTubesTokenizer","repo":"mdivyanshu97/discovr","repo_kind":"official","path":"models/modeling_pretrain.py","file_url":"https://github.com/mdivyanshu97/discovr/blob/HEAD/models/modeling_pretrain.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"df29747ca407c6d4"}},{"code_sha256_prefix":"a260a74207539014","entry":"create_teacher_model","repo":"mdivyanshu97/discovr","repo_kind":"official","path":"models/modeling_pretrain.py","file_url":"https://github.com/mdivyanshu97/discovr/blob/HEAD/models/modeling_pretrain.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"OUTPUT_MISDECLARED","metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"a260a74207539014"}},{"code_sha256_prefix":"44aacaf688b398e0","entry":"get_1d_sincos_pos_embed_from_grid","repo":"mdivyanshu97/discovr","repo_kind":"official","path":"models/modeling_pretrain.py","file_url":"https://github.com/mdivyanshu97/discovr/blob/HEAD/models/modeling_pretrain.py","link_basis":"first_harvest_node","language":"python","status":"ran_fixture","verification_level":1,"contract_check":"RAISES","metamorphic_tier":null,"behaviour_fingerprint":true,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"44aacaf688b398e0"}},{"code_sha256_prefix":"715aab11d1b73ce8","entry":"get_2d_sincos_pos_embed_from_grid","repo":"mdivyanshu97/discovr","repo_kind":"official","path":"models/modeling_pretrain.py","file_url":"https://github.com/mdivyanshu97/discovr/blob/HEAD/models/modeling_pretrain.py","link_basis":"first_harvest_node","language":"python","status":"ran_fixture","verification_level":1,"contract_check":"RAISES","metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"715aab11d1b73ce8"}},{"code_sha256_prefix":"a12beb78a97ac9a5","entry":"get_3d_sincos_pos_embed","repo":"mdivyanshu97/discovr","repo_kind":"official","path":"models/modeling_pretrain.py","file_url":"https://github.com/mdivyanshu97/discovr/blob/HEAD/models/modeling_pretrain.py","link_basis":"first_harvest_node","language":"python","status":"ran_fixture","verification_level":1,"contract_check":"RAISES","metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"a12beb78a97ac9a5"}},{"code_sha256_prefix":"7edd01c772cf84e1","entry":"PretrainVisionTransformer","repo":"mdivyanshu97/discovr","repo_kind":"official","path":"models/modeling_pretrain.py","file_url":"https://github.com/mdivyanshu97/discovr/blob/HEAD/models/modeling_pretrain.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"7edd01c772cf84e1"}},{"code_sha256_prefix":"3ccdc115cd8472ef","entry":"PretrainVisionTransformerEncoder","repo":"mdivyanshu97/discovr","repo_kind":"official","path":"models/modeling_pretrain.py","file_url":"https://github.com/mdivyanshu97/discovr/blob/HEAD/models/modeling_pretrain.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"3ccdc115cd8472ef"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}