{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/provable-guarantees-for-self-supervised-deep","title":"Provable Guarantees for Self-Supervised Deep Learning with Spectral Contrastive Loss","arxiv_id":"2106.04156","date":"2021-06-08","proceeding":"NeurIPS 2021 12","authors":["Jeff Z. HaoChen","Colin Wei","Adrien Gaidon","Tengyu Ma"],"abstract":"Recent works in self-supervised learning have advanced the state-of-the-art by relying on the contrastive learning paradigm, which learns representations by pushing positive pairs, or similar examples from the same class, closer together while keeping negative pairs far apart. Despite the empirical successes, theoretical foundations are limited -- prior analyses assume conditional independence of the positive pairs given the same class label, but recent empirical applications use heavily correlated positive pairs (i.e., data augmentations of the same image). Our work analyzes contrastive learning without assuming conditional independence of positive pairs using a novel concept of the augmentation graph on data. Edges in this graph connect augmentations of the same data, and ground-truth classes naturally form connected sub-graphs. We propose a loss that performs spectral decomposition on the population augmentation graph and can be succinctly written as a contrastive learning objective on neural net representations. Minimizing this objective leads to features with provable accuracy guarantees under linear probe evaluation. By standard generalization bounds, these accuracy guarantees also hold when minimizing the training contrastive loss. Empirically, the features learned by our objective can match or outperform several strong baselines on benchmark vision datasets. In all, this work provides the first provable analysis for contrastive learning where guarantees for linear probe evaluation can apply to realistic empirical settings.","url_abs":"https://arxiv.org/abs/2106.04156v7","url_pdf":"https://arxiv.org/pdf/2106.04156v7.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"provable-guarantees-for-self-supervised-deep","repo_url":"https://github.com/jhaochenz/spectral_contrastive_learning","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null},{"paper_slug":"provable-guarantees-for-self-supervised-deep","repo_url":"https://github.com/jhaochenz96/spectral_contrastive_learning","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"contrastive-learning","task_name":"Contrastive Learning"},{"task_slug":"generalization-bounds","task_name":"Generalization Bounds"},{"task_slug":"self-supervised-learning","task_name":"Self-Supervised Learning"}],"methods":[{"method_slug":"contrastive-learning","method_name":"Contrastive Learning"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2106.04156","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2106.04156"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/jhaochenz/spectral_contrastive_learning","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/jhaochenz96/spectral_contrastive_learning","reach":null}],"summary":{"ran_draft_wrong":1,"ran":2},"by_repo_kind":{"listed":{"samples":3,"ran":3,"repositories":2}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"cd6d57e7c8432808","entry":"D","repo":"jhaochenz96/spectral_contrastive_learning","repo_kind":"listed","path":"models/spectral.py","file_url":"https://github.com/jhaochenz96/spectral_contrastive_learning/blob/HEAD/models/spectral.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"MISDECLARED","metamorphic_tier":"deterministic","behaviour_fingerprint":true,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"cd6d57e7c8432808"}},{"code_sha256_prefix":"86d68d2245cab2f2","entry":"Spectral","repo":"jhaochenz/spectral_contrastive_learning","repo_kind":"listed","path":"models/spectral.py","file_url":"https://github.com/jhaochenz/spectral_contrastive_learning/blob/HEAD/models/spectral.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"86d68d2245cab2f2"}},{"code_sha256_prefix":"55c30097ea15926c","entry":"projection_identity","repo":"jhaochenz/spectral_contrastive_learning","repo_kind":"listed","path":"models/spectral.py","file_url":"https://github.com/jhaochenz/spectral_contrastive_learning/blob/HEAD/models/spectral.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":true,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"55c30097ea15926c"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}