{"url":"/method/contrastive-predictive-coding","slug":"contrastive-predictive-coding","name":"Contrastive Predictive Coding","full_name":"Contrastive Predictive Coding","full_name_withheld":false,"description_markdown":"**Contrastive Predictive Coding (CPC)** learns self-supervised representations by predicting the future in latent space by using powerful autoregressive models. The model uses a probabilistic contrastive loss which induces the latent space to capture information that is maximally useful\r\nto predict future samples.\r\n\r\nFirst, a non-linear encoder $g\\_{enc}$ maps the input sequence of observations $x\\_{t}$ to a sequence of latent representations $z\\_{t} = g\\_{enc}\\left(x\\_{t}\\right)$, potentially with a lower temporal resolution. Next, an autoregressive model $g\\_{ar}$ summarizes all $z\\leq{t}$ in the latent space and produces a context latent representation $c\\_{t} = g\\_{ar}\\left(z\\leq{t}\\right)$.\r\n\r\nA density ratio is modelled which preserves the mutual information between $x\\_{t+k}$ and $c\\_{t}$ as follows:\r\n\r\n$$ f\\_{k}\\left(x\\_{t+k}, c\\_{t}\\right) \\propto \\frac{p\\left(x\\_{t+k}|c\\_{t}\\right)}{p\\left(x\\_{t+k}\\right)} $$\r\n\r\nwhere $\\propto$ stands for ’proportional to’ (i.e. up to a multiplicative constant). Note that the density ratio $f$ can be unnormalized (does not have to integrate to 1). The authors use a simple log-bilinear model:\r\n\r\n$$ f\\_{k}\\left(x\\_{t+k}, c\\_{t}\\right) = \\exp\\left(z^{T}\\_{t+k}W\\_{k}c\\_{t}\\right) $$\r\n\r\nAny type of autoencoder and autoregressive can be used. An example the authors opt for is strided convolutional layers with residual blocks and GRUs.\r\n\r\nThe autoencoder and autoregressive models are trained to minimize an [InfoNCE](https://paperswithcode.com/method/infonce) loss (see components).","description_state":"present","introduced_year":null,"introduced_by":{"title":"Representation Learning with Contrastive Predictive Coding","paper":"/paper/representation-learning-with-contrastive","first_author":"Aaron van den Oord","n_authors":3,"url_abs":null,"archive_paper_url":"https://paperswithcode.com/paper/representation-learning-with-contrastive"},"source":{"url":"http://arxiv.org/abs/1807.03748v2","title":"Representation Learning with Contrastive Predictive Coding","url_on_a_paper_host":true},"code_snippet_url":"https://github.com/jefflai108/Contrastive-Predictive-Coding-PyTorch","code_snippet_url_on_a_code_host":true,"categories":[{"area":"General","area_id":"general","collection":"Semi-Supervised Learning Methods","url":"/methods/category/semi-supervised-learning-methods","pwc_aliases":[]},{"area":"General","area_id":"general","collection":"Self-Supervised Learning","url":"/methods/category/self-supervised-learning","pwc_aliases":[]}],"n_papers_tagged":113,"archive_num_papers":113,"papers_newest_first":[{"paper":null,"title":"Two-Player Zero-Sum Games with Bandit Feedback","date":"2025-06-17","arxiv_id":"2506.14518","n_code_links":0,"syntology":null},{"paper":"/paper/integration-of-contrastive-predictive-coding","title":"Integration of Contrastive Predictive Coding and Spiking Neural Networks","date":"2025-06-10","arxiv_id":"2506.09194","n_code_links":1,"syntology":null},{"paper":null,"title":"Koopman-Based Event-Triggered Control from Data","date":"2025-04-19","arxiv_id":"2504.14334","n_code_links":0,"syntology":null},{"paper":null,"title":"Learning Transformer-based World Models with Contrastive Predictive Coding","date":"2025-03-06","arxiv_id":"2503.04416","n_code_links":0,"syntology":null},{"paper":null,"title":"Contrastive Representation Learning Helps Cross-institutional Knowledge Transfer: A Study in Pediatric Ventilation Management","date":"2025-01-23","arxiv_id":"2501.13587","n_code_links":0,"syntology":null},{"paper":null,"title":"Performance-Barrier Event-Triggered PDE Control of Traffic Flow","date":"2025-01-01","arxiv_id":"2501.00722","n_code_links":0,"syntology":null},{"paper":null,"title":"Automated Toll Management System Using RFID and Image Processing","date":"2024-12-02","arxiv_id":"2412.01728","n_code_links":0,"syntology":null},{"paper":null,"title":"A Contrastive Self-Supervised Learning scheme for beat tracking amenable to few-shot learning","date":"2024-11-06","arxiv_id":"2411.04152","n_code_links":0,"syntology":null},{"paper":null,"title":"Trading through Earnings Seasons using Self-Supervised Contrastive Representation Learning","date":"2024-09-25","arxiv_id":"2409.17392","n_code_links":0,"syntology":null},{"paper":"/paper/context-aware-predictive-coding-a","title":"Context-Aware Predictive Coding: A Representation Learning Framework for WiFi Sensing","date":"2024-09-20","arxiv_id":null,"n_code_links":1,"syntology":null},{"paper":"/paper/context-aware-predictive-coding-a-1","title":"Context-Aware Predictive Coding: A Representation Learning Framework for WiFi Sensing","date":"2024-09-16","arxiv_id":"2410.01825","n_code_links":1,"syntology":null},{"paper":null,"title":"Hierarchical Event-Triggered Systems: Safe Learning of Quasi-Optimal Deadline Policies","date":"2024-09-15","arxiv_id":"2409.09812","n_code_links":0,"syntology":null},{"paper":null,"title":"Speaker and Style Disentanglement of Speech Based on Contrastive Predictive Coding Supported Factorized Variational Autoencoder","date":"2024-09-05","arxiv_id":"2409.03520","n_code_links":0,"syntology":null},{"paper":"/paper/contrastive-representation-learning-for-3","title":"Contrastive Representation Learning for Dynamic Link Prediction in Temporal Networks","date":"2024-08-22","arxiv_id":"2408.12753","n_code_links":1,"syntology":null},{"paper":null,"title":"Performance-Barrier Event-Triggered Control of a Class of Reaction-Diffusion PDEs","date":"2024-07-11","arxiv_id":"2407.08178","n_code_links":0,"syntology":null},{"paper":null,"title":"Contextual Dynamic Pricing: Algorithms, Optimality, and Local Differential Privacy Constraints","date":"2024-06-04","arxiv_id":"2406.02424","n_code_links":0,"syntology":null},{"paper":null,"title":"Causal Contrastive Learning for Counterfactual Regression Over Time","date":"2024-06-01","arxiv_id":"2406.00535","n_code_links":0,"syntology":null},{"paper":"/paper/offline-reinforcement-learning-from-datasets","title":"Offline Reinforcement Learning from Datasets with Structured Non-Stationarity","date":"2024-05-23","arxiv_id":"2405.14114","n_code_links":1,"syntology":{"ran":10,"of":12,"unverified":2,"pointer_only":0}},{"paper":null,"title":"Multilingual Turn-taking Prediction Using Voice Activity Projection","date":"2024-03-11","arxiv_id":"2403.06487","n_code_links":0,"syntology":null},{"paper":null,"title":"Event-Triggered Robust Cooperative Output Regulation for a Class of Linear Multi-Agent Systems with an Unknown Exosystem","date":"2024-03-01","arxiv_id":"2403.00645","n_code_links":0,"syntology":null},{"paper":null,"title":"Revisiting speech segmentation and lexicon learning with better features","date":"2024-01-31","arxiv_id":"2401.17902","n_code_links":0,"syntology":null},{"paper":"/paper/real-time-and-continuous-turn-taking","title":"Real-time and Continuous Turn-taking Prediction Using Voice Activity Projection","date":"2024-01-10","arxiv_id":"2401.04868","n_code_links":2,"syntology":null},{"paper":null,"title":"Replication-proof Bandit Mechanism Design with Bayesian Agents","date":"2023-12-28","arxiv_id":"2312.16896","n_code_links":0,"syntology":null},{"paper":"/paper/bigger-is-not-always-better-the-effect-of","title":"Bigger is not Always Better: The Effect of Context Size on Speech Pre-Training","date":"2023-12-03","arxiv_id":"2312.01515","n_code_links":1,"syntology":null},{"paper":"/paper/representation-learning-in-a-decomposed","title":"Representation Learning in a Decomposed Encoder Design for Bio-inspired Hebbian Learning","date":"2023-11-22","arxiv_id":"2401.08603","n_code_links":1,"syntology":null},{"paper":null,"title":"Self-MI: Efficient Multimodal Fusion via Self-Supervised Multi-Task Learning with Auxiliary Mutual Information Maximization","date":"2023-11-07","arxiv_id":"2311.03785","n_code_links":0,"syntology":null},{"paper":"/paper/contrastive-difference-predictive-coding","title":"Contrastive Difference Predictive Coding","date":"2023-10-31","arxiv_id":"2310.20141","n_code_links":1,"syntology":{"ran":5,"of":5,"unverified":0,"pointer_only":0}},{"paper":null,"title":"Learning-based Scheduling for Information Accuracy and Freshness in Wireless Networks","date":"2023-10-24","arxiv_id":"2310.15705","n_code_links":0,"syntology":null},{"paper":null,"title":"Listen to Minority: Encrypted Traffic Classification for Class Imbalance with Contrastive Pre-Training","date":"2023-08-31","arxiv_id":"2308.16453","n_code_links":0,"syntology":null},{"paper":"/paper/towards-quantitative-precision-for-ecg","title":"Towards quantitative precision for ECG analysis: Leveraging state space models, self-supervision and patient metadata","date":"2023-08-29","arxiv_id":"2308.15291","n_code_links":1,"syntology":{"ran":5,"of":7,"unverified":2,"pointer_only":0}}],"papers_shown":30,"tasks":[{"task":"/task/representation-learning","name":"Representation Learning","papers":23},{"task":"/task/self-supervised-learning","name":"Self-Supervised Learning","papers":17},{"task":"/task/contrastive-learning","name":"Contrastive Learning","papers":10},{"task":"/task/speech-recognition","name":"Speech Recognition","papers":10},{"task":"/task/speech-recognition-1","name":"speech-recognition","papers":10},{"task":"/task/automatic-speech-recognition-2","name":"Automatic Speech Recognition","papers":9},{"task":"/task/automatic-speech-recognition","name":"Automatic Speech Recognition (ASR)","papers":7},{"task":null,"name":"Acoustic Unit Discovery","papers":6},{"task":"/task/transfer-learning","name":"Transfer Learning","papers":6},{"task":"/task/voice-conversion","name":"Voice Conversion","papers":6},{"task":"/task/anomaly-detection","name":"Anomaly Detection","papers":5},{"task":"/task/language-modelling","name":"Language Modelling","papers":5},{"task":"/task/activity-recognition","name":"Activity Recognition","papers":4},{"task":"/task/classification","name":"General Classification","papers":4},{"task":"/task/human-activity-recognition","name":"Human Activity Recognition","papers":4},{"task":"/task/quantization","name":"Quantization","papers":4},{"task":"/task/segmentation","name":"Segmentation","papers":4},{"task":"/task/time-series-1","name":"Time Series","papers":4},{"task":"/task/disentanglement","name":"Disentanglement","papers":3},{"task":"/task/language-modeling","name":"Language Modeling","papers":3}],"tasks_shown":20,"n_tasks":137,"usage_by_year":[{"year":"2018","papers":2},{"year":"2019","papers":7},{"year":"2020","papers":20},{"year":"2021","papers":27},{"year":"2022","papers":21},{"year":"2023","papers":14},{"year":"2024","papers":16},{"year":"2025","papers":6}],"row_source":"methods_table","archive":{"source":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","archive_url":"https://paperswithcode.com/method/contrastive-predictive-coding"},"syntology_read_at":"2026-09-24T18:15:14+00:00"}