{"url":"/method/residual-connection","slug":"residual-connection","name":"Residual Connection","full_name":"Residual Connection","full_name_withheld":false,"description_markdown":"**Residual Connections** are a type of skip-connection that learn residual functions with reference to the layer inputs, instead of learning unreferenced functions. \r\n\r\nFormally, denoting the desired underlying mapping as $\\mathcal{H}({x})$, we let the stacked nonlinear layers fit another mapping of $\\mathcal{F}({x}):=\\mathcal{H}({x})-{x}$. The original mapping is recast into $\\mathcal{F}({x})+{x}$.\r\n\r\nThe intuition is that it is easier to optimize the residual mapping than to optimize the original, unreferenced mapping. To the extreme, if an identity mapping were optimal, it would be easier to push the residual to zero than to fit an identity mapping by a stack of nonlinear layers.","description_state":"present","introduced_year":null,"introduced_by":{"title":null,"paper":null,"first_author":null,"n_authors":0,"url_abs":null,"archive_paper_url":null},"source":{"url":"http://arxiv.org/abs/1512.03385v1","title":"Deep Residual Learning for Image Recognition","url_on_a_paper_host":true},"code_snippet_url":"https://github.com/pytorch/vision/blob/7c077f6a986f05383bcb86b535aedb5a63dd5c4b/torchvision/models/resnet.py#L118","code_snippet_url_on_a_code_host":true,"categories":[{"area":"General","area_id":"general","collection":"Skip Connections","url":"/methods/category/skip-connections","pwc_aliases":[]}],"n_papers_tagged":28401,"archive_num_papers":null,"papers_newest_first":[{"paper":"/paper/constructing-and-evaluating-declarative-rag","title":"Constructing and Evaluating Declarative RAG Pipelines in PyTerrier","date":"2025-06-12","arxiv_id":"2506.10802","n_code_links":1,"syntology":null},{"paper":"/paper/towards-robust-multimodal-emotion-recognition","title":"Towards Robust Multimodal Emotion Recognition under Missing Modalities and Distribution Shifts","date":"2025-06-12","arxiv_id":"2506.10452","n_code_links":1,"syntology":null},{"paper":null,"title":"A Novel Lightweight Transformer with Edge-Aware Fusion for Remote Sensing Image Captioning","date":"2025-06-11","arxiv_id":"2506.09429","n_code_links":0,"syntology":null},{"paper":null,"title":"Auto-Compressing Networks","date":"2025-06-11","arxiv_id":"2506.09714","n_code_links":0,"syntology":null},{"paper":"/paper/learning-efficient-and-generalizable-graph","title":"Learning Efficient and Generalizable Graph Retriever for Knowledge-Graph Question Answering","date":"2025-06-11","arxiv_id":"2506.09645","n_code_links":1,"syntology":null},{"paper":null,"title":"SparseSSM: Efficient Selective Structured State Space Models Can Be Pruned in One-Shot","date":"2025-06-11","arxiv_id":"2506.09613","n_code_links":0,"syntology":null},{"paper":"/paper/2506-08884","title":"InfoDPCCA: Information-Theoretic Dynamic Probabilistic Canonical Correlation Analysis","date":"2025-06-10","arxiv_id":"2506.08884","n_code_links":1,"syntology":null},{"paper":null,"title":"Fine-Grained Spatially Varying Material Selection in Images","date":"2025-06-10","arxiv_id":"2506.09023","n_code_links":0,"syntology":null},{"paper":null,"title":"Hierarchical Neural Collapse Detection Transformer for Class Incremental Object Detection","date":"2025-06-10","arxiv_id":"2506.08562","n_code_links":0,"syntology":null},{"paper":"/paper/hyperspectral-image-classification-via","title":"Hyperspectral Image Classification via Transformer-based Spectral-Spatial Attention Decoupling and Adaptive Gating","date":"2025-06-10","arxiv_id":"2506.08324","n_code_links":1,"syntology":null},{"paper":null,"title":"MedMoE: Modality-Specialized Mixture of Experts for Medical Vision-Language Understanding","date":"2025-06-10","arxiv_id":"2506.08356","n_code_links":0,"syntology":null},{"paper":"/paper/patchguard-adversarially-robust-anomaly-1","title":"PatchGuard: Adversarially Robust Anomaly Detection and Localization through Vision Transformers and Pseudo Anomalies","date":"2025-06-10","arxiv_id":"2506.09237","n_code_links":2,"syntology":null},{"paper":null,"title":"Robust Visual Localization via Semantic-Guided Multi-Scale Transformer","date":"2025-06-10","arxiv_id":"2506.08526","n_code_links":0,"syntology":null},{"paper":"/paper/tactic-translation-agents-with-cognitive","title":"TACTIC: Translation Agents with Cognitive-Theoretic Interactive Collaboration","date":"2025-06-10","arxiv_id":"2506.08403","n_code_links":1,"syntology":null},{"paper":null,"title":"4DGT: Learning a 4D Gaussian Transformer Using Real-World Monocular Videos","date":"2025-06-09","arxiv_id":"2506.08015","n_code_links":0,"syntology":null},{"paper":null,"title":"Lightweight Sequential Transformers for Blood Glucose Level Prediction in Type-1 Diabetes","date":"2025-06-09","arxiv_id":"2506.07864","n_code_links":0,"syntology":null},{"paper":"/paper/llamarec-lkg-rag-a-single-pass-learnable","title":"LlamaRec-LKG-RAG: A Single-Pass, Learnable Knowledge Graph-RAG Framework for LLM-Based Ranking","date":"2025-06-09","arxiv_id":"2506.07449","n_code_links":1,"syntology":null},{"paper":null,"title":"LLM-driven Indoor Scene Layout Generation via Scaled Human-aligned Data Synthesis and Multi-Stage Preference Optimization","date":"2025-06-09","arxiv_id":"2506.07570","n_code_links":0,"syntology":null},{"paper":null,"title":"MADFormer: Mixed Autoregressive and Diffusion Transformers for Continuous Image Generation","date":"2025-06-09","arxiv_id":"2506.07999","n_code_links":0,"syntology":null},{"paper":null,"title":"Quantum Graph Transformer for NLP Sentiment Classification","date":"2025-06-09","arxiv_id":"2506.07937","n_code_links":0,"syntology":null},{"paper":null,"title":"SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding","date":"2025-06-09","arxiv_id":"2506.07600","n_code_links":0,"syntology":null},{"paper":null,"title":"SpikeSMOKE: Spiking Neural Networks for Monocular 3D Object Detection with Cross-Scale Gated Coding","date":"2025-06-09","arxiv_id":"2506.07737","n_code_links":0,"syntology":null},{"paper":null,"title":"Quality-Diversity Red-Teaming: Automated Generation of High-Quality and Diverse Attackers for Large Language Models","date":"2025-06-08","arxiv_id":"2506.07121","n_code_links":0,"syntology":null},{"paper":null,"title":"Breaking Data Silos: Towards Open and Scalable Mobility Foundation Models via Generative Continual Learning","date":"2025-06-07","arxiv_id":"2506.06694","n_code_links":0,"syntology":null},{"paper":null,"title":"Can In-Context Reinforcement Learning Recover From Reward Poisoning Attacks?","date":"2025-06-07","arxiv_id":"2506.06891","n_code_links":0,"syntology":null},{"paper":null,"title":"BEAST: Efficient Tokenization of B-Splines Encoded Action Sequences for Imitation Learning","date":"2025-06-06","arxiv_id":"2506.06072","n_code_links":0,"syntology":null},{"paper":null,"title":"Direct Behavior Optimization: Unlocking the Potential of Lightweight LLMs","date":"2025-06-06","arxiv_id":"2506.06401","n_code_links":0,"syntology":null},{"paper":null,"title":"Textile Analysis for Recycling Automation using Transfer Learning and Zero-Shot Foundation Models","date":"2025-06-06","arxiv_id":"2506.06569","n_code_links":0,"syntology":null},{"paper":"/paper/when-to-use-graphs-in-rag-a-comprehensive","title":"When to use Graphs in RAG: A Comprehensive Analysis for Graph Retrieval-Augmented Generation","date":"2025-06-06","arxiv_id":"2506.05690","n_code_links":1,"syntology":{"ran":0,"of":5,"unverified":5,"pointer_only":0}},{"paper":null,"title":"A Multi-Dataset Evaluation of Models for Automated Vulnerability Repair","date":"2025-06-05","arxiv_id":"2506.04987","n_code_links":0,"syntology":null}],"papers_shown":30,"tasks":[{"task":"/task/language-modelling","name":"Language Modelling","papers":3004},{"task":"/task/language-modeling","name":"Language Modeling","papers":2347},{"task":"/task/retrieval","name":"Retrieval","papers":1822},{"task":"/task/question-answering","name":"Question Answering","papers":1485},{"task":"/task/decoder","name":"Decoder","papers":1483},{"task":"/task/sentence","name":"Sentence","papers":1340},{"task":"/task/rag","name":"RAG","papers":1293},{"task":"/task/semantic-segmentation","name":"Semantic Segmentation","papers":1199},{"task":"/task/translation","name":"Translation","papers":1159},{"task":"/task/retrieval-augmented-generation","name":"Retrieval-augmented Generation","papers":1126},{"task":"/task/image-classification","name":"Image Classification","papers":1121},{"task":"/task/object-detection","name":"Object Detection","papers":1073},{"task":"/task/object-detection-1","name":"object-detection","papers":955},{"task":"/task/transfer-learning","name":"Transfer Learning","papers":933},{"task":"/task/machine-translation","name":"Machine Translation","papers":932},{"task":"/task/image-classification","name":"image-classification","papers":872},{"task":"/task/large-language-model","name":"Large Language Model","papers":834},{"task":"/task/segmentation","name":"Segmentation","papers":822},{"task":"/task/classification-1","name":"Classification","papers":784},{"task":"/task/representation-learning","name":"Representation Learning","papers":767}],"tasks_shown":20,"n_tasks":2748,"usage_by_year":[{"year":"2015","papers":1},{"year":"2016","papers":46},{"year":"2017","papers":160},{"year":"2018","papers":414},{"year":"2019","papers":1556},{"year":"2020","papers":2835},{"year":"2021","papers":3816},{"year":"2022","papers":3936},{"year":"2023","papers":5815},{"year":"2024","papers":7243},{"year":"2025","papers":2579}],"row_source":"embedded","archive":{"source":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","archive_url":"https://paperswithcode.com/method/residual-connection"},"syntology_read_at":"2026-09-24T18:15:14+00:00"}