{"url":"/method/tam","slug":"tam","name":"TAM","full_name":"Temporal Adaptive Module","full_name_withheld":false,"description_markdown":"TAM is designed to capture complex temporal relationships both  efficiently and  flexibly,\r\nIt adopts an adaptive kernel instead of self-attention to capture  global contextual information, with lower time complexity \r\nthan GLTR.\r\n\r\nTAM has two branches, a local branch and a global branch. Given the input feature map $X\\in \\mathbb{R}^{C\\times T\\times H\\times W}$,  global spatial average pooling $\\text{GAP}$ is first applied to the feature map to ensure TAM has a low computational cost. Then the local branch in TAM employs several 1D convolutions with ReLU nonlinearity across the temporal domain to produce location-sensitive importance maps for enhancing frame-wise features.\r\nThe local branch can be written as\r\n\\begin{align}\r\n    s &= \\sigma(\\text{Conv1D}(\\delta(\\text{Conv1D}(\\text{GAP}(X)))))\r\n\\end{align}\r\n\\begin{align}\r\n    X^1 &= s X\r\n\\end{align}\r\nUnlike the local branch, the global branch is location invariant and focuses on generating a channel-wise adaptive kernel based on global temporal information in each channel. For the $c$-th channel, the  kernel can be written as\r\n\r\n\\begin{align}\r\n    \\Theta_c = \\text{Softmax}(\\text{FC}_2(\\delta(\\text{FC}_1(\\text{GAP}(X)_c)))) \r\n\\end{align}\r\n\r\nwhere $\\Theta_c \\in \\mathbb{R}^{K}$ and $K$ is the adaptive kernel size. Finally, TAM  convolves the adaptive kernel $\\Theta$ with $ X_\\text{out}^1$:\r\n\\begin{align}\r\n    Y = \\Theta \\otimes  X^1\r\n\\end{align}\r\n\r\nWith the help of the local branch and global branch,\r\nTAM can capture the complex temporal structures in video and \r\nenhance per-frame features at low computational cost.\r\nDue to its flexibility and lightweight design,\r\nTAM can be added to any existing 2D CNNs.","description_state":"present","introduced_year":null,"introduced_by":{"title":"TAM: Temporal Adaptive Module for Video Recognition","paper":"/paper/tam-temporal-adaptive-module-for-video","first_author":"Zhao-Yang Liu","n_authors":5,"url_abs":null,"archive_paper_url":"https://paperswithcode.com/paper/tam-temporal-adaptive-module-for-video"},"source":{"url":"https://arxiv.org/abs/2005.06803v3","title":"TAM: Temporal Adaptive Module for Video Recognition","url_on_a_paper_host":true},"code_snippet_url":null,"code_snippet_url_on_a_code_host":false,"categories":[{"area":"General","area_id":"general","collection":"Attention Mechanisms","url":"/methods/category/attention-mechanisms","pwc_aliases":["attention-mechanisms-1"]}],"n_papers_tagged":32,"archive_num_papers":32,"papers_newest_first":[{"paper":"/paper/topology-aware-modeling-for-unsupervised","title":"Topology-Aware Modeling for Unsupervised Simulation-to-Reality Point Cloud Recognition","date":"2025-06-26","arxiv_id":"2506.21165","n_code_links":1,"syntology":null},{"paper":null,"title":"Motion-enhancement to Echocardiography Segmentation via Inserting a Temporal Attention Module: An Efficient, Adaptable, and Scalable Approach","date":"2025-01-24","arxiv_id":"2501.14929","n_code_links":0,"syntology":null},{"paper":null,"title":"Threshold Attention Network for Semantic Segmentation of Remote Sensing Images","date":"2025-01-14","arxiv_id":"2501.07984","n_code_links":0,"syntology":null},{"paper":null,"title":"Torque-Aware Momentum","date":"2024-12-25","arxiv_id":"2412.18790","n_code_links":0,"syntology":null},{"paper":null,"title":"EntityCLIP: Entity-Centric Image-Text Matching via Multimodal Attentive Contrastive Learning","date":"2024-10-23","arxiv_id":"2410.17810","n_code_links":0,"syntology":null},{"paper":null,"title":"Parameter Estimation in Optimal Tolling for Traffic Networks Under the Markovian Traffic Equilibrium","date":"2024-09-29","arxiv_id":"2409.19765","n_code_links":0,"syntology":null},{"paper":"/paper/user-story-tutor-ust-to-support-agile","title":"User Story Tutor (UST) to Support Agile Software Developers","date":"2024-06-24","arxiv_id":"2406.16259","n_code_links":2,"syntology":null},{"paper":null,"title":"Digital Health and Indoor Air Quality: An IoT-Driven Human-Centred Visualisation Platform for Behavioural Change and Technology Acceptance","date":"2024-05-20","arxiv_id":"2405.13064","n_code_links":0,"syntology":null},{"paper":null,"title":"Enhancing Multivariate Time Series Forecasting with Mutual Information-driven Cross-Variable and Temporal Modeling","date":"2024-03-01","arxiv_id":"2403.00869","n_code_links":0,"syntology":null},{"paper":null,"title":"Minimizing Energy Consumption in MU-MIMO via Antenna Muting by Neural Networks with Asymmetric Loss","date":"2023-06-08","arxiv_id":"2306.05162","n_code_links":0,"syntology":null},{"paper":"/paper/truncated-affinity-maximization-one-class-1","title":"Truncated Affinity Maximization: One-class Homophily Modeling for Graph Anomaly Detection","date":"2023-05-29","arxiv_id":"2306.00006","n_code_links":1,"syntology":{"ran":1,"of":3,"unverified":2,"pointer_only":3}},{"paper":null,"title":"Improve Video Representation with Temporal Adversarial Augmentation","date":"2023-04-28","arxiv_id":"2304.14601","n_code_links":0,"syntology":null},{"paper":null,"title":"Arc-based Traffic Assignment: Equilibrium Characterization and Learning","date":"2023-04-10","arxiv_id":"2304.04705","n_code_links":0,"syntology":null},{"paper":"/paper/spectral-gap-regularization-of-neural","title":"Spectral Gap Regularization of Neural Networks","date":"2023-04-06","arxiv_id":"2304.03096","n_code_links":0,"syntology":{"ran":0,"of":1,"unverified":1,"pointer_only":1}},{"paper":"/paper/fgahoi-fine-grained-anchors-for-human-object","title":"FGAHOI: Fine-Grained Anchors for Human-Object Interaction Detection","date":"2023-01-08","arxiv_id":"2301.04019","n_code_links":1,"syntology":null},{"paper":null,"title":"MangngalApp -- An integrated package of technology for COVID-19 response and rural development: Acceptability and usability using TAM","date":"2023-01-07","arxiv_id":"2301.02893","n_code_links":0,"syntology":null},{"paper":null,"title":"Consumer acceptance of the use of artificial intelligence in online shopping: evidence from Hungary","date":"2022-12-26","arxiv_id":"2301.01277","n_code_links":0,"syntology":null},{"paper":"/paper/domain-alignment-and-temporal-aggregation-for","title":"Dual Prototype Attention for Unsupervised Video Object Segmentation","date":"2022-11-22","arxiv_id":"2211.12036","n_code_links":1,"syntology":null},{"paper":null,"title":"Trust in AI and Its Role in the Acceptance of AI Technologies","date":"2022-03-23","arxiv_id":"2203.12687","n_code_links":0,"syntology":null},{"paper":"/paper/slow-fast-visual-tempo-learning-for-video","title":"Motion-driven Visual Tempo Learning for Video-based Action Recognition","date":"2022-02-24","arxiv_id":"2202.12116","n_code_links":2,"syntology":null},{"paper":null,"title":"A Tech Hybrid-Recommendation Engine and Personalized Notification: An integrated tool to assist users through Recommendations (Project ATHENA)","date":"2022-02-13","arxiv_id":"2202.06248","n_code_links":0,"syntology":null},{"paper":null,"title":"CPSeg: Cluster-free Panoptic Segmentation of 3D LiDAR Point Clouds","date":"2021-11-02","arxiv_id":"2111.01723","n_code_links":0,"syntology":null},{"paper":null,"title":"ACFNet: Adaptively-Cooperative Fusion Network for RGB-D Salient Object Detection","date":"2021-09-10","arxiv_id":"2109.04627","n_code_links":0,"syntology":null},{"paper":null,"title":"Sequential Attention Module for Natural Language Processing","date":"2021-09-07","arxiv_id":"2109.03009","n_code_links":0,"syntology":null},{"paper":"/paper/tvt-transferable-vision-transformer-for","title":"TVT: Transferable Vision Transformer for Unsupervised Domain Adaptation","date":"2021-08-12","arxiv_id":"2108.05988","n_code_links":1,"syntology":{"ran":2,"of":2,"unverified":0,"pointer_only":0}},{"paper":null,"title":"A Trigger-Aware Multi-Task Learning for Chinese Event Entity Recognition","date":"2021-08-01","arxiv_id":null,"n_code_links":0,"syntology":null},{"paper":null,"title":"Genre determining prediction: Non-standard TAM marking in football language","date":"2021-06-30","arxiv_id":"2106.15872","n_code_links":0,"syntology":null},{"paper":null,"title":"Helping users discover perspectives: Enhancing opinion mining with joint topic models","date":"2020-10-23","arxiv_id":"2010.12505","n_code_links":0,"syntology":null},{"paper":null,"title":"Temporal Attention-Augmented Graph Convolutional Network for Efficient Skeleton-Based Human Action Recognition","date":"2020-10-23","arxiv_id":"2010.12221","n_code_links":0,"syntology":null},{"paper":null,"title":"Tense, aspect and mood based event extraction for situation analysis and crisis management","date":"2020-08-01","arxiv_id":"2008.01555","n_code_links":0,"syntology":null}],"papers_shown":30,"tasks":[{"task":"/task/arc","name":"ARC","papers":3},{"task":"/task/action-recognition-in-videos","name":"Action Recognition","papers":3},{"task":"/task/semantic-segmentation","name":"Semantic Segmentation","papers":3},{"task":"/task/contrastive-learning","name":"Contrastive Learning","papers":2},{"task":"/task/domain-adaptation","name":"Domain Adaptation","papers":2},{"task":"/task/language-modelling","name":"Language Modelling","papers":2},{"task":"/task/object","name":"Object","papers":2},{"task":"/task/segmentation","name":"Segmentation","papers":2},{"task":"/task/unsupervised-domain-adaptation","name":"Unsupervised Domain Adaptation","papers":2},{"task":"/task/anatomy","name":"Anatomy","papers":1},{"task":"/task/anomaly-detection","name":"Anomaly Detection","papers":1},{"task":"/task/autonomous-driving","name":"Autonomous Driving","papers":1},{"task":"/task/classification-1","name":"Classification","papers":1},{"task":"/task/clustering","name":"Clustering","papers":1},{"task":"/task/collaborative-filtering","name":"Collaborative Filtering","papers":1},{"task":"/task/combinatorial-optimization","name":"Combinatorial Optimization","papers":1},{"task":"/task/culture","name":"Cultural Vocal Bursts Intensity Prediction","papers":1},{"task":"/task/decoder","name":"Decoder","papers":1},{"task":"/task/depth-completion","name":"Depth Completion","papers":1},{"task":"/task/descriptive","name":"Descriptive","papers":1}],"tasks_shown":20,"n_tasks":60,"usage_by_year":[{"year":"2020","papers":5},{"year":"2021","papers":6},{"year":"2022","papers":5},{"year":"2023","papers":7},{"year":"2024","papers":6},{"year":"2025","papers":3}],"row_source":"methods_table","archive":{"source":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","archive_url":"https://paperswithcode.com/method/tam"},"syntology_read_at":"2026-09-24T18:15:14+00:00"}