Methods › Computer Vision › Generative Video Models › TimeSformer
TimeSformer
Introduced by Gedas Bertasius et al. in Is Space-Time Attention All You Need for Video Understanding?
archive 2025-07-28 Description, source and code snippet are the archive's method entry.
TimeSformer is a convolution-free approach to video classification built exclusively on self-attention over space and time. It adapts the standard Transformer architecture to video by enabling spatiotemporal feature learning directly from a sequence of frame-level patches. Specifically, the method adapts the image model [Vision Transformer](https://paperswithcode.com/method/vision-transformer) (ViT) to video by extending the self-attention mechanism from the image space to the space-time 3D volume. As in ViT, each patch is linearly mapped into an embedding and augmented with positional information. This makes it possible to interpret the resulting sequence of vector
Papers archive 2025-07-28
18 shown of 18, newest first. Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code; it is per sample and not a correctness claim.
-
DualX-VSR: Dual Axial Spatial×Temporal Transformer for Real-World Video Super-Resolution without Motion Compensation 5 Jun 2025 · 0 repositories · arXiv:2506.04830
-
Fine-Tuning Video Transformers for Word-Level Bangla Sign Language: A Comparative Analysis for Classification Tasks 4 Jun 2025 · 0 repositories · arXiv:2506.04367
-
SkillFormer: Unified Multi-View Video Understanding for Proficiency Estimation 13 May 2025 · 0 repositories · arXiv:2505.08665
-
FedDA-TSformer: Federated Domain Adaptation with Vision TimeSformer for Left Ventricle Segmentation on Gated Myocardial Perfusion SPECT Image 23 Feb 2025 · 0 repositories · arXiv:2502.16709
-
EITNet: An IoT-Enhanced Framework for Real-Time Basketball Action Recognition 13 Oct 2024 · 0 repositories · arXiv:2410.09954
-
3D-LSPTM: An Automatic Framework with 3D-Large-Scale Pretrained Model for Laryngeal Cancer Detection Using Laryngoscopic Videos 2 Sep 2024 · 0 repositories · arXiv:2409.01459
-
Motion meets Attention: Video Motion Prompts 3 Jul 2024 · 1 repository · arXiv:2407.03179Syntology ran 9 of 11 samples · 2 unverified
-
Pig aggression classification using CNN, Transformers and Recurrent Networks 13 Mar 2024 · 0 repositories · arXiv:2403.08528
-
P-Age: Pexels Dataset for Robust Spatio-Temporal Apparent Age Classification 4 Nov 2023 · 0 repositories · arXiv:2311.02432
-
Scalable and Accurate Self-supervised Multimodal Representation Learning without Aligned Video and Text Data 4 Apr 2023 · 0 repositories · arXiv:2304.02080
-
Video Question Answering Using CLIP-Guided Visual-Text Attention 6 Mar 2023 · 0 repositories · arXiv:2303.03131
-
CholecTriplet2022: Show me a tool and tell me the triplet -- an endoscopic vision challenge for surgical action triplet detection 13 Feb 2023 · 2 repositories · arXiv:2302.06294
-
MINTIME: Multi-Identity Size-Invariant Video Deepfake Detection 20 Nov 2022 · 1 repository · arXiv:2211.10996
-
One Model is Not Enough: Ensembles for Isolated Sign Language Recognition 4 Jul 2022 · 0 repositories
-
Context-aware Proposal Network for Temporal Action Detection 18 Jun 2022 · 0 repositories · arXiv:2206.09082
-
VIDI: A Video Dataset of Incidents 26 May 2022 · 1 repository · arXiv:2205.13277
-
IA-RED²: Interpretability-Aware Redundancy Reduction for Vision Transformers 23 Jun 2021 · 0 repositories · arXiv:2106.12620
-
Is Space-Time Attention All You Need for Video Understanding? 9 Feb 2021 · 16 repositories · arXiv:2102.05095Syntology ran 35 of 43 samples · 8 unverified · 14 pointer-only (licence)
Tasks archive 2025-07-28
20 shown of 46 tasks the archive attaches to papers tagged with this method, by distinct papers. A task without a page in the catalog is plain text.
Usage over time archive 2025-07-28
Components: the archive holds no method-to-method composition, so PwC's Components table cannot be rebuilt; the Papers list carries no Results column for the same reason (the archive does not join its leaderboard rows to method tags).
Categories archive 2025-07-28
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections