Datasets › Charades

Charades

Introduced by Gunnar A. Sigurdsson et al. in Hollywood in Homes: Crowdsourcing Data Collection for Activity Understanding1 Jan 2016 archive 2025-07-28

The Charades dataset is composed of 9,848 videos of daily indoors activities with an average length of 30 seconds, involving interactions with 46 objects classes in 15 types of indoor scenes and containing a vocabulary of 30 verbs leading to 157 action classes. Each video in this dataset is annotated by multiple free-text descriptions, action labels, action intervals and classes of interacting objects. 267 different users were presented with a sentence, which includes objects and actions from a fixed vocabulary, and they recorded a video acting out the sentence. In total, the dataset contains 66,500 temporal annotations for 157 action classes, 41,104 labels for 46 object classes, and 27,847 textual descriptions of the videos. In the standard split there are7,986 training video and 1,863 validation video.

Source: Temporal Reasoning Graph for Activity Recognition

Benchmarks archive 2025-07-28

All 6 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

Papers archive 2025-07-28

30 shown of 53 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 428. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
UniMD: Towards Unifying Moment Retrieval and Temporal Action Detection 1 1 7 Apr 2024 not harvested
Towards Weakly Supervised End-to-end Learning for Long-video Action Recognition 0 4 28 Nov 2023 not harvested
PAT: Position-Aware Transformer for Dense Multi-Label Action Detection 0 1 9 Aug 2023 not harvested
Actor-agnostic Multi-label Action Recognition with Multi-modal Query 1 2 20 Jul 2023 not harvested
VicTR: Video-conditioned Text Representations for Activity Recognition 0 1 5 Apr 2023 not harvested
MAtch, eXpand and Improve: Unsupervised Finetuning for Zero-Shot Action Recognition with Language Knowledge 1 1 15 Mar 2023 ran 3 of 4 samples (1 unverified; 4 pointer-only for licence)
Bidirectional Cross-Modal Knowledge Exploration for Video Recognition with Pre-trained Vision-Language Models 5 1 31 Dec 2022 not harvested
VideoCoCa: Video-Text Modeling with Zero-Shot Transfer from Contrastive Captioners 0 1 9 Dec 2022 not harvested
Rethinking Video ViTs: Sparse Video Tubes for Joint Image and Video Learning 1 1 6 Dec 2022 ran 1 of 3 samples (2 unverified)
Token Turing Machines 1 1 16 Nov 2022 not harvested
A CLIP-Hitchhiker's Guide to Long Video Retrieval 1 1 17 May 2022 not harvested
MS-TCT: Multi-Scale Temporal ConvTransformer for Action Detection 1 1 7 Dec 2021 ran 1 of 1 samples (0 unverified; 1 pointer-only for licence)
Weakly-guided Self-supervised Pretraining for Temporal Activity Detection 1 1 26 Nov 2021 ran 2 of 11 samples (9 unverified)
Revisiting spatio-temporal layouts for compositional action recognition 1 1 2 Nov 2021 not harvested
CTRN: Class-Temporal Relational Network for Action Detection 0 1 26 Oct 2021 not harvested
ActionCLIP: A New Paradigm for Video Action Recognition 2 1 17 Sep 2021 ran 5 of 9 samples (4 unverified)
TokenLearner: What Can 8 Learned Tokens Do for Images and Videos? 11 1 21 Jun 2021 ran 3 of 3 samples (0 unverified)
Continual 3D Convolutional Neural Networks for Real-time Processing of Videos 1 3 31 May 2021 not harvested
VidTr: Video Transformer Without Convolutions 0 2 23 Apr 2021 not harvested
Multiscale Vision Transformers 8 6 22 Apr 2021 ran 13 of 26 samples (13 unverified; 5 pointer-only for licence)
MoViNets: Mobile Video Networks for Efficient Video Recognition 3 3 21 Mar 2021 ran 8 of 13 samples (5 unverified)
Modeling Multi-Label Action Dependencies for Temporal Action Localization 1 1 4 Mar 2021 ran 5 of 5 samples (0 unverified; 5 pointer-only for licence)
Coarse-Fine Networks for Temporal Activity Detection in Videos 1 1 1 Mar 2021 not harvested
PDAN: Pyramid Dilated Attention Network for Action Detection 1 1 5 Jan 2021 not harvested
Pose And Joint-Aware Action Recognition 1 2 16 Oct 2020 not harvested
AssembleNet++: Assembling Modality Representations via Attention Connections 1 2 18 Aug 2020 not harvested
AViD Dataset: Anonymized Videos from Diverse Countries 1 2 10 Jul 2020 not harvested
Self-supervising Action Recognition by Statistical Moment and Subspace Descriptors 0 2 14 Jan 2020 not harvested
A Multigrid Method for Efficiently Training Video Models 3 1 2 Dec 2019 ran 2 of 10 samples (8 unverified; 3 pointer-only for licence)
Hallucinating IDT Descriptors and I3D Optical Flow Features for Action Recognition with CNNs 0 1 13 Jun 2019 not harvested

The full list of 53 is in the JSON twin.

Dataset loaders archive 2025-07-28

2 loaders as listed in the archive; links are outbound and not re-checked here.

Tasks archive 2025-07-28

License archive 2025-07-28

Custom (non-commercial)

Modalities archive 2025-07-28

Languages archive 2025-07-28

No language tagged.

Variants archive 2025-07-28

  • Charades

1 variant name, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections