Datasets › A2D Sentences

A2D Sentences (Sentences for the Actor-Action Dataset (A2D))

Introduced by Kirill Gavrilyuk et al. in Actor and Action Video Segmentation from a Sentence20 Mar 2018 archive 2025-07-28

The Actor-Action Dataset (A2D) by Xu et al. [29] serves as the largest video dataset for the general actor and action segmentation task. It contains 3,782 videos from YouTube with pixel-level labeled actors and their actions. The dataset includes eight different actions, while a total of seven actor classes are considered to perform those actions. We follow [29], who split the dataset into 3,036 training videos and 746 testing videos.

As we are interested in pixel-level actor and action segmentation from sentences, we augment the videos in A2D with natural language descriptions about what each actor is doing in the videos. Following the guidelines set forth in [12], we ask our annotators for a discriminative referring expression of each actor instance if multiple objects are considered in a video. The annotation process resulted in a total of 6,656 sentences, including 811 different nouns, 225 verbs and 189 adjectives. Our sentences enrich the actor and action pairs from the A2D dataset with finer granularities. For example, the actor adult in A2D may be annotated with man, woman, person and player in our sentences, while action rolling may also refer to flipping, sliding, moving and running when describing different actors in different scenarios. Our sentences contain on average more words than the ReferIt dataset [12] (7.3 vs 4.7), even when we leave out prepositions, articles and linking verbs (4.5 vs 3.6). This makes sense as our sentences contain a variety of verbs while existing referring expression datasets mostly ignore verbs.

Benchmarks archive 2025-07-28

All 1 leaderboard whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

First row (archive order)PaperCode
Referring Expression Segmentation A2D Sentences SgMg (Video-Swin-B) AP 0.585 Spectrum-guided Multi-granularity Referring Video Object... bo-miao/sgmg 27 Compare

Papers archive 2025-07-28

22 shown of 22 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 31. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
Spectrum-guided Multi-granularity Referring Video Object Segmentation 1 1 25 Jul 2023 ran 6 of 9 samples (3 unverified; 9 pointer-only for licence)
SOC: Semantic-Assisted Object Cluster for Referring Video Object Segmentation 1 2 26 May 2023 not harvested
Multi-Attention Network for Compressed Video Referring Object Segmentation 1 1 26 Jul 2022 not harvested
Modeling Motion with Multi-Modal Features for Text-Based Video Segmentation 1 1 6 Apr 2022 not harvested
Deeply Interleaved Two-Stream Encoder for Referring Video Segmentation 0 1 30 Mar 2022 not harvested
Local-Global Context Aware Transformer for Language-Guided Video Segmentation 1 1 18 Mar 2022 ran 1 of 1 samples (0 unverified)
Language as Queries for Referring Video Object Segmentation 1 1 3 Jan 2022 ran 7 of 8 samples (1 unverified; 8 pointer-only for licence)
End-to-End Referring Video Object Segmentation with Multimodal Transformers 2 2 29 Nov 2021 ran 6 of 11 samples (5 unverified)
Hierarchical interaction network for video object segmentation from referring expressions 0 2 22 Nov 2021 not harvested
Cross-Modal Progressive Comprehension for Referring Segmentation 1 2 15 May 2021 not harvested
Collaborative Spatial-Temporal Modeling for Language-Queried Video Actor Segmentation 0 1 14 May 2021 not harvested
ClawCraneNet: Leveraging Object-level Relation for Text-based Video Segmentation 0 1 19 Mar 2021 not harvested
Referring Segmentation in Images and Videos with Cross-Modal Self-Attention Network 0 1 9 Feb 2021 not harvested
Actor and Action Modular Network for Text-based Video Segmentation 0 1 2 Nov 2020 not harvested
RefVOS: A Closer Look at Referring Expressions for Video Object Segmentation 2 1 1 Oct 2020 not harvested
Polar Relative Positional Encoding for Video-Language Segmentation 0 1 20 Jul 2020 not harvested
Visual-Textual Capsule Routing for Text-Based Video Segmentation 0 1 1 Jun 2020 not harvested
Context Modulated Dynamic Networks for Actor and Action Video Segmentation with Language Queries 0 1 3 Apr 2020 not harvested
Asymmetric Cross-Guided Attention Network for Actor and Action Video Segmentation From Natural Language Query 1 1 1 Oct 2019 not harvested
Actor and Action Video Segmentation from a Sentence 1 2 20 Mar 2018 not harvested
Tracking by Natural Language Specification 0 1 1 Jul 2017 not harvested
Segmentation from Natural Language Expressions 4 1 20 Mar 2016 not harvested

Dataset loaders archive 2025-07-28

1 loader as listed in the archive; links are outbound and not re-checked here.

Tasks archive 2025-07-28

License archive 2025-07-28

No licence recorded in the archive. Absence here is not a statement about the dataset's terms.

Modalities archive 2025-07-28

Languages archive 2025-07-28

Variants archive 2025-07-28

  • A2D Sentences

1 variant name, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections