Datasets › COIN

COIN

Introduced by Yansong Tang et al. in COIN: A Large-scale Dataset for Comprehensive Instructional Video Analysis archive 2025-07-28

The COIN dataset (a large-scale dataset for COmprehensive INstructional video analysis) consists of 11,827 videos related to 180 different tasks in 12 domains (e.g., vehicles, gadgets, etc.) related to our daily life. The videos are all collected from YouTube. The average length of a video is 2.36 minutes. Each video is labelled with 3.91 step segments, where each segment lasts 14.91 seconds on average. In total, the dataset contains videos of 476 hours, with 46,354 annotated segments.

Source: COIN: A Large-scale Dataset for Comprehensive Instructional Video Analysis

Benchmarks archive 2025-07-28

All 2 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

Papers archive 2025-07-28

15 shown of 15 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 105. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
HERMES: temporal-coHERent long-forM understanding with Episodes and Semantics 1 1 30 Aug 2024 ran 7 of 9 samples (2 unverified)
MA-LMM: Memory-Augmented Large Multimodal Model for Long-Term Video Understanding 1 1 8 Apr 2024 ran 7 of 9 samples (2 unverified)
Multi-granularity Correspondence Learning from Long-term Noisy Videos 1 1 30 Jan 2024 ran 2 of 3 samples (1 unverified)
UnLoc: A Unified Framework for Video Localization Tasks 1 1 21 Aug 2023 not harvested
Selective Structured State-Spaces for Long-Form Video Understanding 0 1 25 Mar 2023 not harvested
Efficient Movie Scene Detection using State-Space Transformers 1 1 29 Dec 2022 ran 3 of 3 samples (0 unverified)
Long Movie Clip Classification with State-Space Video Models 1 1 4 Apr 2022 ran 11 of 18 samples (7 unverified; 2 pointer-only for licence)
Learning To Recognize Procedural Activities with Distant Supervision 1 1 26 Jan 2022 not harvested
VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding 2 1 28 Sep 2021 not harvested
TACo: Token-aware Cascade Contrastive Learning for Video-Text Alignment 0 1 23 Aug 2021 not harvested
VLM: Task-agnostic Video-Language Model Pre-training for Video Understanding 1 1 20 May 2021 not harvested
ActBERT: Learning Global-Local Video-Text Representations 1 1 14 Nov 2020 not harvested
UniVL: A Unified Video and Language Pre-Training Model for Multimodal Understanding and Generation 2 1 15 Feb 2020 ran 1 of 4 samples (3 unverified; 1 pointer-only for licence)
End-to-End Learning of Visual Representations from Uncurated Instructional Videos 4 2 13 Dec 2019 ran 3 of 5 samples (2 unverified)
Temporal Segment Networks for Action Recognition in Videos 11 1 8 May 2017 not harvested

Dataset loaders archive 2025-07-28

No loader listed in the archive.

Tasks archive 2025-07-28

License archive 2025-07-28

Custom

Modalities archive 2025-07-28

Languages archive 2025-07-28

No language tagged.

Variants archive 2025-07-28

  • COIN

1 variant name, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections