Datasets › Youtube INRIA Instructional
Youtube INRIA Instructional (Unsupervised learning from narrated instruction videos)
We address the problem of automatically learning the main steps to complete a certain task, such as changing a car tire, from a set of narrated instruction videos. The contributions of this paper are three-fold. First, we develop a new unsupervised learning approach that takes advantage of the complementary nature of the input video and the associated narration. The method solves two clustering problems, one in text and one in video, applied one after each other and linked by joint constraints to obtain a single coherent sequence of steps in both modalities. Second, we collect and annotate a new challenging dataset of real-world instruction videos from the Internet. The dataset contains about 800,000 frames for five different tasks (How to : change a car tire, perform CardioPulmonary resuscitation (CPR), jump cars, repot a plant and make coffee) that include complex interactions between people and objects, and are captured in a variety of indoor and outdoor settings. Third, we experimentally demonstrate that the proposed method can automatically discover, in an unsupervised manner , the main steps to achieve the task and locate the steps in the input videos.
This video presents our results of automatically discovering the scenario for the two following task : changing a tire and performing CardioPulmonary Resuscitation (CPR). At the bottom of the videos, there are three bars. The first one corresponds to our ground truth annotation. The second one corresponds to our time interval prediction in video. Finally the third one corresponds to the constraints that we obtain from the text domain. On the right, there is a list of label. They corresponds to the label recovered by our NLP method in an unsupervised manner.
Benchmarks archive 2025-07-28
All 2 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.
| First row (archive order) | Paper | Code | ||||
|---|---|---|---|---|---|---|
| Unsupervised Action Segmentation | Youtube INRIA Instructional | LSTM+AL F1 39.7 | A Perceptual Prediction Framework for Self Supervised... | CVPRUSFTampa/EventSegmentation | 8 | Compare |
| Action Segmentation | Youtube INRIA Instructional | TSA (FINCH) Acc 62.4 | Leveraging triplet loss for unsupervised action segmentation | elenabbbuenob/tsa-actionseg | 2 | Compare |
Papers archive 2025-07-28
8 shown of 8 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 8. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.
| Date | Samples run Syntology | |||
|---|---|---|---|---|
| Hierarchical Vector Quantization for Unsupervised Action Segmentation | 1 | 1 | 23 Dec 2024 | not harvested |
| Temporally Consistent Unbalanced Optimal Transport for Unsupervised Action Segmentation | 1 | 1 | 1 Apr 2024 | not harvested |
| Permutation-Aware Action Segmentation via Unsupervised Frame-to-Segment Alignment | 1 | 1 | 31 May 2023 | not harvested |
| Leveraging triplet loss for unsupervised action segmentation | 1 | 2 | 13 Apr 2023 | not harvested |
| Unsupervised Action Segmentation by Joint Representation Learning and Online Clustering | 1 | 2 | 27 May 2021 | ran 0 of 2 samples (2 unverified) |
| Action Shuffle Alternating Learning for Unsupervised Action Segmentation | 0 | 1 | 5 Apr 2021 | not harvested |
| Unsupervised learning of action classes with continuous temporal embedding | 2 | 1 | 8 Apr 2019 | not harvested |
| A Perceptual Prediction Framework for Self Supervised Event Segmentation | 1 | 1 | 12 Nov 2018 | ran 0 of 1 samples (1 unverified) |
Dataset loaders archive 2025-07-28
No loader listed in the archive.
Tasks archive 2025-07-28
License archive 2025-07-28
No licence recorded in the archive. Absence here is not a statement about the dataset's terms.
Modalities archive 2025-07-28
No modality tagged.
Languages archive 2025-07-28
Variants archive 2025-07-28
- Youtube INRIA Instructional
1 variant name, as the archive lists them.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections