Papers › Shuffle and Learn: Unsupervised Learning using Temporal Order Verification
Shuffle and Learn: Unsupervised Learning using Temporal Order Verification
Ishan Misra, C. Lawrence Zitnick, Martial Hebert
In this paper, we present an approach for learning a visual representation from the raw spatiotemporal signals in videos. Our representation is learned without supervision from semantic labels. We formulate our method as an unsupervised sequential verification task, i.e., we determine whether a sequence of frames from a video is in the correct temporal order. With this simple task and no semantic labels, we learn a powerful visual representation using a Convolutional Neural Network (CNN). The representation contains complementary information to that learned from supervised image datasets like ImageNet. Qualitative results show that our method captures information that is temporally varying, such as human pose. When used as pre-training for action recognition, our method gives significant gains over learning without external data on benchmark datasets like UCF101 and HMDB51. To demonstrate its sensitivity to human pose, we show results for pose estimation on the FLIC and MPII datasets that are competitive, or better than approaches using significantly more supervision. Our method can be combined with supervised representations to provide an additional boost in accuracy.
Code
No code repository is listed for this paper in the archive or in Syntology's graph.
Code Syntology ran Syntology
Not run by Syntology. Nothing on this page verifies that the listed code works.
Tasks
Results from the paper archive 2025-07-28
| Task | Dataset | Model | Metric | Value | Rank at snapshot | Leaderboard | Report |
|---|---|---|---|---|---|---|---|
| Self-Supervised Action Recognition | HMDB51 | Shuffle and Learn (AlexNet) | Frozen | false | #48 of 48 | Archive leaderboard | report |
| Self-Supervised Action Recognition | HMDB51 | Shuffle and Learn (AlexNet) | Pre-Training Dataset | UCF101 | #48 of 48 | Archive leaderboard | report |
| Self-Supervised Action Recognition | HMDB51 | Shuffle and Learn (AlexNet) | Top-1 Accuracy | 19.8 | #48 of 48 | Archive leaderboard | report |
| Self-Supervised Action Recognition | UCF101 | Shuffle and Learn (AlexNet) | 3-fold Accuracy | 50.9 | #52 of 53 | Archive leaderboard | report |
| Self-Supervised Action Recognition | UCF101 | Shuffle and Learn (AlexNet) | Frozen | false | #52 of 53 | Archive leaderboard | report |
| Self-Supervised Action Recognition | UCF101 | Shuffle and Learn (AlexNet) | Pre-Training Dataset | UCF101 | #52 of 53 | Archive leaderboard | report |
Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections