Papers › Large-Scale Video Classification with Convolutional Neural Networks

Large-Scale Video Classification with Convolutional Neural Networks

23 Jun 20142014 IEEE Conference on Computer Vision and Pattern Recognition 2014 6archive 2025-07-28

Andrej Karpathy, George Toderici, Sanketh Shetty, Thomas Leung, Rahul Sukthankar, Li Fei-Fei

Convolutional Neural Networks (CNNs) have been established as a powerful class of models for image recognition problems. Encouraged by these results, we provide an extensive empirical evaluation of CNNs on large-scale video classification using a new dataset of 1 million YouTube videos belonging to 487 classes. We study multiple approaches for extending the connectivity of a CNN in time domain to take advantage of local spatio-temporal information and suggest a multiresolution, foveated architecture as a promising way of speeding up the training. Our best spatio-temporal networks display significant performance improvements compared to strong feature-based baselines (55.3% to 63.9%), but only a surprisingly modest improvement compared to single-frame models (59.3% to 60.9%). We further study the generalization performance of our best model by retraining the top layers on the UCF-101 Action Recognition dataset and observe significant performance improvements compared to the UCF-101 baseline model (63.3% up from 43.9%).

PaperPDFConference PDFCode

Code

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Action RecognitionClassificationGeneral ClassificationSkeleton Based Action RecognitionVideo Classification

Datasets

Introduced by this paper, per the archive.

Sports-1M

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Action Recognition Sports-1M DeepVideo’s Slow Fusion Clip Hit@1 41.9 #9 of 9 Archive leaderboard report
Action Recognition Sports-1M DeepVideo’s Slow Fusion Video hit@1 60.9 #9 of 9 Archive leaderboard report
Action Recognition Sports-1M DeepVideo’s Slow Fusion Video hit@5 80.2 #9 of 9 Archive leaderboard report
Action Recognition UCF101 Slow Fusion + Finetune top 3 layers 3-fold Accuracy 65.4 #85 of 91 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections