Methods › Computer Vision › Vision and Language Pre-Trained Models › InternVideo

InternVideo: General Video Foundation Models via Generative and Discriminative Learning

InternVideo

5 papers tagged archive 2025-07-28

archive 2025-07-28 Description, source and code snippet are the archive's method entry.

The foundation models have recently shown excellent performance on a variety of downstream tasks in computer vision. However, most existing vision foundation models simply focus on image-level pretraining and adpation, which are limited for dynamic and complex video-level understanding tasks. To fill the gap, we present general video foundation models, InternVideo, by taking advantage of both generative and discriminative self-supervised video learning. Specifically, InternVideo efficiently explores masked video modeling and video-language contrastive learning as the pretraining objectives, and selectively coordinates video representations of these two complementary frameworks in a learnable manner to boost various video applications. Without bells and whistles, InternVideo achieves state-of-the-art performance on 39 video datasets from extensive tasks including video action recognition/detection, video-language alignment, and open-world video applications. Especially, our methods can obtain 91.1% and 77.2% top-1 accuracy on the challenging Kinetics-400 and Something-Something V2 benchmarks, respectively. All of these results effectively show the generality of our InternVideo for video understanding. The code will be released at https://github.com/OpenGVLab/InternVideo.

Papers archive 2025-07-28

5 shown of 5, newest first. Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code; it is per sample and not a correctness claim.

Tasks archive 2025-07-28

20 shown of 31 tasks the archive attaches to papers tagged with this method, by distinct papers. A task without a page in the catalog is plain text.

TaskPapers
Video Question Answering2
Video Understanding2
Action Classification1
Action Recognition1
Contrastive Learning1
Denoising1
Descriptive1
Future Hand Prediction1
Language Modelling1
Large Language Model1
Moment Queries1
Natural Language Queries1
Object1
Object Detection1
Open Set Action Recognition1
Question Answering1
Representation Learning1
Sentence1
Short-term Object Interaction Anticipation1
Spatio-Temporal Action Localization1

Usage over time archive 2025-07-28

Papers per year tagged with InternVideo: 2022 to 2023, peak 3 3 0 2022: 2 papers 2022 2023: 3 papers 2023
Papers per year the archive tags with this method, by the paper's archive date (5 dated). Bars are counts, not a trend claim.

Components: the archive holds no method-to-method composition, so PwC's Components table cannot be rebuilt; the Papers list carries no Results column for the same reason (the archive does not join its leaderboard rows to method tags).

Categories archive 2025-07-28

Vision and Language Pre-Trained Models

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections