Methods › Computer Vision › Vision and Language Pre-Trained Models › InternVideo
InternVideo: General Video Foundation Models via Generative and Discriminative Learning
InternVideo
archive 2025-07-28 Description, source and code snippet are the archive's method entry.
The foundation models have recently shown excellent performance on a variety of downstream tasks in computer vision. However, most existing vision foundation models simply focus on image-level pretraining and adpation, which are limited for dynamic and complex video-level understanding tasks. To fill the gap, we present general video foundation models, InternVideo, by taking advantage of both generative and discriminative self-supervised video learning. Specifically, InternVideo efficiently explores masked video modeling and video-language contrastive learning as the pretraining objectives, and selectively coordinates video representations of these two complementary frameworks in a learnable manner to boost various video applications. Without bells and whistles, InternVideo achieves state-of-the-art performance on 39 video datasets from extensive tasks including video action recognition/detection, video-language alignment, and open-world video applications. Especially, our methods can obtain 91.1% and 77.2% top-1 accuracy on the challenging Kinetics-400 and Something-Something V2 benchmarks, respectively. All of these results effectively show the generality of our InternVideo for video understanding. The code will be released at https://github.com/OpenGVLab/InternVideo.
Papers archive 2025-07-28
5 shown of 5, newest first. Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code; it is per sample and not a correctness claim.
-
Multi-Sentence Grounding for Long-term Instructional Video 21 Dec 2023 · 0 repositories · arXiv:2312.14055
-
Harvest Video Foundation Models via Efficient Post-Pretraining 30 Oct 2023 · 1 repository · arXiv:2310.19554
-
TVTSv2: Learning Out-of-the-box Spatiotemporal Visual Representations at Scale 23 May 2023 · 1 repository · arXiv:2305.14173
-
InternVideo: General Video Foundation Models via Generative and Discriminative Learning 6 Dec 2022 · 2 repositories · arXiv:2212.03191Syntology ran 3 of 3 samples · 0 unverified
-
InternVideo-Ego4D: A Pack of Champion Solutions to Ego4D Challenges 17 Nov 2022 · 2 repositories · arXiv:2211.09529
Tasks archive 2025-07-28
20 shown of 31 tasks the archive attaches to papers tagged with this method, by distinct papers. A task without a page in the catalog is plain text.
Usage over time archive 2025-07-28
Components: the archive holds no method-to-method composition, so PwC's Components table cannot be rebuilt; the Papers list carries no Results column for the same reason (the archive does not join its leaderboard rows to method tags).
Categories archive 2025-07-28
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections