Papers › InternVideo: General Video Foundation Models via Generative and Discriminative Learning

InternVideo: General Video Foundation Models via Generative and Discriminative Learning

6 Dec 2022arXiv:2212.03191archive 2025-07-28

Yi Wang, Kunchang Li, Yizhuo Li, Yinan He, Bingkun Huang, Zhiyu Zhao, Hongjie Zhang, Jilan Xu, Yi Liu, Zun Wang, Sen Xing, Guo Chen, Junting Pan, Jiashuo Yu, Yali Wang, LiMin Wang, Yu Qiao

The foundation models have recently shown excellent performance on a variety of downstream tasks in computer vision. However, most existing vision foundation models simply focus on image-level pretraining and adpation, which are limited for dynamic and complex video-level understanding tasks. To fill the gap, we present general video foundation models, InternVideo, by taking advantage of both generative and discriminative self-supervised video learning. Specifically, InternVideo efficiently explores masked video modeling and video-language contrastive learning as the pretraining objectives, and selectively coordinates video representations of these two complementary frameworks in a learnable manner to boost various video applications. Without bells and whistles, InternVideo achieves state-of-the-art performance on 39 video datasets from extensive tasks including video action recognition/detection, video-language alignment, and open-world video applications. Especially, our methods can obtain 91.1% and 77.2% top-1 accuracy on the challenging Kinetics-400 and Something-Something V2 benchmarks, respectively. All of these results effectively show the generality of our InternVideo for video understanding. The code will be released at https://github.com/OpenGVLab/InternVideo .

PaperPDFCodeCode Syntology ran

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

For agents, Syntology's MCP tool lists every function and class Syntology harvested from this paper and whether it ran (how to connect): get_harvested_code_for_paper(arxiv_id="2212.03191")

Code

Syntology Ran 3 of 3 code samples harvested from 1 repository linked to this paper; 0 have no recorded run. Of those that ran: 2 ran · our draft was wrong; 1 ran · fixture could not drive it.

By repository: official repository: 3 samples from 1 repository, 3 ran. The run record, sample by sample. “Ran” means executed on a synthesized input, not that the code is correct or reproduces the paper.

opengvlab/internvideo officialmentioned in papermentioned on GitHubpytorch report
yingsen1/unimd mentioned on GitHubpytorch report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

3 samples harvested; 3 ran; 0 honoured the contract we drafted; 0 have no recorded run. Read from Syntology's graph 2026-09-24; that is when this build read the record, not when the samples ran.

2ran · our draft was wrong
1ran · fixture could not drive it

Licence: 0 of the 3 samples are pointer only, meaning Syntology does not serve that copy's text. This page shows no code text for any sample; each one links to its file in the repository.

Harvested from opengvlab/internvideo. “Ran” means the sample executed on a synthesized input. It does not mean the output is correct, and nothing here reproduces the paper's results. “Honoured” and “violated” refer to a contract Syntology drafted from the code itself; “our draft was wrong” and “fixture could not drive it” are failures of Syntology's instrument, not of the code.

Each sample ends with its code_sha256, Syntology's identity for that exact code. An agent fetches the stored sample with Syntology's MCP tool get_code(code_sha256="…") (how to connect); click an identity to copy that call.

Repository labels, per sample. official repository: The archive marks this repository official for the paper. named in the paper: The archive records that the paper mentions this repository; it is not marked official. community (archive-listed): In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper. found in paper text by Syntology: Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted. community: Not in the archive's code links for this paper; a community repository Syntology harvested. Samples from a repository marked official are listed first. Licence labels name the repository's licence as recorded at harvest. “Pointer only” means Syntology does not serve that copy's text, for one of four reasons: no licence file was found; the licence was not identified; the licence is recorded as permissive but that copy's record is not marked cleared; or the licence is outside the permissive list Syntology serves text under (MIT, Apache-2.0, BSD and similar). Some licences outside that list permit redistribution, such as WTFPL, and GPL-3.0 under its conditions; they are simply not on the list. Hover a licence label for the reason. File links open the file on GitHub at the default branch, which may have changed since the harvest.

compute_acc opengvlab/internvideo/Data/InternVid/utils/basic_utils.py official repository ran · fixture could not drive it fingerprinted Apache-2.0 (permissive) · aa81d03f07735ab8 · report
compute_n_params opengvlab/internvideo/Data/InternVid/utils/basic_utils.py official repository ran · our draft was wrong Apache-2.0 (permissive) · ad12a494674d23fb · report
load_json opengvlab/internvideo/Data/InternVid/utils/basic_utils.py official repository ran · our draft was wrong Apache-2.0 (permissive) · 2d946250dd2f5a4f · report

Tasks

Action ClassificationAction RecognitionContrastive LearningOpen Set Action RecognitionSpatio-Temporal Action LocalizationTemporal Action LocalizationVideo Question AnsweringVideo RetrievalVideo UnderstandingVisual Question Answering (VQA)Zero-Shot Video Question AnswerZero-Shot Video Retrieval

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Action Classification Kinetics-400 InternVideo Acc@1 91.1 #5 of 207 Archive leaderboard report
Action Classification Kinetics-600 InternVideo-T Top-1 Accuracy 91.3 #5 of 65 Archive leaderboard report
Action Classification Kinetics-700 InternVideo-T Top-1 Accuracy 84.0 #3 of 36 Archive leaderboard report
Action Recognition AVA v2.2 InternVideo mAP 41.01 #6 of 38 Archive leaderboard report
Action Recognition Something-Something V1 InternVideo Top 1 Accuracy 70.0 #1 of 74 Archive leaderboard report
Action Recognition Something-Something V2 InternVideo Top-1 Accuracy 77.2 #3 of 123 Archive leaderboard report
Open Set Action Recognition UCF-HMDB InternVideo AUROC 85.48 #1 of 1 Archive leaderboard report
Open Set Action Recognition UCF101-MiTv2 InternVideo AUROC 91.85 #1 of 1 Archive leaderboard report
Spatio-Temporal Action Localization AVA-Kinetics InternVideo val mAP 41.01 #3 of 7 Archive leaderboard report
Temporal Action Localization ActivityNet-1.3 InternVideo mAP 39.00 #9 of 33 Archive leaderboard report
Temporal Action Localization FineAction InternVideo mAP 17.57 #6 of 9 Archive leaderboard report
Temporal Action Localization HACS InternVideo Average-mAP 41.55 #7 of 12 Archive leaderboard report
Temporal Action Localization THUMOS’14 ActionFormer (InternVideo features) Avg mAP (0.3:0.7) 71.58 #5 of 42 Archive leaderboard report
Video Question Answering STAR Benchmark InternVideo Average Accuracy 58.7 #4 of 17 Archive leaderboard report
Video Retrieval ActivityNet InternVideo text-to-video R@1 62.2 #8 of 31 Archive leaderboard report
Video Retrieval ActivityNet InternVideo video-to-text R@1 62.8 #8 of 31 Archive leaderboard report
Video Retrieval DiDeMo InternVideo text-to-video R@1 57.9 #10 of 40 Archive leaderboard report
Video Retrieval DiDeMo InternVideo video-to-text R@1 59.1 #10 of 40 Archive leaderboard report
Video Retrieval LSMDC InternVideo text-to-video R@1 34.0 #8 of 38 Archive leaderboard report
Video Retrieval LSMDC InternVideo video-to-text R@1 34.9 #8 of 38 Archive leaderboard report
Video Retrieval MSR-VTT InternVideo text-to-video R@1 55.2 #8 of 40 Archive leaderboard report
Video Retrieval MSR-VTT InternVideo video-to-text R@1 57.9 #8 of 40 Archive leaderboard report
Video Retrieval MSVD InternVideo text-to-video R@1 58.4 #3 of 24 Archive leaderboard report
Video Retrieval MSVD InternVideo video-to-text R@1 76.3 #3 of 24 Archive leaderboard report
Video Retrieval VATEX InternVideo text-to-video R@1 71.1 #6 of 13 Archive leaderboard report
Video Retrieval VATEX InternVideo video-to-text R@1 87.2 #6 of 13 Archive leaderboard report
Visual Question Answering (VQA) MSRVTT-QA InternVideo Accuracy 0.471 #6 of 34 Archive leaderboard report
Visual Question Answering (VQA) MSVD-QA InternVideo Accuracy 0.555 #12 of 36 Archive leaderboard report
Visual Question Answering (VQA) TGIF-QA InternVideo Accuracy 0.722 #2 of 2 Archive leaderboard report
Zero-Shot Video Question Answer EgoSchema (fullset) InternVideo Accuracy 32.1 #25 of 29 Archive leaderboard report
Zero-Shot Video Question Answer STAR Benchmark InternVideo Accuracy 41.6 #4 of 4 Archive leaderboard report
Zero-Shot Video Question Answer TVQA InternVideo (no speech) Accuracy 35.9 #8 of 9 Archive leaderboard report
Zero-Shot Video Retrieval ActivityNet InternVideo text-to-video R@1 30.7 #11 of 12 Archive leaderboard report
Zero-Shot Video Retrieval ActivityNet InternVideo video-to-text R@1 31.4 #11 of 12 Archive leaderboard report
Zero-Shot Video Retrieval DiDeMo InternVideo text-to-video R@1 31.5 #16 of 26 Archive leaderboard report
Zero-Shot Video Retrieval DiDeMo InternVideo text-to-video R@10 68.2 #16 of 26 Archive leaderboard report
Zero-Shot Video Retrieval DiDeMo InternVideo text-to-video R@5 57.6 #16 of 26 Archive leaderboard report
Zero-Shot Video Retrieval DiDeMo InternVideo video-to-text R@1 33.5 #16 of 26 Archive leaderboard report
Zero-Shot Video Retrieval DiDeMo InternVideo video-to-text R@10 71.1 #16 of 26 Archive leaderboard report
Zero-Shot Video Retrieval DiDeMo InternVideo video-to-text R@5 60.3 #16 of 26 Archive leaderboard report
Zero-Shot Video Retrieval LSMDC InternVideo text-to-video R@1 17.6 #8 of 16 Archive leaderboard report
Zero-Shot Video Retrieval LSMDC InternVideo text-to-video R@10 40.2 #8 of 16 Archive leaderboard report
Zero-Shot Video Retrieval LSMDC InternVideo text-to-video R@5 32.4 #8 of 16 Archive leaderboard report
Zero-Shot Video Retrieval LSMDC InternVideo video-to-text R@1 13.2 #8 of 16 Archive leaderboard report
Zero-Shot Video Retrieval LSMDC InternVideo video-to-text R@10 34.9 #8 of 16 Archive leaderboard report
Zero-Shot Video Retrieval LSMDC InternVideo video-to-text R@5 27.8 #8 of 16 Archive leaderboard report
Zero-Shot Video Retrieval MSR-VTT InternVideo text-to-video R@1 40.7 #14 of 41 Archive leaderboard report
Zero-Shot Video Retrieval MSR-VTT InternVideo video-to-text R@1 39.6 #14 of 41 Archive leaderboard report
Zero-Shot Video Retrieval MSVD InternVideo text-to-video R@1 43.4 #11 of 14 Archive leaderboard report
Zero-Shot Video Retrieval MSVD InternVideo video-to-text R@1 67.6 #11 of 14 Archive leaderboard report
Zero-Shot Video Retrieval VATEX InternVideo text-to-video R@1 49.5 #5 of 5 Archive leaderboard report
Zero-Shot Video Retrieval VATEX InternVideo video-to-text R@1 69.5 #5 of 5 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

Contrastive LearningInternVideo

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections