{"url":"/method/internvideo","slug":"internvideo","name":"InternVideo","full_name":"InternVideo: General Video Foundation Models via Generative and Discriminative Learning","full_name_withheld":false,"description_markdown":"The foundation models have recently shown excellent performance on a variety of downstream tasks in computer vision. However, most existing vision foundation models simply focus on image-level pretraining and adpation, which are limited for dynamic and complex video-level understanding tasks. To fill the gap, we present general video foundation models, InternVideo, by taking advantage of both generative and discriminative self-supervised video learning. Specifically, InternVideo efficiently explores masked video modeling and video-language contrastive learning as the pretraining objectives, and selectively coordinates video representations of these two complementary frameworks in a learnable manner to boost various video applications. Without bells and whistles, InternVideo achieves state-of-the-art performance on 39 video datasets from extensive tasks including video action recognition/detection, video-language alignment, and open-world video applications. Especially, our methods can obtain 91.1% and 77.2% top-1 accuracy on the challenging Kinetics-400 and Something-Something V2 benchmarks, respectively. All of these results effectively show the generality of our InternVideo for video understanding. The code will be released at https://github.com/OpenGVLab/InternVideo.","description_state":"present","introduced_year":null,"introduced_by":{"title":null,"paper":null,"first_author":null,"n_authors":0,"url_abs":null,"archive_paper_url":null},"source":{"url":null,"title":null,"url_on_a_paper_host":false},"code_snippet_url":null,"code_snippet_url_on_a_code_host":false,"categories":[{"area":"Computer Vision","area_id":"computer-vision","collection":"Vision and Language Pre-Trained Models","url":"/methods/category/vision-and-language-pre-trained-models","pwc_aliases":[]}],"n_papers_tagged":5,"archive_num_papers":5,"papers_newest_first":[{"paper":null,"title":"Multi-Sentence Grounding for Long-term Instructional Video","date":"2023-12-21","arxiv_id":"2312.14055","n_code_links":0,"syntology":null},{"paper":"/paper/harvest-video-foundation-models-via-efficient","title":"Harvest Video Foundation Models via Efficient Post-Pretraining","date":"2023-10-30","arxiv_id":"2310.19554","n_code_links":1,"syntology":null},{"paper":"/paper/tvtsv2-learning-out-of-the-box-spatiotemporal","title":"TVTSv2: Learning Out-of-the-box Spatiotemporal Visual Representations at Scale","date":"2023-05-23","arxiv_id":"2305.14173","n_code_links":1,"syntology":null},{"paper":"/paper/internvideo-general-video-foundation-models","title":"InternVideo: General Video Foundation Models via Generative and Discriminative Learning","date":"2022-12-06","arxiv_id":"2212.03191","n_code_links":2,"syntology":{"ran":3,"of":3,"unverified":0,"pointer_only":0}},{"paper":"/paper/internvideo-ego4d-a-pack-of-champion","title":"InternVideo-Ego4D: A Pack of Champion Solutions to Ego4D Challenges","date":"2022-11-17","arxiv_id":"2211.09529","n_code_links":2,"syntology":null}],"papers_shown":5,"tasks":[{"task":"/task/video-question-answering","name":"Video Question Answering","papers":2},{"task":"/task/video-understanding","name":"Video Understanding","papers":2},{"task":"/task/action-classification","name":"Action Classification","papers":1},{"task":"/task/action-recognition-in-videos","name":"Action Recognition","papers":1},{"task":"/task/contrastive-learning","name":"Contrastive Learning","papers":1},{"task":"/task/denoising","name":"Denoising","papers":1},{"task":"/task/descriptive","name":"Descriptive","papers":1},{"task":"/task/future-hand-prediction","name":"Future Hand Prediction","papers":1},{"task":"/task/language-modelling","name":"Language Modelling","papers":1},{"task":"/task/large-language-model","name":"Large Language Model","papers":1},{"task":"/task/moment-queries","name":"Moment Queries","papers":1},{"task":"/task/natural-language-queries","name":"Natural Language Queries","papers":1},{"task":"/task/object","name":"Object","papers":1},{"task":"/task/object-detection","name":"Object Detection","papers":1},{"task":"/task/open-set-action-recognition","name":"Open Set Action Recognition","papers":1},{"task":"/task/question-answering","name":"Question Answering","papers":1},{"task":"/task/representation-learning","name":"Representation Learning","papers":1},{"task":"/task/sentence","name":"Sentence","papers":1},{"task":"/task/short-term-object-interaction-anticipation","name":"Short-term Object Interaction Anticipation","papers":1},{"task":"/task/spatio-temporal-action-localization","name":"Spatio-Temporal Action Localization","papers":1}],"tasks_shown":20,"n_tasks":31,"usage_by_year":[{"year":"2022","papers":2},{"year":"2023","papers":3}],"row_source":"methods_table","archive":{"source":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","archive_url":"https://paperswithcode.com/method/internvideo"},"syntology_read_at":"2026-09-24T18:15:14+00:00"}