{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/videoeval-comprehensive-benchmark-suite-for","title":"VideoEval: Comprehensive Benchmark Suite for Low-Cost Evaluation of Video Foundation Model","arxiv_id":"2407.06491","date":"2024-07-09","proceeding":null,"authors":["Xinhao Li","Zhenpeng Huang","Jing Wang","Kunchang Li","LiMin Wang"],"abstract":"With the growth of high-quality data and advancement in visual pre-training paradigms, Video Foundation Models (VFMs) have made significant progress recently, demonstrating their remarkable performance on traditional video understanding benchmarks. However, the existing benchmarks (e.g. Kinetics) and their evaluation protocols are often limited by relatively poor diversity, high evaluation costs, and saturated performance metrics. In this paper, we build a comprehensive benchmark suite to address these issues, namely VideoEval. Specifically, we establish the Video Task Adaption Benchmark (VidTAB) and the Video Embedding Benchmark (VidEB) from two perspectives: evaluating the task adaptability of VFMs under few-shot conditions and assessing their representation power by directly applying to downstream tasks. With VideoEval, we conduct a large-scale study on 20 popular open-source vision foundation models. Our study reveals some insightful findings on VFMs: 1) overall, current VFMs exhibit weak generalization across diverse tasks, 2) increasing video data, whether labeled or weakly-labeled video-text pairs, does not necessarily improve task performance, 3) the effectiveness of some pre-training paradigms may not be fully validated in previous benchmarks, and 4) combining different pre-training paradigms can help improve the generalization capabilities. We believe this study serves as an important complement to the current evaluation for VFMs and offers valuable insights for the future research.","url_abs":"https://arxiv.org/abs/2407.06491v1","url_pdf":"https://arxiv.org/pdf/2407.06491v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"videoeval-comprehensive-benchmark-suite-for","repo_url":"https://github.com/leexinhao/VideoEval","is_official":1,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok"}}],"tasks":[{"task_slug":"video-understanding","task_name":"Video Understanding"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2407.06491","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2407.06491"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/leexinhao/VideoEval","reach":{"status":"ok"}}],"summary":{"ran_violates":1,"ran_draft_wrong":1,"ran":4,"unverified":5},"by_repo_kind":{"official":{"samples":11,"ran":6,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":11,"samples":[{"code_sha256_prefix":"7bff6f3b3cb560e7","entry":"convert_to_custom_text_state_dict","repo":"leexinhao/VideoEval","repo_kind":"official","path":"VidTAB_Zeroshot/eva_clip/model.py","file_url":"https://github.com/leexinhao/VideoEval/blob/HEAD/VidTAB_Zeroshot/eva_clip/model.py","link_basis":"plan_row","language":"python","status":"ran_violates","verification_level":1,"contract_check":"VIOLATES","metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"7bff6f3b3cb560e7"}},{"code_sha256_prefix":"dcd422d66b0581d8","entry":"get_cast_dtype","repo":"leexinhao/VideoEval","repo_kind":"official","path":"VidTAB_Zeroshot/eva_clip/model.py","file_url":"https://github.com/leexinhao/VideoEval/blob/HEAD/VidTAB_Zeroshot/eva_clip/model.py","link_basis":"harvester_set","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"OUTPUT_MISDECLARED","metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"dcd422d66b0581d8"}},{"code_sha256_prefix":"5770f110b4cb1743","entry":"prompt_generate","repo":"leexinhao/VideoEval","repo_kind":"official","path":"VidTAB_Zeroshot/vid_prompt_gen.py","file_url":"https://github.com/leexinhao/VideoEval/blob/HEAD/VidTAB_Zeroshot/vid_prompt_gen.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"5770f110b4cb1743"}},{"code_sha256_prefix":"31baa455fe73dd62","entry":"read_text","repo":"leexinhao/VideoEval","repo_kind":"official","path":"VidTAB_Zeroshot/vid_prompt_gen.py","file_url":"https://github.com/leexinhao/VideoEval/blob/HEAD/VidTAB_Zeroshot/vid_prompt_gen.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"31baa455fe73dd62"}},{"code_sha256_prefix":"eba4b88f42e1982c","entry":"read_video_decord","repo":"leexinhao/VideoEval","repo_kind":"official","path":"VidTAB_Zeroshot/datasets.py","file_url":"https://github.com/leexinhao/VideoEval/blob/HEAD/VidTAB_Zeroshot/datasets.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"eba4b88f42e1982c"}},{"code_sha256_prefix":"d3cb626155896501","entry":"sample_frame_indices","repo":"leexinhao/VideoEval","repo_kind":"official","path":"VidTAB_Zeroshot/datasets.py","file_url":"https://github.com/leexinhao/VideoEval/blob/HEAD/VidTAB_Zeroshot/datasets.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"d3cb626155896501"}},{"code_sha256_prefix":"b34d6c17c6e09a04","entry":"build_model_from_openai_state_dict","repo":"leexinhao/VideoEval","repo_kind":"official","path":"VidTAB_Zeroshot/eva_clip/model.py","file_url":"https://github.com/leexinhao/VideoEval/blob/HEAD/VidTAB_Zeroshot/eva_clip/model.py","link_basis":"plan_row","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"b34d6c17c6e09a04"}},{"code_sha256_prefix":"ddcbd45e940484ee","entry":"gather_features","repo":"leexinhao/VideoEval","repo_kind":"official","path":"VidTAB_Zeroshot/eva_clip/loss.py","file_url":"https://github.com/leexinhao/VideoEval/blob/HEAD/VidTAB_Zeroshot/eva_clip/loss.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"ddcbd45e940484ee"}},{"code_sha256_prefix":"5b77a4135c6e26ae","entry":"prompt_generate","repo":"leexinhao/VideoEval","repo_kind":"official","path":"VidTAB_Zeroshot/img_prompt_gen.py","file_url":"https://github.com/leexinhao/VideoEval/blob/HEAD/VidTAB_Zeroshot/img_prompt_gen.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"5b77a4135c6e26ae"}},{"code_sha256_prefix":"1992dca569248e15","entry":"read_text","repo":"leexinhao/VideoEval","repo_kind":"official","path":"VidTAB_Zeroshot/img_prompt_gen.py","file_url":"https://github.com/leexinhao/VideoEval/blob/HEAD/VidTAB_Zeroshot/img_prompt_gen.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"1992dca569248e15"}},{"code_sha256_prefix":"2a377da4a76a2d44","entry":"register_pooler","repo":"leexinhao/VideoEval","repo_kind":"official","path":"VidTAB_Zeroshot/eva_clip/hf_model.py","file_url":"https://github.com/leexinhao/VideoEval/blob/HEAD/VidTAB_Zeroshot/eva_clip/hf_model.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"2a377da4a76a2d44"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}