{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/videoxum-cross-modal-visual-and-textural","title":"VideoXum: Cross-modal Visual and Textural Summarization of Videos","arxiv_id":"2303.12060","date":"2023-03-21","proceeding":null,"authors":["Jingyang Lin","Hang Hua","Ming Chen","Yikang Li","Jenhao Hsiao","Chiuman Ho","Jiebo Luo"],"abstract":"Video summarization aims to distill the most important information from a source video to produce either an abridged clip or a textual narrative. Traditionally, different methods have been proposed depending on whether the output is a video or text, thus ignoring the correlation between the two semantically related tasks of visual summarization and textual summarization. We propose a new joint video and text summarization task. The goal is to generate both a shortened video clip along with the corresponding textual summary from a long video, collectively referred to as a cross-modal summary. The generated shortened video clip and text narratives should be semantically well aligned. To this end, we first build a large-scale human-annotated dataset -- VideoXum (X refers to different modalities). The dataset is reannotated based on ActivityNet. After we filter out the videos that do not meet the length requirements, 14,001 long videos remain in our new dataset. Each video in our reannotated dataset has human-annotated video summaries and the corresponding narrative summaries. We then design a novel end-to-end model -- VTSUM-BILP to address the challenges of our proposed task. Moreover, we propose a new metric called VT-CLIPScore to help evaluate the semantic consistency of cross-modality summary. The proposed model achieves promising performance on this new task and establishes a benchmark for future research.","url_abs":"https://arxiv.org/abs/2303.12060v3","url_pdf":"https://arxiv.org/pdf/2303.12060v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"videoxum-cross-modal-visual-and-textural","repo_url":"https://github.com/jylins/videoxum","is_official":1,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"text-summarization","task_name":"Text Summarization"},{"task_slug":"video-summarization","task_name":"Video Summarization"}],"methods":[{"method_slug":"clip","method_name":"CLIP"}],"datasets_introduced":[{"slug":"videoxum-1","name":"VideoXum","full_name":""}],"methods_introduced":[],"results":[{"leaderboard":"/sota/video-summarization-on-videoxum","task":"Video Summarization","dataset":"videoxum","model":"VTSUM-BLIP","rank_in_archive_order":1,"of":1,"metrics":{"1 shot Micro-F1":"23.5"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2303.12060","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2303.12060"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/jylins/videoxum","reach":null}],"summary":{"ran_violates":1},"by_repo_kind":{"official":{"samples":1,"ran":1,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":1,"samples":[{"code_sha256_prefix":"4d1d9c47619920ba","entry":"is_url","repo":"jylins/videoxum","repo_kind":"official","path":"models/vtsum_blip.py","file_url":"https://github.com/jylins/videoxum/blob/HEAD/models/vtsum_blip.py","link_basis":"first_harvest_node","language":"python","status":"ran_violates","verification_level":1,"contract_check":"VIOLATES","metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"4d1d9c47619920ba"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}