Papers › HiTeA: Hierarchical Temporal-Aware Video-Language Pre-training

HiTeA: Hierarchical Temporal-Aware Video-Language Pre-training

30 Dec 2022ICCV 2023 1arXiv:2212.14546archive 2025-07-28

Qinghao Ye, Guohai Xu, Ming Yan, Haiyang Xu, Qi Qian, Ji Zhang, Fei Huang

Video-language pre-training has advanced the performance of various downstream video-language tasks. However, most previous methods directly inherit or adapt typical image-language pre-training paradigms to video-language pre-training, thus not fully exploiting the unique characteristic of video, i.e., temporal. In this paper, we propose a Hierarchical Temporal-Aware video-language pre-training framework, HiTeA, with two novel pre-training tasks for modeling cross-modal alignment between moments and texts as well as the temporal relations of video-text pairs. Specifically, we propose a cross-modal moment exploration task to explore moments in videos, which results in detailed video moment representation. Besides, the inherent temporal relations are captured by aligning video-text pairs as a whole in different time resolutions with multi-modal temporal relation exploration task. Furthermore, we introduce the shuffling test to evaluate the temporal reliance of datasets and video-language pre-training models. We achieve state-of-the-art results on 15 well-established video-language understanding and generation tasks, especially on temporal-oriented datasets (e.g., SSv2-Template and SSv2-Label) with 8.6% and 11.1% improvement respectively. HiTeA also demonstrates strong generalization ability when directly transferred to downstream tasks in a zero-shot manner. Models and demo will be available on ModelScope.

PaperPDFConference PDF

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

No code repository is listed for this paper in the archive or in Syntology's graph.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

TGIF-ActionTGIF-FrameTGIF-TransitionVideo CaptioningVideo Question AnsweringVideo RetrievalVisual Question Answering (VQA)Zero-Shot LearningZero-Shot Video Retrievalcross-modal alignment

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Video Captioning MSR-VTT HiTeA BLEU-4 49.2 #11 of 24 Archive leaderboard report
Video Captioning MSR-VTT HiTeA CIDEr 65.1 #11 of 24 Archive leaderboard report
Video Captioning MSR-VTT HiTeA METEOR 30.7 #11 of 24 Archive leaderboard report
Video Captioning MSR-VTT HiTeA ROUGE-L 65.0 #11 of 24 Archive leaderboard report
Video Captioning MSVD HiTeA BLEU-4 71.0 #7 of 16 Archive leaderboard report
Video Captioning MSVD HiTeA CIDEr 146.9 #7 of 16 Archive leaderboard report
Video Captioning MSVD HiTeA METEOR 45.3 #7 of 16 Archive leaderboard report
Video Captioning MSVD HiTeA ROUGE-L 81.4 #7 of 16 Archive leaderboard report
Video Question Answering MSRVTT-MC HiTeA Accuracy 97.4 #2 of 7 Archive leaderboard report
Video Question Answering NExT-QA HiTeA Accuracy 63.1 #33 of 47 Archive leaderboard report
Video Retrieval ActivityNet HiTeA text-to-video R@1 49.7 #17 of 31 Archive leaderboard report
Video Retrieval ActivityNet HiTeA text-to-video R@10 86.7 #17 of 31 Archive leaderboard report
Video Retrieval ActivityNet HiTeA text-to-video R@5 77.1 #17 of 31 Archive leaderboard report
Video Retrieval DiDeMo HiTeA text-to-video R@1 56.5 #13 of 40 Archive leaderboard report
Video Retrieval DiDeMo HiTeA text-to-video R@10 89.7 #13 of 40 Archive leaderboard report
Video Retrieval DiDeMo HiTeA text-to-video R@5 81.7 #13 of 40 Archive leaderboard report
Video Retrieval LSMDC HiTeA text-to-video R@1 28.7 #12 of 38 Archive leaderboard report
Video Retrieval LSMDC HiTeA text-to-video R@10 59.0 #12 of 38 Archive leaderboard report
Video Retrieval LSMDC HiTeA text-to-video R@5 50.3 #12 of 38 Archive leaderboard report
Video Retrieval MSR-VTT-1kA HiTeA text-to-video R@1 46.8 #33 of 63 Archive leaderboard report
Video Retrieval MSR-VTT-1kA HiTeA text-to-video R@10 81.9 #33 of 63 Archive leaderboard report
Video Retrieval MSR-VTT-1kA HiTeA text-to-video R@5 71.2 #33 of 63 Archive leaderboard report
Video Retrieval SSv2-label retrieval HiTeA text-to-video R@1 55.2 #3 of 5 Archive leaderboard report
Video Retrieval SSv2-label retrieval HiTeA text-to-video R@10 81.4 #3 of 5 Archive leaderboard report
Video Retrieval SSv2-label retrieval HiTeA text-to-video R@5 89.1 #3 of 5 Archive leaderboard report
Video Retrieval SSv2-template retrieval HiTeA text-to-video R@1 85.6 #3 of 5 Archive leaderboard report
Video Retrieval SSv2-template retrieval HiTeA text-to-video R@10 100 #3 of 5 Archive leaderboard report
Video Retrieval SSv2-template retrieval HiTeA text-to-video R@5 100 #3 of 5 Archive leaderboard report
Visual Question Answering (VQA) MSRVTT-QA HiTeA Accuracy 0.459 #12 of 34 Archive leaderboard report
Visual Question Answering (VQA) MSVD-QA HiTeA Accuracy 0.556 #11 of 36 Archive leaderboard report
Visual Question Answering (VQA) TGIF-QA HiTeA Accuracy 0.732 #1 of 2 Archive leaderboard report
Zero-Shot Learning MSRVTT-QA HiTeA Accuracy 21.7 #1 of 1 Archive leaderboard report
Zero-Shot Learning MSVD-QA HiTeA Accuracy 37.4 #1 of 1 Archive leaderboard report
Zero-Shot Video Retrieval DiDeMo HiTeA-17M text-to-video R@1 43.2 #8 of 26 Archive leaderboard report
Zero-Shot Video Retrieval DiDeMo HiTeA-17M text-to-video R@10 79.0 #8 of 26 Archive leaderboard report
Zero-Shot Video Retrieval DiDeMo HiTeA-17M text-to-video R@5 69.3 #8 of 26 Archive leaderboard report
Zero-Shot Video Retrieval DiDeMo HiTeA-5M text-to-video R@1 36.1 #13 of 26 Archive leaderboard report
Zero-Shot Video Retrieval DiDeMo HiTeA-5M text-to-video R@10 70.3 #13 of 26 Archive leaderboard report
Zero-Shot Video Retrieval DiDeMo HiTeA-5M text-to-video R@5 60.1 #13 of 26 Archive leaderboard report
Zero-Shot Video Retrieval LSMDC HiTeA-17M text-to-video R@1 18.3 #7 of 16 Archive leaderboard report
Zero-Shot Video Retrieval LSMDC HiTeA-17M text-to-video R@10 44.2 #7 of 16 Archive leaderboard report
Zero-Shot Video Retrieval LSMDC HiTeA-17M text-to-video R@5 36.7 #7 of 16 Archive leaderboard report
Zero-Shot Video Retrieval LSMDC HiTeA-5M text-to-video R@1 15.5 #11 of 16 Archive leaderboard report
Zero-Shot Video Retrieval LSMDC HiTeA-5M text-to-video R@10 39.8 #11 of 16 Archive leaderboard report
Zero-Shot Video Retrieval LSMDC HiTeA-5M text-to-video R@5 31.1 #11 of 16 Archive leaderboard report
Zero-Shot Video Retrieval MSR-VTT HiTeA-17M text-to-video R@1 34.4 #19 of 41 Archive leaderboard report
Zero-Shot Video Retrieval MSR-VTT HiTeA-17M text-to-video R@10 69.9 #19 of 41 Archive leaderboard report
Zero-Shot Video Retrieval MSR-VTT HiTeA-17M text-to-video R@5 60.0 #19 of 41 Archive leaderboard report
Zero-Shot Video Retrieval MSR-VTT HiTeA-5M text-to-video R@1 29.9 #23 of 41 Archive leaderboard report
Zero-Shot Video Retrieval MSR-VTT HiTeA-5M text-to-video R@10 62.9 #23 of 41 Archive leaderboard report
Zero-Shot Video Retrieval MSR-VTT HiTeA-5M text-to-video R@5 54.2 #23 of 41 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

Test

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections