Browse State-of-the-Art › Video Alignment
Video Alignment
43 papers with code · 2 benchmarks · 4 datasets archive 2025-07-28
Benchmarks archive 2025-07-28
2 leaderboard tables shown for this task, 2 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted.
| Dataset | Best model (first row in archive order) | Paper | Code | Syntology | Compare |
|---|---|---|---|---|---|
| MSU Video Alignment and Retrieval Benchmark Suite (4 rows) | VQMT3D | ACCURATE METHOD OF TEMPORAL-SHIFT ESTIMATION FOR 3D VIDEO | — | — | Compare |
| UPenn Action (4 rows) | TCC + TCN | Temporal Cycle-Consistency Learning | code | Syntology ran 0 of 12 samples · 12 unverified | Compare |
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
4 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
No subtask under this task in the archive's task tree.
Parent tasks archive 2025-07-28
Most implemented papers archive 2025-07-28
30 shown of 43 papers with code (83 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
23 Apr 2017 7 repositories listedWhile representations are learned from an unlabeled collection of task-related videos, robot behaviors such as pouring are learned by watching a single 3rd-person demonstration by a human.
-
3 Dec 2024 2 repositories listed Syntology ran 18 of 27 samples · 9 unverified · 27 pointer-only (licence)In this report, we introduce HunyuanVideo, an innovative open-source video foundation model that demonstrates performance in video generation comparable to, or even surpassing, that of leading closed-source models.
-
21 Aug 2024 2 repositories listedTo the best of our knowledge, VE-Bench introduces the first quality assessment dataset for video editing and an effective subjective-aligned quantitative metric for this domain.
-
12 Aug 2024 2 repositories listed Syntology ran 7 of 7 samples · 0 unverifiedWe present CogVideoX, a large-scale text-to-video generation model based on diffusion transformer, which can generate 10-second continuous videos aligned with text prompt, with a frame rate of 16 fps and resolution of…
-
8 Jul 2024 2 repositories listedSora's high-motion intensity and long consistent videos have significantly impacted the field of video generation, attracting unprecedented attention.
-
3 Jan 2024 2 repositories listed Syntology ran 2 of 2 samples · 0 unverifiedTo establish a unified evaluation framework for video generation tasks, our benchmark includes 11 metrics spanning four dimensions to assess algorithm performance.
-
23 Oct 2020 2 repositories listedRecognition of human poses and actions is crucial for autonomous systems to interact smoothly with people.
-
2 Dec 2019 2 repositories listedDepictions of similar human body configurations can vary with changing viewpoints.
-
16 Apr 2019 2 repositories listed Syntology ran 0 of 12 samples · 12 unverifiedWe introduce a self-supervised representation learning method based on the task of temporal alignment between videos.
-
27 Jul 2017 2 repositories listedDiscriminative clustering has been successfully applied to a number of weakly-supervised learning tasks.
-
10 Jun 2025 1 repository listed Syntology ran 1 of 2 samples · 1 unverifiedThe parameter-efficient adaptation of the image-text pretraining model CLIP for video-text retrieval is a prominent area of research.
-
29 May 2025 1 repository listedGenerating highly dynamic and photorealistic portrait animations driven by audio and skeletal motion remains challenging due to the need for precise lip synchronization, natural facial expressions, and high-fidelity…
-
17 May 2025 1 repository listedTo this end, we present AIGVE-60K, a comprehensive dataset and benchmark for AI-Generated Video Evaluation, which features (i) comprehensive tasks, encompassing 3, 050 extensive prompts across 20 fine-grained task…
-
7 May 2025 1 repository listed Syntology ran 2 of 2 samples · 0 unverified · 2 pointer-only (licence)Customized video generation aims to produce videos featuring specific subjects under flexible user-defined conditions, yet existing methods often struggle with identity consistency and limited input modalities.
-
5 Apr 2025 1 repository listedWith the ability of 4D and video generation, Video4DGen offers a powerful tool for applications in virtual reality, animation, and beyond.
-
11 Mar 2025 1 repository listedWe propose a new task, video referring matting, which obtains the alpha matte of a specified instance by inputting a referring caption.
-
5 Mar 2025 1 repository listedThe objective of this work is to align asynchronous subtitles in sign language videos with limited labelled data.
-
31 Jan 2025 1 repository listed Syntology ran 7 of 10 samples · 3 unverifiedThe remarkable progress in text-to-video diffusion models enables photorealistic generations, although the contents of the generated video often include unnatural movement or deformation, reverse playback, and…
-
1 Jan 2025 1 repository listedHowever, existing visual-to-visual and visual-to-textual Ego-Exo video alignment methods struggle with the problem that there could be non-visual overlap for the same activity.
-
22 Nov 2024 1 repository listedAs these models become prevalent, various metrics and benchmarks have emerged to evaluate the quality of the generated videos.
-
8 Oct 2024 1 repository listed Syntology ran 8 of 8 samples · 0 unverified · 8 pointer-only (licence)In this paper, we focus on enhancing a diffusion-based text-to-video (T2V) model during the post-training phase by distilling a highly capable consistency model from a pretrained T2V model.
-
8 Sep 2024 1 repository listedMaTAV is with the advantages of aligning unimodal features to ensure consistency across different modalities and handling long input sequences to better capture contextual multimodal information.
-
6 Sep 2024 1 repository listedRobust frame-wise embeddings are essential to perform video analysis and understanding tasks.
-
1 Jul 2024 1 repository listed Syntology ran 3 of 8 samples · 5 unverifiedMeanwhile, the temporal controller incorporates an onset detector and a timestampbased adapter to achieve precise audio-video alignment.
-
20 Jun 2024 1 repository listed Syntology ran 0 of 2 samples · 2 unverified · 2 pointer-only (licence)To mitigate the risk of harmful outputs from large vision models (LVMs), we introduce the SafeSora dataset to promote research on aligning text-to-video generation with human values.
-
21 Apr 2024 1 repository listedOur approach exhibits an improved ability to leverage the video modality by using the audio modality as a bridge with the language modality.
-
18 Mar 2024 1 repository listed Syntology ran 5 of 5 samples · 0 unverified · 5 pointer-only (licence)Based on T2VQA-DB, we propose a novel transformer-based model for subjective-aligned Text-to-Video Quality Assessment (T2VQA).
-
18 Mar 2024 1 repository listedTo this end, this paper proposes a novel text-guided video inpainting model that achieves better consistency, controllability and compatibility.
-
17 Oct 2023 1 repository listed Syntology ran 5 of 11 samples · 6 unverified · 11 pointer-only (licence)For video generation, various open-sourced models and public-available services have been developed to generate high-quality videos.
-
27 Sep 2023 1 repository listed Syntology ran 1 of 1 samples · 0 unverified · 1 pointer-only (licence)In this paper, we are the first to propose a hybrid model, dubbed as Show-1, which marries pixel-based and latent-based VDMs for text-to-video generation.
Syntology lines on 13 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections