Browse State-of-the-Art › Video Understanding › Papers, page 7
Video Understanding
Papers archive 2025-07-28
archive papers tagged: 1,149 · with a code link: 542 · where Syntology ran a sample: 218 (182 with a run with no instrument failure, 36 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (218 of 1,149 tagged: 182 with a run with no instrument failure, 36 where every run was a failure of Syntology's instrument)
Page 7 of 12: papers 601 to 700 of 1,149, in archive order: by repositories listed in the archive (most first), then newest first, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers that list no repository come after every paper that lists one.
Papers without a page here are shown as plain text. A Syntology line reads “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the figure “where Syntology's instrument failed” counts failures of Syntology's instrument, not of the code. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code); hover it for the repositories the samples that ran came from. Abstracts are on each paper's page.
-
DyMU: Dynamic Merging and Virtual Unmerging for Efficient VLMs23 Apr 2025 0 repositories listed
-
An LMM for Efficient Video Understanding via Reinforced Compression of Video Cubes21 Apr 2025 0 repositories listed
-
Grounding-MD: Grounded Video-language Pre-training for Open-World Moment Detection20 Apr 2025 0 repositories listed
-
OmniV-Med: Scaling Medical Vision-Language Model for Universal Visual Understanding20 Apr 2025 0 repositories listed
-
ResNetVLLM -- Multi-modal Vision LLM for the Video Understanding Task20 Apr 2025 0 repositories listed
-
How Well Can General Vision-Language Models Learn Medicine By Watching Public Educational Videos?19 Apr 2025 0 repositories listed
-
Prototypes are Balanced Units for Efficient and Effective Partially Relevant Video Retrieval17 Apr 2025 0 repositories listed
-
16 Apr 2025 0 repositories listed
-
OmniVDiff: Omni Controllable Video Diffusion for Generation and Understanding15 Apr 2025 0 repositories listed
-
PVUW 2025 Challenge Report: Advances in Pixel-level Understanding of Complex Videos in the Wild15 Apr 2025 0 repositories listed
-
Mavors: Multi-granularity Video Representation for Multimodal Large Language Model14 Apr 2025 0 repositories listed
-
Towards Efficient and Robust Moment Retrieval System: A Unified Framework for Multi-Granularity Models and Temporal Reranking11 Apr 2025 0 repositories listed
-
How Can Objects Help Video-Language Understanding?10 Apr 2025 0 repositories listed
-
SF2T: Self-supervised Fragment Finetuning of Video-LLMs for Fine-Grained Understanding10 Apr 2025 0 repositories listed
-
VideoExpert: Augmented LLM for Temporal-Sensitive Video Understanding10 Apr 2025 0 repositories listed
-
From 128K to 4M: Efficient Training of Ultra-Long Context Large Language Models8 Apr 2025 0 repositories listed
-
From Broadcast to Minimap: Achieving State-of-the-Art SoccerNet Game State Reconstruction8 Apr 2025 0 repositories listed
-
InstructionBench: An Instructional Video Understanding Benchmark7 Apr 2025 0 repositories listed
-
Learning Audio-guided Video Representation with Gated Attention for Video-Text Retrieval3 Apr 2025 0 repositories listed
-
Moment Quantization for Video Temporal Grounding3 Apr 2025 0 repositories listed
-
Aligned Better, Listen Better for Audio-Visual Large Language Models2 Apr 2025 0 repositories listed
-
Is Temporal Prompting All We Need For Limited Labeled Action Recognition?2 Apr 2025 0 repositories listed
-
TimeSearch: Hierarchical Video Search with Spotlight and Reflection for Human-like Long Video Understanding2 Apr 2025 0 repositories listed
-
DANTE-AD: Dual-Vision Attention Network for Long-Term Audio Description31 Mar 2025 0 repositories listed
-
H2VU-Benchmark: A Comprehensive Benchmark for Hierarchical Holistic Video Understanding31 Mar 2025 0 repositories listed
-
30 Mar 2025 0 repositories listed
-
OmniMMI: A Comprehensive Multi-modal Interaction Benchmark in Streaming Video Contexts29 Mar 2025 0 repositories listed
-
From Trial to Triumph: Advancing Long Video Understanding via Visual Context Sample Scaling and Self-reward Alignment26 Mar 2025 0 repositories listed
-
Self-ReS: Self-Reflection in Large Vision-Language Models for Long Video Understanding26 Mar 2025 0 repositories listed
-
Breaking the Encoder Barrier for Seamless Video-Language Understanding24 Mar 2025 0 repositories listed
-
CRCL: Causal Representation Consistency Learning for Anomaly Detection in Surveillance Videos24 Mar 2025 0 repositories listed
-
SlowFast-LLaVA-1.5: A Family of Token-Efficient Video Large Language Models for Long-Form Video Understanding24 Mar 2025 0 repositories listed
-
Unbiasing through Textual Descriptions: Mitigating Representation Bias in Video Benchmarks24 Mar 2025 0 repositories listed
-
Video-XL-Pro: Reconstructive Token Compression for Extremely Long Video Understanding24 Mar 2025 0 repositories listed
-
Collaborative Temporal Consistency Learning for Point-supervised Natural Language Video Localization22 Mar 2025 0 repositories listed
-
PVChat: Personalized Video Chat with One-Shot Learning21 Mar 2025 0 repositories listed
-
Temporal Action Detection Model Compression by Progressive Block Drop21 Mar 2025 0 repositories listed
-
DocVideoQA: Towards Comprehensive Understanding of Document-Centric Videos through Question Answering20 Mar 2025 0 repositories listed
-
MASH-VLM: Mitigating Action-Scene Hallucination in Video-LLMs through Disentangled Spatial-Temporal Representations20 Mar 2025 0 repositories listed
-
What can Off-the-Shelves Large Multi-Modal Models do for Dynamic Scene Graph Generation?20 Mar 2025 0 repositories listed
-
FAVOR-Bench: A Comprehensive Benchmark for Fine-Grained Video Motion Understanding19 Mar 2025 0 repositories listed
-
Impossible Videos18 Mar 2025 0 repositories listed
-
Improving LLM Video Understanding with 16 Frames Per Second18 Mar 2025 0 repositories listed
-
SpaceVLLM: Endowing Multimodal Large Language Model with Spatio-Temporal Video Grounding Capability18 Mar 2025 0 repositories listed
-
Logic-in-Frames: Dynamic Keyframe Search via Visual Semantic-Logical Verification for Long Video Understanding17 Mar 2025 0 repositories listed
-
Long-VMNet: Accelerating Long-Form Video Understanding via Fixed Memory17 Mar 2025 0 repositories listed
-
Towards Scalable Modeling of Compressed Videos for Efficient Action Recognition17 Mar 2025 0 repositories listed
-
LLaVA-MLB: Mitigating and Leveraging Attention Bias for Training-Free Video LLMs14 Mar 2025 0 repositories listed
-
V-STaR: Benchmarking Video-LLMs on Video Spatio-Temporal Reasoning14 Mar 2025 0 repositories listed
-
Vamba: Understanding Hour-Long Videos with Hybrid Mamba-Transformers14 Mar 2025 0 repositories listed
-
Watch and Learn: Leveraging Expert Knowledge and Language for Surgical Video Understanding14 Mar 2025 0 repositories listed
-
13 Mar 2025 0 repositories listed Syntology 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
TIME: Temporal-sensitive Multi-dimensional Instruction Tuning and Benchmarking for Video-LLMs13 Mar 2025 0 repositories listed
-
Everything Can Be Described in Words: A Simple Unified Multi-Modal Framework with Semantic and Temporal Alignment12 Mar 2025 0 repositories listed
-
Exo2Ego: Exocentric Knowledge Guided MLLM for Egocentric Video Understanding12 Mar 2025 0 repositories listed
-
FaVChat: Unlocking Fine-Grained Facail Video Understanding with Multimodal Large Language Models12 Mar 2025 0 repositories listed
-
Generative Frame Sampler for Long Video Understanding12 Mar 2025 0 repositories listed
-
Measure Twice, Cut Once: Grasping Video Structures and Event Semantics with LLMs for Video Temporal Localization12 Mar 2025 0 repositories listed
-
Memory-enhanced Retrieval Augmentation for Long Video Understanding12 Mar 2025 0 repositories listed
-
On the Limitations of Vision-Language Models in Understanding Image Transforms12 Mar 2025 0 repositories listed
-
Reasoning is All You Need for Video Generalization: A Counterfactual Benchmark with Sub-question Evaluation12 Mar 2025 0 repositories listed
-
VideoScan: Enabling Efficient Streaming Video Understanding via Frame-level Semantic Carriers12 Mar 2025 0 repositories listed
-
ALLVB: All-in-One Long Video Understanding Benchmark10 Mar 2025 0 repositories listed
-
BEARCUBS: A benchmark for computer-using web agents10 Mar 2025 0 repositories listed
-
Towards Fine-Grained Video Question Answering10 Mar 2025 0 repositories listed
-
Towards Visual Discrimination and Reasoning of Real-World Physical Dynamics: Physics-Grounded Anomaly Detection5 Mar 2025 0 repositories listed
-
PreMind: Multi-Agent Video Understanding for Advanced Indexing of Presentation-style Videos28 Feb 2025 0 repositories listed
-
HAIC: Improving Human Action Understanding and Generation with Better Captions for Multi-modal Large Language Models28 Feb 2025 0 repositories listed
-
M-LLM Based Video Frame Selection for Efficient Video Understanding27 Feb 2025 0 repositories listed
-
InternVQA: Advancing Compressed Video Quality Assessment with Distilling Large Foundation Model26 Feb 2025 0 repositories listed
-
An Analysis of Data Transformation Effects on Segment Anything 225 Feb 2025 0 repositories listed
-
Fine-Grained Video Captioning through Scene Graph Consolidation23 Feb 2025 0 repositories listed
-
LongCaptioning: Unlocking the Power of Long Caption Generation in Large Multimodal Models21 Feb 2025 0 repositories listed
-
AVD2: Accident Video Diffusion for Accident Video Description20 Feb 2025 0 repositories listed
-
MomentSeeker: A Task-Oriented Benchmark For Long-Video Moment Retrieval18 Feb 2025 0 repositories listed
-
iMOVE: Instance-Motion-Aware Video Understanding17 Feb 2025 0 repositories listed
-
15 Feb 2025 0 repositories listed
-
Optimizing GPT for Video Understanding: Zero-Shot Performance and Prompt Engineering13 Feb 2025 0 repositories listed
-
A Survey on Mamba Architecture for Vision Applications11 Feb 2025 0 repositories listed
-
Enhancing Video Understanding: Deep Neural Networks for Spatiotemporal Analysis11 Feb 2025 0 repositories listed
-
A Survey on Video Analytics in Cloud-Edge-Terminal Collaborative Systems10 Feb 2025 0 repositories listed
-
CoS: Chain-of-Shot Prompting for Long Video Understanding10 Feb 2025 0 repositories listed
-
WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs6 Feb 2025 0 repositories listed
-
A Decade of Action Quality Assessment: Largest Systematic Survey of Trends, Challenges, and Future Directions5 Feb 2025 0 repositories listed
-
MaxInfo: A Training-Free Key-Frame Selection Method Using Maximum Volume for Enhanced Video Understanding5 Feb 2025 0 repositories listed
-
LV-XAttn: Distributed Cross-Attention for Long Visual Inputs in Multimodal Large Language Models4 Feb 2025 0 repositories listed
-
Exploring the Role of Explicit Temporal Modeling in Multimodal Large Language Models for Video Understanding28 Jan 2025 0 repositories listed
-
Understanding Long Videos via LLM-Powered Entity Relation Graphs27 Jan 2025 0 repositories listed
-
HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding25 Jan 2025 0 repositories listed
-
Temporal Preference Optimization for Long-Form Video Understanding23 Jan 2025 0 repositories listed
-
HFGCN:Hypergraph Fusion Graph Convolutional Networks for Skeleton-Based Action Recognition19 Jan 2025 0 repositories listed
-
Omni-RGPT: Unifying Image and Video Region-level Understanding via Token Marks14 Jan 2025 0 repositories listed
-
Video Quality Assessment for Online Processing: From Spatial to Temporal Sampling13 Jan 2025 0 repositories listed
-
X-LeBench: A Benchmark for Extremely Long Egocentric Video Understanding12 Jan 2025 0 repositories listed
-
Zero-shot Shark Tracking and Biometrics from Aerial Imagery10 Jan 2025 0 repositories listed
-
LLaVA-Octopus: Unlocking Instruction-Driven Adaptive Projector Fusion for Video Understanding9 Jan 2025 0 repositories listed
-
LongViTU: Instruction Tuning for Long-Form Video Understanding9 Jan 2025 0 repositories listed
-
Building a Mind Palace: Structuring Environment-Grounded Semantic Graphs for Effective Long Video Analysis with LLMs8 Jan 2025 0 repositories listed
-
H-MBA: Hierarchical MamBa Adaptation for Multi-Modal Video Understanding in Autonomous Driving8 Jan 2025 0 repositories listed
-
MotionBench: Benchmarking and Improving Fine-grained Video Motion Understanding for Vision Language Models6 Jan 2025 0 repositories listed
Syntology lines on 1 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced; each line links to that paper's sample list. Syntology's record for this page has not changed since , the first build that kept a record date for it; when this build read Syntology's graph is in the build record. For agents: get_harvested_code_for_paper(arxiv_id) lists each paper's samples; how to connect.