Browse State-of-the-Art › Video Understanding › Papers, page 8
Video Understanding
Papers archive 2025-07-28
archive papers tagged: 1,149 · with a code link: 542 · where Syntology ran a sample: 218 (182 with a run with no instrument failure, 36 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (218 of 1,149 tagged: 182 with a run with no instrument failure, 36 where every run was a failure of Syntology's instrument)
Page 8 of 12: papers 701 to 800 of 1,149, in archive order: by repositories listed in the archive (most first), then newest first, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers that list no repository come after every paper that lists one.
Papers without a page here are shown as plain text. A Syntology line reads “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the figure “where Syntology's instrument failed” counts failures of Syntology's instrument, not of the code. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code); hover it for the repositories the samples that ran came from. Abstracts are on each paper's page.
-
AdaCM^2: On Understanding Extremely Long-Term Video with Adaptive Cross-Modality Memory Reduction1 Jan 2025 0 repositories listed
-
Adapting Pre-trained 3D Models for Point Cloud Video Understanding via Cross-frame Spatio-temporal Perception1 Jan 2025 0 repositories listed
-
Efficient Motion-Aware Video MLLM1 Jan 2025 0 repositories listed
-
Flexible Frame Selection for Efficient Video Reasoning1 Jan 2025 0 repositories listed
-
HierarQ: Task-Aware Hierarchical Q-Former for Enhanced Video Understanding1 Jan 2025 0 repositories listed
-
HuMoCon: Concept Discovery for Human Motion Understanding1 Jan 2025 0 repositories listed
-
Q-Bench-Video: Benchmark the Video Quality Understanding of LMMs1 Jan 2025 0 repositories listed
-
VEU-Bench: Towards Comprehensive Understanding of Video Editing1 Jan 2025 0 repositories listed
-
Video Language Model Pretraining with Spatio-temporal Masking1 Jan 2025 0 repositories listed
-
Weakly Supervised Temporal Action Localization via Dual-Prior Collaborative Learning Guided by Multimodal Large Language Models1 Jan 2025 0 repositories listed
-
Embodied VideoAgent: Persistent Memory from Egocentric Videos and Embodied Sensors Enables Dynamic Scene Understanding31 Dec 2024 0 repositories listed
-
CaReBench: A Fine-Grained Benchmark for Video Captioning and Retrieval31 Dec 2024 0 repositories listed
-
OV-HHIR: Open Vocabulary Human Interaction Recognition Using Cross-modal Integration of Large Language Models31 Dec 2024 0 repositories listed
-
MVTamperBench: Evaluating Robustness of Vision-Language Models27 Dec 2024 0 repositories listed
-
Perceive, Query & Reason: Enhancing Video QA with Question-Guided Temporal Queries26 Dec 2024 0 repositories listed
-
Video Domain Incremental Learning for Human Action Recognition in Home Environments22 Dec 2024 0 repositories listed
-
FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering17 Dec 2024 0 repositories listed
-
ShotVL: Human-Centric Highlight Frame Retrieval via Language Queries17 Dec 2024 0 repositories listed
-
CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding16 Dec 2024 0 repositories listed
-
Overview of TREC 2024 Medical Video Question Answering (MedVidQA) Track15 Dec 2024 0 repositories listed
-
Apollo: An Exploration of Video Understanding in Large Multimodal Models13 Dec 2024 0 repositories listed
-
IQViC: In-context, Question Adaptive Vision Compressor for Long-term Video Understanding LMMs13 Dec 2024 0 repositories listed
-
PVC: Progressive Visual Token Compression for Unified Image and Video Processing in Large Vision-Language Models12 Dec 2024 0 repositories listed
-
VCA: Video Curious Agent for Long Video Understanding12 Dec 2024 0 repositories listed
-
ViCaS: A Dataset for Combining Holistic and Pixel-level Video Understanding using Captions with Grounded Segmentation12 Dec 2024 0 repositories listed
-
COEF-VQ: Cost-Efficient Video Quality Understanding through a Cascaded Multimodal LLM Framework11 Dec 2024 0 repositories listed
-
3DSRBench: A Comprehensive 3D Spatial Reasoning Benchmark10 Dec 2024 0 repositories listed
-
GEXIA: Granularity Expansion and Iterative Approximation for Scalable Multi-grained Video-language Learning10 Dec 2024 0 repositories listed
-
Multi-Scale Contrastive Learning for Video Temporal Grounding10 Dec 2024 0 repositories listed
-
Towards Long Video Understanding via Fine-detailed Video Story Generation9 Dec 2024 0 repositories listed
-
Beyond Boxes: Mask-Guided Spatio-Temporal Feature Aggregation for Video Object Detection6 Dec 2024 0 repositories listed
-
Espresso: High Compression For Rich Extraction From Videos for Your Vision-Language Model6 Dec 2024 0 repositories listed
-
VidHalluc: Evaluating Temporal Hallucinations in Multimodal Large Language Models for Video Understanding4 Dec 2024 0 repositories listed
-
Progress-Aware Video Frame Captioning3 Dec 2024 0 repositories listed
-
SEAL: Semantic Attention Learning for Long Video Representation2 Dec 2024 0 repositories listed
-
VideoSAVi: Self-Aligned Video Language Models without Human Supervision1 Dec 2024 0 repositories listed
-
VISTA: Enhancing Long-Duration and High-Resolution Video Understanding by Video Spatiotemporal Augmentation1 Dec 2024 0 repositories listed
-
Look Every Frame All at Once: Video-Ma²mba for Efficient Long-form Video Understanding with Multi-Axis Gradient Checkpointing29 Nov 2024 0 repositories listed
-
Perception Test 2024: Challenge Summary and a Novel Hour-Long VideoQA Benchmark29 Nov 2024 0 repositories listed
-
STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training29 Nov 2024 0 repositories listed
-
SAVEn-Vid: Synergistic Audio-Visual Integration for Enhanced Understanding in Long Video Context25 Nov 2024 0 repositories listed
-
ReWind: Understanding Long Videos with Instructed Learnable Memory23 Nov 2024 0 repositories listed
-
Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding21 Nov 2024 0 repositories listed
-
20 Nov 2024 0 repositories listed
-
Principles of Visual Tokens for Efficient Video Understanding20 Nov 2024 0 repositories listed
-
VideoAutoArena: An Automated Arena for Evaluating Large Multimodal Models in Video Analysis through User Simulation20 Nov 2024 0 repositories listed
-
AdaCM²: On Understanding Extremely Long-Term Video with Adaptive Cross-Modality Memory Reduction19 Nov 2024 0 repositories listed
-
DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding19 Nov 2024 0 repositories listed
-
ViBe: A Text-to-Video Benchmark for Evaluating Hallucination in Large Multimodal Models16 Nov 2024 0 repositories listed
-
Can MLLMs Guide Weakly-Supervised Temporal Action Localization Tasks?13 Nov 2024 0 repositories listed
-
EVQAScore: Efficient Video Question Answering Data Evaluation11 Nov 2024 0 repositories listed
-
Video RWKV:Video Action Recognition Based RWKV8 Nov 2024 0 repositories listed
-
Personalized Video Summarization by Multimodal Video Understanding5 Nov 2024 0 repositories listed
-
Video Token Merging for Long-form Video Understanding31 Oct 2024 0 repositories listed
-
Zero-Shot Action Recognition in Surveillance Videos28 Oct 2024 0 repositories listed
-
Egocentric and Exocentric Methods: A Short Survey27 Oct 2024 0 repositories listed
-
Adaptive Video Understanding Agent: Enhancing efficiency with dynamic frame sampling and feedback-driven reasoning26 Oct 2024 0 repositories listed
-
FIOVA: A Multi-Annotator Benchmark for Human-Aligned Video Captioning20 Oct 2024 0 repositories listed
-
ContextDet: Temporal Action Detection with Adaptive Context Aggregation20 Oct 2024 0 repositories listed
-
EVA: An Embodied World Model for Future Video Anticipation20 Oct 2024 0 repositories listed
-
Making Every Frame Matter: Continuous Video Understanding for Large Models via Adaptive State Modeling19 Oct 2024 0 repositories listed
-
Zero-shot Action Localization via the Confidence of Large Vision-Language Models18 Oct 2024 0 repositories listed
-
VidCompress: Memory-Enhanced Temporal Compression for Video Understanding in Large Language Models15 Oct 2024 0 repositories listed
-
ViFi-ReID: A Two-Stream Vision-WiFi Multimodal Approach for Person Re-identification13 Oct 2024 0 repositories listed
-
Prompting Video-Language Foundation Models with Domain-specific Fine-grained Heuristics for Video Question Answering12 Oct 2024 0 repositories listed
-
10 Oct 2024 0 repositories listed
-
Enhancing Multimodal LLM for Detailed and Accurate Video Captioning using Multi-Round Preference Optimization9 Oct 2024 0 repositories listed
-
MM-Ego: Towards Building Egocentric Multimodal LLMs9 Oct 2024 0 repositories listed
-
AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark4 Oct 2024 0 repositories listed
-
Frame-Voyager: Learning to Query Frames for Video Large Language Models4 Oct 2024 0 repositories listed
-
3 Oct 2024 0 repositories listed
-
DTVLT: A Multi-modal Diverse Text Benchmark for Visual Language Tracking Based on LLM3 Oct 2024 0 repositories listed
-
Deep learning for action spotting in association football videos2 Oct 2024 0 repositories listed
-
30 Sep 2024 0 repositories listed
-
Q-Bench-Video: Benchmarking the Video Quality Understanding of LMMs30 Sep 2024 0 repositories listed
-
Visual Context Window Extension: A New Perspective for Long Video Understanding30 Sep 2024 0 repositories listed
-
Temporal2Seq: A Unified Framework for Temporal Video Understanding Tasks27 Sep 2024 0 repositories listed
-
EAGLE: Egocentric AGgregated Language-video Engine26 Sep 2024 0 repositories listed
-
LLM4Brain: Training a Large Language Model for Brain Video Understanding26 Sep 2024 0 repositories listed
-
Can CLIP Count Stars? An Empirical Study on Quantity Bias in CLIP23 Sep 2024 0 repositories listed
-
First Place Solution to the Multiple-choice Video QA Track of The Second Perception Test Challenge20 Sep 2024 0 repositories listed
-
Towards Child-Inclusive Clinical Video Understanding for Autism Spectrum Disorder20 Sep 2024 0 repositories listed
-
Interpretable Action Recognition on Hard to Classify Actions19 Sep 2024 0 repositories listed
-
AMEGO: Active Memory from long EGOcentric videos17 Sep 2024 0 repositories listed
-
HAVANA: Hierarchical stochastic neighbor embedding for Accelerated Video ANnotAtions16 Sep 2024 0 repositories listed
-
Enhancing Long Video Understanding via Hierarchical Event-Based Memory10 Sep 2024 0 repositories listed
-
VidLPRO: A Video-Language Pre-training Framework for Robotic and Laparoscopic Surgery7 Sep 2024 0 repositories listed
-
TC-LLaVA: Rethinking the Transfer from Image to Video Understanding with Temporal Considerations5 Sep 2024 0 repositories listed
-
VideoLLaMB: Long-context Video Understanding with Recurrent Memory Bridges2 Sep 2024 0 repositories listed
-
StimuVAR: Spatiotemporal Stimuli-aware Video Affective Reasoning with Multimodal Large Language Models31 Aug 2024 0 repositories listed
-
Streamlining Forest Wildfire Surveillance: AI-Enhanced UAVs Utilizing the FLAME Aerial Video Dataset for Lightweight and Efficient Monitoring31 Aug 2024 0 repositories listed
-
DLM-VMTL:A Double Layer Mapper for heterogeneous data video Multi-task prompt learning29 Aug 2024 0 repositories listed
-
Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input28 Aug 2024 0 repositories listed
-
Attend-Fusion: Efficient Audio-Visual Fusion for Video Classification26 Aug 2024 0 repositories listed
-
LMM-VQA: Advancing Video Quality Assessment with Large Multimodal Models26 Aug 2024 0 repositories listed
-
Flatten: Video Action Recognition is an Image Classification task17 Aug 2024 0 repositories listed
-
Disentangle and denoise: Tackling context misalignment for video moment retrieval14 Aug 2024 0 repositories listed
-
Spherical World-Locking for Audio-Visual Localization in Egocentric Videos9 Aug 2024 0 repositories listed
-
FE-Adapter: Adapting Image-based Emotion Classifiers to Videos5 Aug 2024 0 repositories listed
-
Multimodal Fusion and Coherence Modeling for Video Topic Segmentation1 Aug 2024 0 repositories listed