Browse State-of-the-Art › Video Captioning › Papers, page 3
Video Captioning
Papers archive 2025-07-28
archive papers tagged: 473 · with a code link: 211 · where Syntology ran a sample: 64 (56 with a run with no instrument failure, 8 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (64 of 473 tagged: 56 with a run with no instrument failure, 8 where every run was a failure of Syntology's instrument)
Page 3 of 5: papers 201 to 300 of 473, in archive order: by repositories listed in the archive (most first), then newest first, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers that list no repository come after every paper that lists one.
Papers without a page here are shown as plain text. A Syntology line reads “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the figure “where Syntology's instrument failed” counts failures of Syntology's instrument, not of the code. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code); hover it for the repositories the samples that ran came from. Abstracts are on each paper's page.
-
21 Mar 2018 1 repository listed
-
28 Feb 2018 1 repository listed Syntology official: harvested, nothing ran · 0 ran · 1 unverified (of 1 harvested sample)
-
18 Nov 2017 1 repository listed
-
21 Dec 2016 1 repository listed
-
22 Sep 2016 1 repository listed
-
17 Aug 2016 1 repository listed
-
12 Apr 2016 1 repository listed
-
17 Nov 2015 1 repository listed
-
14 Nov 2015 1 repository listed
-
15 Dec 2014 1 repository listed
-
Dense Video Captioning using Graph-based Sentence Summarization25 Jun 2025 0 repositories listed
-
Show, Tell and Summarize: Dense Video Captioning Using Visual Cue Aided Sentence Summarization25 Jun 2025 0 repositories listed
-
VersaVid-R1: A Versatile Video Understanding and Reasoning Model from Question Answering to Captioning Tasks10 Jun 2025 0 repositories listed
-
ARGUS: Hallucination and Omission Evaluation in Video-LLMs9 Jun 2025 0 repositories listed
-
Temporal Object Captioning for Street Scene Videos from LiDAR Tracks22 May 2025 0 repositories listed
-
TimeSoccer: An End-to-End Multimodal Large Language Model for Soccer Commentary Generation24 Apr 2025 0 repositories listed
-
Describe Anything: Detailed Localized Image and Video Captioning22 Apr 2025 0 repositories listed
-
The Devil is in the Distributions: Explicit Modeling of Scene Content is Key in Zero-Shot Video Captioning31 Mar 2025 0 repositories listed
-
Watch and Learn: Leveraging Expert Knowledge and Language for Surgical Video Understanding14 Mar 2025 0 repositories listed
-
Get In Video: Add Anything You Want to the Video8 Mar 2025 0 repositories listed
-
Fine-Grained Video Captioning through Scene Graph Consolidation23 Feb 2025 0 repositories listed
-
LongCaptioning: Unlocking the Power of Long Caption Generation in Large Multimodal Models21 Feb 2025 0 repositories listed
-
Capturing Rich Behavior Representations: A Dynamic Action Semantic-Aware Graph Transformer for Video Captioning19 Feb 2025 0 repositories listed
-
MAMS: Model-Agnostic Module Selection Framework for Video Captioning30 Jan 2025 0 repositories listed
-
Classifier-Guided Captioning Across Modalities3 Jan 2025 0 repositories listed
-
AdaCM^2: On Understanding Extremely Long-Term Video with Adaptive Cross-Modality Memory Reduction1 Jan 2025 0 repositories listed
-
Event-Equalized Dense Video Captioning1 Jan 2025 0 repositories listed
-
CaReBench: A Fine-Grained Benchmark for Video Captioning and Retrieval31 Dec 2024 0 repositories listed
-
PolySmart @ TRECVid 2024 Video Captioning (VTT)20 Dec 2024 0 repositories listed
-
Exploring Temporal Event Cues for Dense Video Captioning in Cyclic Co-learning16 Dec 2024 0 repositories listed
-
Bridging Vision and Language: Modeling Causality and Temporality in Video Narratives14 Dec 2024 0 repositories listed
-
Agent-based Video Trimming12 Dec 2024 0 repositories listed
-
ViCaS: A Dataset for Combining Holistic and Pixel-level Video Understanding using Captions with Grounded Segmentation12 Dec 2024 0 repositories listed
-
Video LLMs for Temporal Reasoning in Long Videos4 Dec 2024 0 repositories listed
-
Progress-Aware Video Frame Captioning3 Dec 2024 0 repositories listed
-
HyperGLM: HyperGraph for Video Scene Graph Generation and Anticipation27 Nov 2024 0 repositories listed
-
Seq2Time: Sequential Knowledge Transfer for Video LLM Temporal Grounding25 Nov 2024 0 repositories listed
-
FINECAPTION: Compositional Image Captioning Focusing on Wherever You Want at Any Granularity23 Nov 2024 0 repositories listed
-
Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning22 Nov 2024 0 repositories listed
-
AdaCM²: On Understanding Extremely Long-Term Video with Adaptive Cross-Modality Memory Reduction19 Nov 2024 0 repositories listed
-
Multi-Modal interpretable automatic video captioning11 Nov 2024 0 repositories listed
-
SPECTRUM: Semantic Processing and Emotion-informed video-Captioning Through Retrieval and Understanding Modalities4 Nov 2024 0 repositories listed
-
Technical Report for Soccernet 2023 -- Dense Video Captioning31 Oct 2024 0 repositories listed
-
EVC-MF: End-to-end Video Captioning Network with Multi-scale Features22 Oct 2024 0 repositories listed
-
FIOVA: A Multi-Annotator Benchmark for Human-Aligned Video Captioning20 Oct 2024 0 repositories listed
-
It's Just Another Day: Unique Video Captioning by Discriminative Prompting15 Oct 2024 0 repositories listed
-
13 Oct 2024 0 repositories listed
-
Enhancing Multimodal LLM for Detailed and Accurate Video Captioning using Multi-Round Preference Optimization9 Oct 2024 0 repositories listed
-
Fine-grained length controllable video captioning with ordinal embeddings27 Aug 2024 0 repositories listed
-
Wolf: Captioning Everything with a World Summarization Framework26 Jul 2024 0 repositories listed
-
EVLM: An Efficient Vision-Language Model for Visual Understanding19 Jul 2024 0 repositories listed
-
Reexamining Racial Disparities in Automatic Speech Recognition Performance: The Role of Confounding by Provenance19 Jul 2024 0 repositories listed
-
https://arxiv.org/abs/2407.006342 Jul 2024 0 repositories listed
-
Directed Domain Fine-Tuning: Tailoring Separate Modalities for Specific Training Tasks24 Jun 2024 0 repositories listed
-
GUI Action Narrator: Where and When Did That Action Take Place?19 Jun 2024 0 repositories listed
-
Towards Holistic Language-video Representation: the language model-enhanced MSR-Video to Text Dataset19 Jun 2024 0 repositories listed
-
10 Jun 2024 0 repositories listed
-
Story Generation from Visual Inputs: Techniques, Related Tasks, and Challenges4 Jun 2024 0 repositories listed
-
RETTA: Retrieval-Enhanced Test-Time Adaptation for Zero-Shot Video Captioning11 May 2024 0 repositories listed
-
A Toolchain for Comprehensive Audio/Video Analysis Using Deep Learning Based Multimodal Approach (A use case of riot or violent context detection)2 May 2024 0 repositories listed
-
15 Apr 2024 0 repositories listed
-
DIBS: Enhancing Dense Video Captioning with Unlabeled Videos via Pseudo Boundary Enrichment and Online Refinement3 Apr 2024 0 repositories listed
-
Empowering LLMs with Pseudo-Untrimmed Videos for Audio-Visual Temporal Understanding24 Mar 2024 0 repositories listed
-
Sora as an AGI World Model? A Complete Survey on Text-to-Video Generation8 Mar 2024 0 repositories listed
-
MCF-VC: Mitigate Catastrophic Forgetting in Class-Incremental Learning for Multimodal Video Captioning27 Feb 2024 0 repositories listed
-
Knowledge Guided Entity-aware Video Captioning and A Basketball Benchmark25 Jan 2024 0 repositories listed
-
SnapCap: Efficient Snapshot Compressive Video Captioning10 Jan 2024 0 repositories listed
-
On Scaling Up a Multilingual Vision and Language Model1 Jan 2024 0 repositories listed
-
Retrieval-Augmented Egocentric Video Captioning1 Jan 2024 0 repositories listed
-
Set Prediction Guided by Semantic Concepts for Diverse Video Captioning25 Dec 2023 0 repositories listed
-
SOVC: Subject-Oriented Video Captioning20 Dec 2023 0 repositories listed
-
Attention Based Encoder Decoder Model for Video Captioning in Nepali (2023)12 Dec 2023 0 repositories listed
-
Exo2EgoDVC: Dense Video Captioning of Egocentric Procedural Activities Using Web Instructional Videos28 Nov 2023 0 repositories listed
-
Incorporating granularity bias as the margin into contrastive loss for video captioning25 Nov 2023 0 repositories listed
-
Dense Video Captioning: A Survey of Techniques, Datasets and Evaluation Protocols5 Nov 2023 0 repositories listed
-
Nepali Video Captioning using CNN-RNN Architecture5 Nov 2023 0 repositories listed
-
Learning Interactive Real-World Simulators9 Oct 2023 0 repositories listed
-
5 Oct 2023 0 repositories listed
-
Human-centric Behavior Description in Videos: New Benchmark and Model4 Oct 2023 0 repositories listed
-
Encoder-Decoder Based Long Short-Term Memory (LSTM) Model for Video Captioning2 Oct 2023 0 repositories listed
-
Towards Surveillance Video-and-Language Understanding: New Dataset, Baselines, and Challenges25 Sep 2023 0 repositories listed
-
Collaborative Three-Stream Transformers for Video Captioning18 Sep 2023 0 repositories listed
-
Video Captioning with Aggregated Features Based on Dual Graphs and Gated Fusion13 Aug 2023 0 repositories listed
-
Zero-Shot Dense Video Captioning by Jointly Optimizing Text and Moment5 Jul 2023 0 repositories listed
-
Style-transfer based Speech and Audio-visual Scene Understanding for Robot Action Sequence Acquisition from Videos27 Jun 2023 0 repositories listed
-
Exploring the Role of Audio in Video Captioning21 Jun 2023 0 repositories listed
-
Knowledge Distillation for Efficient Audio-Visual Video Captioning16 Jun 2023 0 repositories listed
-
22 May 2023 0 repositories listed
-
VideoOFA: Two-Stage Pre-Training for Video-to-Text Generation4 May 2023 0 repositories listed
-
TCR: Short Video Title Generation and Cover Selection with Attention Refinement25 Apr 2023 0 repositories listed
-
A Review of Deep Learning for Video Captioning22 Apr 2023 0 repositories listed
-
LASER: A Neuro-Symbolic Framework for Learning Spatial-Temporal Scene Graphs with Weak Supervision15 Apr 2023 0 repositories listed
-
Scalable and Accurate Self-supervised Multimodal Representation Learning without Aligned Video and Text Data4 Apr 2023 0 repositories listed
-
26 Mar 2023 0 repositories listed
-
22 Mar 2023 0 repositories listed
-
Implicit and Explicit Commonsense for Multi-sentence Video Captioning14 Mar 2023 0 repositories listed
-
Models See Hallucinations: Evaluating the Factuality in Video Captioning6 Mar 2023 0 repositories listed
-
STOA-VLP: Spatial-Temporal Modeling of Object and Action for Video-Language Pre-training20 Feb 2023 0 repositories listed
-
Temporal Perceiving Video-Language Pre-training18 Jan 2023 0 repositories listed
Syntology lines on 2 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced; each line links to that paper's sample list. Syntology's record for this page has not changed since , the first build that kept a record date for it; when this build read Syntology's graph is in the build record. For agents: get_harvested_code_for_paper(arxiv_id) lists each paper's samples; how to connect.