Papers › NarrativeBridge: Enhancing Video Captioning with Causal-Temporal Narrative

NarrativeBridge: Enhancing Video Captioning with Causal-Temporal Narrative

10 Jun 2024arXiv:2406.06499archive 2025-07-28

Asmar Nadeem, Faegheh Sardari, Robert Dawes, Syed Sameed Husain, Adrian Hilton, Armin Mustafa

Existing video captioning benchmarks and models lack causal-temporal narrative, which is sequences of events linked through cause and effect, unfolding over time and driven by characters or agents. This lack of narrative restricts models' ability to generate text descriptions that capture the causal and temporal dynamics inherent in video content. To address this gap, we propose NarrativeBridge, an approach comprising of: (1) a novel Causal-Temporal Narrative (CTN) captions benchmark generated using a large language model and few-shot prompting, explicitly encoding cause-effect temporal relationships in video descriptions; and (2) a Cause-Effect Network (CEN) with separate encoders for capturing cause and effect dynamics, enabling effective learning and generation of captions with causal-temporal narrative. Extensive experiments demonstrate that CEN significantly outperforms state-of-the-art models in articulating the causal and temporal aspects of video content: 17.88 and 17.44 CIDEr on the MSVD-CTN and MSRVTT-CTN datasets, respectively. Cross-dataset evaluations further showcase CEN's strong generalization capabilities. The proposed framework understands and generates nuanced text descriptions with intricate causal-temporal narrative structures present in videos, addressing a critical limitation in video captioning. For project details, visit https://narrativebridge.github.io/.

PaperPDF

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

No code repository is listed for this paper in the archive or in Syntology's graph.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Language ModellingLarge Language ModelVideo Captioning

Datasets

Introduced by this paper, per the archive.

MSRVTT-CTNMSVD-CTN

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Video Captioning MSRVTT-CTN CEN CIDEr 49.87 #1 of 4 Archive leaderboard report
Video Captioning MSRVTT-CTN CEN ROUGE-L 27.90 #1 of 4 Archive leaderboard report
Video Captioning MSRVTT-CTN CEN SPICE 15.76 #1 of 4 Archive leaderboard report
Video Captioning MSVD-CTN CEN CIDEr 63.51 #1 of 4 Archive leaderboard report
Video Captioning MSVD-CTN CEN ROUGE-L 31.46 #1 of 4 Archive leaderboard report
Video Captioning MSVD-CTN CEN SPICE 19.25 #1 of 4 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections