Papers › Tuning Large Multimodal Models for Videos using Reinforcement Learning from AI Feedback

Tuning Large Multimodal Models for Videos using Reinforcement Learning from AI Feedback

6 Feb 2024arXiv:2402.03746archive 2025-07-28

Daechul Ahn, Yura Choi, Youngjae Yu, Dongyeop Kang, Jonghyun Choi

Recent advancements in large language models have influenced the development of video large multimodal models (VLMMs). The previous approaches for VLMMs involved Supervised Fine-Tuning (SFT) with instruction-tuned datasets, integrating LLM with visual encoders, and adding additional learnable modules. Video and text multimodal alignment remains challenging, primarily due to the deficient volume and quality of multimodal instruction-tune data compared to text-only data. We present a novel alignment strategy that employs multimodal AI system to oversee itself called Reinforcement Learning from AI Feedback (RLAIF), providing self-preference feedback to refine itself and facilitating the alignment of video and text modalities. In specific, we propose context-aware reward modeling by providing detailed video descriptions as context during the generation of preference feedback in order to enrich the understanding of video content. Demonstrating enhanced performance across diverse video benchmarks, our multimodal RLAIF approach, VLM-RLAIF, outperforms existing approaches, including the SFT model. We commit to open-sourcing our code, models, and datasets to foster further research in this area.

PaperPDFCode

In Syntology View this paper on Syntology: its repositories, every harvested function with whether it ran, its licence and the call to fetch it.

Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

yonseivnl/vlm-rlaif officialmentioned in papermentioned on GitHubpytorchApache-2.0 report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Video-based Generative Performance Benchmarking

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Video-based Generative Performance Benchmarking VideoInstruct VLM-RLAIF Consistency 3.32 #2 of 23 Archive leaderboard report
Video-based Generative Performance Benchmarking VideoInstruct VLM-RLAIF Contextual Understanding 4 #2 of 23 Archive leaderboard report
Video-based Generative Performance Benchmarking VideoInstruct VLM-RLAIF Correctness of Information 3.63 #2 of 23 Archive leaderboard report
Video-based Generative Performance Benchmarking VideoInstruct VLM-RLAIF Detail Orientation 3.25 #2 of 23 Archive leaderboard report
Video-based Generative Performance Benchmarking VideoInstruct VLM-RLAIF Temporal Understanding 3.23 #2 of 23 Archive leaderboard report
Video-based Generative Performance Benchmarking VideoInstruct VLM-RLAIF mean 3.49 #2 of 23 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

Introduced by this paper: RLAIF

RLAIFSFT

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections