Papers › Tuning Large Multimodal Models for Videos using Reinforcement Learning from AI Feedback
Tuning Large Multimodal Models for Videos using Reinforcement Learning from AI Feedback
Daechul Ahn, Yura Choi, Youngjae Yu, Dongyeop Kang, Jonghyun Choi
Recent advancements in large language models have influenced the development of video large multimodal models (VLMMs). The previous approaches for VLMMs involved Supervised Fine-Tuning (SFT) with instruction-tuned datasets, integrating LLM with visual encoders, and adding additional learnable modules. Video and text multimodal alignment remains challenging, primarily due to the deficient volume and quality of multimodal instruction-tune data compared to text-only data. We present a novel alignment strategy that employs multimodal AI system to oversee itself called Reinforcement Learning from AI Feedback (RLAIF), providing self-preference feedback to refine itself and facilitating the alignment of video and text modalities. In specific, we propose context-aware reward modeling by providing detailed video descriptions as context during the generation of preference feedback in order to enrich the understanding of video content. Demonstrating enhanced performance across diverse video benchmarks, our multimodal RLAIF approach, VLM-RLAIF, outperforms existing approaches, including the SFT model. We commit to open-sourcing our code, models, and datasets to foster further research in this area.
In Syntology View this paper on Syntology: its repositories, every harvested function with whether it ran, its licence and the call to fetch it.
Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.
Code
Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.
Code Syntology ran Syntology
Not run by Syntology. Nothing on this page verifies that the listed code works.
Tasks
Results from the paper archive 2025-07-28
| Task | Dataset | Model | Metric | Value | Rank at snapshot | Leaderboard | Report |
|---|---|---|---|---|---|---|---|
| Video-based Generative Performance Benchmarking | VideoInstruct | VLM-RLAIF | Consistency | 3.32 | #2 of 23 | Archive leaderboard | report |
| Video-based Generative Performance Benchmarking | VideoInstruct | VLM-RLAIF | Contextual Understanding | 4 | #2 of 23 | Archive leaderboard | report |
| Video-based Generative Performance Benchmarking | VideoInstruct | VLM-RLAIF | Correctness of Information | 3.63 | #2 of 23 | Archive leaderboard | report |
| Video-based Generative Performance Benchmarking | VideoInstruct | VLM-RLAIF | Detail Orientation | 3.25 | #2 of 23 | Archive leaderboard | report |
| Video-based Generative Performance Benchmarking | VideoInstruct | VLM-RLAIF | Temporal Understanding | 3.23 | #2 of 23 | Archive leaderboard | report |
| Video-based Generative Performance Benchmarking | VideoInstruct | VLM-RLAIF | mean | 3.49 | #2 of 23 | Archive leaderboard | report |
Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.
Methods
Introduced by this paper: RLAIF
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections