Papers › ReLaX-VQA: Residual Fragment and Layer Stack Extraction for Enhancing Video Quality Assessment

ReLaX-VQA: Residual Fragment and Layer Stack Extraction for Enhancing Video Quality Assessment

16 Jul 2024arXiv:2407.11496archive 2025-07-28

Xinyi Wang, Angeliki Katsenou, David Bull

With the rapid growth of User-Generated Content (UGC) exchanged between users and sharing platforms, the need for video quality assessment in the wild is increasingly evident. UGC is typically acquired using consumer devices and undergoes multiple rounds of compression (transcoding) before reaching the end user. Therefore, traditional quality metrics that employ the original content as a reference are not suitable. In this paper, we propose ReLaX-VQA, a novel No-Reference Video Quality Assessment (NR-VQA) model that aims to address the challenges of evaluating the quality of diverse video content without reference to the original uncompressed videos. ReLaX-VQA uses frame differences to select spatio-temporal fragments intelligently together with different expressions of spatial features associated with the sampled frames. These are then used to better capture spatial and temporal variabilities in the quality of neighbouring frames. Furthermore, the model enhances abstraction by employing layer-stacking techniques in deep neural network features from Residual Networks and Vision Transformers. Extensive testing across four UGC datasets demonstrates that ReLaX-VQA consistently outperforms existing NR-VQA methods, achieving an average SRCC of 0.8658 and PLCC of 0.8873. Open-source code and trained models that will facilitate further research and applications of NR-VQA can be found at https://github.com/xinyiW915/ReLaX-VQA.

PaperPDFCode

Code

xinyiw915/relax-vqa officialmentioned in papermentioned on GitHubpytorch report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Optical Flow EstimationVideo CompressionVideo Quality AssessmentVisual Question Answering (VQA)

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Video Quality Assessment KoNViD-1k ReLaX-VQA (finetuned on KoNViD-1k) PLCC 0.8668 #5 of 21 Archive leaderboard report
Video Quality Assessment KoNViD-1k ReLaX-VQA PLCC 0.8473 #11 of 21 Archive leaderboard report
Video Quality Assessment KoNViD-1k ReLaX-VQA (trained on LSVQ only) PLCC 0.8427 #12 of 21 Archive leaderboard report
Video Quality Assessment LIVE-VQC ReLaX-VQA (finetuned on LIVE-VQC) PLCC 0.8876 #1 of 20 Archive leaderboard report
Video Quality Assessment LIVE-VQC ReLaX-VQA (trained on LSVQ only) PLCC 0.8242 #11 of 20 Archive leaderboard report
Video Quality Assessment LIVE-VQC ReLaX-VQA PLCC 0.8079 #14 of 20 Archive leaderboard report
Video Quality Assessment YouTube-UGC ReLaX-VQA (finetuned on YouTube-UGC) PLCC 0.8652 #2 of 17 Archive leaderboard report
Video Quality Assessment YouTube-UGC ReLaX-VQA (trained on LSVQ only) PLCC 0.8354 #8 of 17 Archive leaderboard report
Video Quality Assessment YouTube-UGC ReLaX-VQA PLCC 0.8204 #10 of 17 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections