Papers › No-Reference Image Quality Assessment via Transformers, Relative Ranking, and Self-Consistency

No-Reference Image Quality Assessment via Transformers, Relative Ranking, and Self-Consistency

16 Aug 2021arXiv:2108.06858archive 2025-07-28

S. Alireza Golestaneh, Saba Dadsetan, Kris M. Kitani

The goal of No-Reference Image Quality Assessment (NR-IQA) is to estimate the perceptual image quality in accordance with subjective evaluations, it is a complex and unsolved problem due to the absence of the pristine reference image. In this paper, we propose a novel model to address the NR-IQA task by leveraging a hybrid approach that benefits from Convolutional Neural Networks (CNNs) and self-attention mechanism in Transformers to extract both local and non-local features from the input image. We capture local structure information of the image via CNNs, then to circumvent the locality bias among the extracted CNNs features and obtain a non-local representation of the image, we utilize Transformers on the extracted features where we model them as a sequential input to the Transformer model. Furthermore, to improve the monotonicity correlation between the subjective and objective scores, we utilize the relative distance information among the images within each batch and enforce the relative ranking among them. Last but not least, we observe that the performance of NR-IQA models degrades when we apply equivariant transformations (e.g. horizontal flipping) to the inputs. Therefore, we propose a method that leverages self-consistency as a source of self-supervision to improve the robustness of NRIQA models. Specifically, we enforce self-consistency between the outputs of our quality assessment model for each image and its transformation (horizontally flipped) to utilize the rich self-supervisory information and reduce the uncertainty of the model. To demonstrate the effectiveness of our work, we evaluate it on seven standard IQA datasets (both synthetic and authentic) and show that our model achieves state-of-the-art results on various datasets.

PaperPDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

isalirezag/tres officialmentioned in papermentioned on GitHubpytorchNOASSERTION report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Image Quality AssessmentNo-Reference Image Quality Assessment

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
No-Reference Image Quality Assessment CSIQ TReS PLCC 0.942 #7 of 8 Archive leaderboard report
No-Reference Image Quality Assessment CSIQ TReS SRCC 0.922 #7 of 8 Archive leaderboard report
No-Reference Image Quality Assessment KADID-10k TReS PLCC 0.858 #6 of 9 Archive leaderboard report
No-Reference Image Quality Assessment KADID-10k TReS SRCC 0.859 #6 of 9 Archive leaderboard report
No-Reference Image Quality Assessment TID2013 TReS PLCC 0.883 #3 of 8 Archive leaderboard report
No-Reference Image Quality Assessment TID2013 TReS SRCC 0.863 #3 of 8 Archive leaderboard report
Video Quality Assessment MSU SR-QA Dataset TReS trained on KONIQ KLCC 0.49004 #15 of 60 Archive leaderboard report
Video Quality Assessment MSU SR-QA Dataset TReS trained on KONIQ PLCC 0.56226 #15 of 60 Archive leaderboard report
Video Quality Assessment MSU SR-QA Dataset TReS trained on KONIQ SROCC 0.62578 #15 of 60 Archive leaderboard report
Video Quality Assessment MSU SR-QA Dataset TReS trained on KONIQ Type NR #15 of 60 Archive leaderboard report
Video Quality Assessment MSU SR-QA Dataset TReS KLCC 0.48901 #16 of 60 Archive leaderboard report
Video Quality Assessment MSU SR-QA Dataset TReS PLCC 0.56277 #16 of 60 Archive leaderboard report
Video Quality Assessment MSU SR-QA Dataset TReS SROCC 0.62496 #16 of 60 Archive leaderboard report
Video Quality Assessment MSU SR-QA Dataset TReS Type NR #16 of 60 Archive leaderboard report
Video Quality Assessment MSU SR-QA Dataset TReS trained on FLIVE KLCC 0.39398 #38 of 60 Archive leaderboard report
Video Quality Assessment MSU SR-QA Dataset TReS trained on FLIVE PLCC 0.50005 #38 of 60 Archive leaderboard report
Video Quality Assessment MSU SR-QA Dataset TReS trained on FLIVE SROCC 0.48882 #38 of 60 Archive leaderboard report
Video Quality Assessment MSU SR-QA Dataset TReS trained on FLIVE Type NR #38 of 60 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

Absolute Position EncodingsAdamAttentionBPEDense ConnectionsDropoutLabel SmoothingLayer NormalizationLinear LayerMulti-Head AttentionPosition-Wise Feed-Forward LayerResidual ConnectionSoftmaxTransformer

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections