Papers › EVP: Enhanced Visual Perception using Inverse Multi-Attentive Feature Refinement and...

EVP: Enhanced Visual Perception using Inverse Multi-Attentive Feature Refinement and Regularized Image-Text Alignment

13 Dec 2023arXiv:2312.08548archive 2025-07-28

Mykola Lavreniuk, Shariq Farooq Bhat, Matthias Müller, Peter Wonka

This work presents the network architecture EVP (Enhanced Visual Perception). EVP builds on the previous work VPD which paved the way to use the Stable Diffusion network for computer vision tasks. We propose two major enhancements. First, we develop the Inverse Multi-Attentive Feature Refinement (IMAFR) module which enhances feature learning capabilities by aggregating spatial information from higher pyramid levels. Second, we propose a novel image-text alignment module for improved feature extraction of the Stable Diffusion backbone. The resulting architecture is suitable for a wide variety of tasks and we demonstrate its performance in the context of single-image depth estimation with a specialized decoder using classification-based bins and referring segmentation with an off-the-shelf decoder. Comprehensive experiments conducted on established datasets show that EVP achieves state-of-the-art results in single-image depth estimation for indoor (NYU Depth v2, 11.8% RMSE improvement over VPD) and outdoor (KITTI) environments, as well as referring segmentation (RefCOCO, 2.53 IoU improvement over ReLA). The code and pre-trained models are publicly available at https://github.com/Lavreniuk/EVP.

PaperPDFCode

Code

lavreniuk/evp officialmentioned in papermentioned on GitHubpytorch report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

DecoderDepth EstimationMonocular Depth EstimationReferring Expression Segmentation

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Depth Estimation NYU-Depth V2 EVP RMS 0.224 #1 of 17 Archive leaderboard report
Monocular Depth Estimation KITTI Eigen split EVP Delta < 1.25 0.980 #14 of 79 Archive leaderboard report
Monocular Depth Estimation KITTI Eigen split EVP Delta < 1.25^2 0.998 #14 of 79 Archive leaderboard report
Monocular Depth Estimation KITTI Eigen split EVP Delta < 1.25^3 1.000 #14 of 79 Archive leaderboard report
Monocular Depth Estimation KITTI Eigen split EVP RMSE 2.015 #14 of 79 Archive leaderboard report
Monocular Depth Estimation KITTI Eigen split EVP RMSE log 0.073 #14 of 79 Archive leaderboard report
Monocular Depth Estimation KITTI Eigen split EVP Sq Rel 0.136 #14 of 79 Archive leaderboard report
Monocular Depth Estimation KITTI Eigen split EVP absolute relative error 0.048 #14 of 79 Archive leaderboard report
Monocular Depth Estimation NYU-Depth V2 EVP Delta < 1.25 0.976 #16 of 85 Archive leaderboard report
Monocular Depth Estimation NYU-Depth V2 EVP Delta < 1.25^2 0.997 #16 of 85 Archive leaderboard report
Monocular Depth Estimation NYU-Depth V2 EVP Delta < 1.25^3 0.999 #16 of 85 Archive leaderboard report
Monocular Depth Estimation NYU-Depth V2 EVP RMSE 0.224 #16 of 85 Archive leaderboard report
Monocular Depth Estimation NYU-Depth V2 EVP absolute relative error 0.061 #16 of 85 Archive leaderboard report
Monocular Depth Estimation NYU-Depth V2 EVP log 10 0.027 #16 of 85 Archive leaderboard report
Referring Expression Segmentation RefCOCO EVP IoU 77.61 #3 of 4 Archive leaderboard report
Referring Expression Segmentation RefCOCO EVP IoU (%) 77.61 #3 of 4 Archive leaderboard report
Referring Expression Segmentation RefCOCO testA EVP Overall IoU 78.75 #9 of 13 Archive leaderboard report
Referring Expression Segmentation RefCOCO testB EVP Overall IoU 72.94 #9 of 13 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

AttentionDiffusionLinear LayerMulti-Head AttentionTransformer

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections