{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/evp-enhanced-visual-perception-using-inverse","title":"EVP: Enhanced Visual Perception using Inverse Multi-Attentive Feature Refinement and Regularized Image-Text Alignment","arxiv_id":"2312.08548","date":"2023-12-13","proceeding":null,"authors":["Mykola Lavreniuk","Shariq Farooq Bhat","Matthias Müller","Peter Wonka"],"abstract":"This work presents the network architecture EVP (Enhanced Visual Perception). EVP builds on the previous work VPD which paved the way to use the Stable Diffusion network for computer vision tasks. We propose two major enhancements. First, we develop the Inverse Multi-Attentive Feature Refinement (IMAFR) module which enhances feature learning capabilities by aggregating spatial information from higher pyramid levels. Second, we propose a novel image-text alignment module for improved feature extraction of the Stable Diffusion backbone. The resulting architecture is suitable for a wide variety of tasks and we demonstrate its performance in the context of single-image depth estimation with a specialized decoder using classification-based bins and referring segmentation with an off-the-shelf decoder. Comprehensive experiments conducted on established datasets show that EVP achieves state-of-the-art results in single-image depth estimation for indoor (NYU Depth v2, 11.8% RMSE improvement over VPD) and outdoor (KITTI) environments, as well as referring segmentation (RefCOCO, 2.53 IoU improvement over ReLA). The code and pre-trained models are publicly available at https://github.com/Lavreniuk/EVP.","url_abs":"https://arxiv.org/abs/2312.08548v1","url_pdf":"https://arxiv.org/pdf/2312.08548v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"evp-enhanced-visual-perception-using-inverse","repo_url":"https://github.com/lavreniuk/evp","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"decoder","task_name":"Decoder"},{"task_slug":"depth-estimation","task_name":"Depth Estimation"},{"task_slug":"monocular-depth-estimation","task_name":"Monocular Depth Estimation"},{"task_slug":"referring-expression-segmentation","task_name":"Referring Expression Segmentation"}],"methods":[{"method_slug":"attention","method_name":"Attention"},{"method_slug":"diffusion","method_name":"Diffusion"},{"method_slug":"linear-layer","method_name":"Linear Layer"},{"method_slug":"multi-head-attention","method_name":"Multi-Head Attention"},{"method_slug":"transformer","method_name":"Transformer"}],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/depth-estimation-on-nyu-depth-v2","task":"Depth Estimation","dataset":"NYU-Depth V2","model":"EVP","rank_in_archive_order":1,"of":17,"metrics":{"RMS":"0.224"},"uses_additional_data":false},{"leaderboard":"/sota/monocular-depth-estimation-on-kitti-eigen","task":"Monocular Depth Estimation","dataset":"KITTI Eigen split","model":"EVP","rank_in_archive_order":14,"of":79,"metrics":{"Delta < 1.25":"0.980","Delta < 1.25^2":"0.998","Delta < 1.25^3":"1.000","RMSE":"2.015","RMSE log":"0.073","Sq Rel":"0.136","absolute relative error":"0.048"},"uses_additional_data":false},{"leaderboard":"/sota/monocular-depth-estimation-on-nyu-depth-v2","task":"Monocular Depth Estimation","dataset":"NYU-Depth V2","model":"EVP","rank_in_archive_order":16,"of":85,"metrics":{"Delta < 1.25":"0.976","Delta < 1.25^2":"0.997","Delta < 1.25^3":"0.999","RMSE":"0.224","absolute relative error":"0.061","log 10":"0.027"},"uses_additional_data":false},{"leaderboard":"/sota/referring-expression-segmentation-on-refcoco-6","task":"Referring Expression Segmentation","dataset":"RefCOCO","model":"EVP","rank_in_archive_order":3,"of":4,"metrics":{"IoU":"77.61","IoU (%)":"77.61"},"uses_additional_data":false},{"leaderboard":"/sota/referring-expression-segmentation-on-refcoco-8","task":"Referring Expression Segmentation","dataset":"RefCOCO testA","model":"EVP","rank_in_archive_order":9,"of":13,"metrics":{"Overall IoU":"78.75"},"uses_additional_data":false},{"leaderboard":"/sota/referring-expression-segmentation-on-refcoco-9","task":"Referring Expression Segmentation","dataset":"RefCOCO testB","model":"EVP","rank_in_archive_order":9,"of":13,"metrics":{"Overall IoU":"72.94"},"uses_additional_data":false}],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}