Browse State-of-the-Art › Natural Language Visual Grounding
Natural Language Visual Grounding
30 papers with code · 1 benchmark · 7 datasets archive 2025-07-28
Benchmarks archive 2025-07-28
1 leaderboard table shown for this task, 1 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted.
| Dataset | Best model (first row in archive order) | Paper | Code | Syntology | Compare |
|---|---|---|---|---|---|
| ScreenSpot (18 rows) | UGround-V1-7B | Navigating the Digital World as Humans Do: Universal Visual... | code | Syntology ran 0 of 1 samples · 1 unverified | Compare |
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
7 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
No subtask under this task in the archive's task tree.
Most implemented papers archive 2025-07-28
30 shown of 30 papers with code (32 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
3 Dec 2019 11 repositories listed Syntology ran 0 of 5 samples · 5 unverifiedWe present ALFRED (Action Learning From Realistic Environments and Directives), a benchmark for learning a mapping from natural language instructions and egocentric vision to sequences of actions for household tasks.
-
18 Sep 2024 8 repositories listed Syntology ran 8 of 12 samples · 4 unverifiedWe present the Qwen2-VL Series, an advanced upgrade of the previous Qwen-VL models that redefines the conventional predetermined-resolution approach in visual processing.
-
14 Dec 2023 3 repositories listed Syntology ran 12 of 18 samples · 6 unverifiedPeople are spending an enormous amount of time on digital devices through graphical user interfaces (GUIs), e.
-
12 Nov 2015 3 repositories listedWe propose a novel approach which learns grounding by reconstructing a given phrase using an attention mechanism, which can be either latent or optimized directly.
-
30 Oct 2024 2 repositories listed Syntology ran 1 of 1 samples · 0 unverifiedExisting efforts in building GUI agents heavily rely on the availability of robust commercial Vision-Language Models (VLMs) such as GPT-4o and GeminiProVision.
-
14 Oct 2023 2 repositories listed Syntology ran 3 of 3 samples · 0 unverified · 1 pointer-only (licence)Motivated by this, we target to build a unified interface for completing many vision-language tasks including image description, visual question answering, and visual grounding, among others.
-
Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond24 Aug 2023 2 repositories listed Syntology ran 0 of 2 samples · 2 unverified · 2 pointer-only (licence)In this work, we introduce the Qwen-VL series, a set of large-scale vision-language models (LVLMs) designed to perceive and understand both texts and images.
-
16 Feb 2021 2 repositories listedControlling robots to perform tasks via natural language is one of the most challenging topics in human-robot interaction.
-
8 Oct 2020 2 repositories listed Syntology ran 11 of 14 samples · 3 unverifiedALFWorld enables the creation of a new BUTLER agent whose abstract knowledge, learned in TextWorld, corresponds directly to concrete, visually grounded actions.
-
10 Jan 2019 2 repositories listed Syntology ran 0 of 6 samples · 6 unverifiedThe Vision-and-Language Navigation (VLN) task entails an agent following navigational instruction in photo-realistic unknown environments.
-
20 Dec 2024 1 repository listed Syntology ran 2 of 2 samples · 0 unverified · 2 pointer-only (licence)Digital agents for automating tasks across different platforms by directly manipulating the GUIs are increasingly important.
-
5 Dec 2024 1 repository listedAutomating GUI tasks remains challenging due to reliance on textual representations, platform-specific action spaces, and limited reasoning capabilities.
-
26 Nov 2024 1 repository listed Syntology ran 1 of 16 samples · 15 unverifiedIn this work, we develop a vision-language-action model in digital world, namely ShowUI, which features the following innovations: (i) UI-Guided Visual Token Selection to reduce computational costs by formulating…
-
18 Nov 2024 1 repository listedGraphical User Interface (GUI) grounding plays a crucial role in enhancing the capabilities of Vision-Language Model (VLM) agents.
-
7 Oct 2024 1 repository listed Syntology ran 0 of 1 samples · 1 unverifiedThe key is visual grounding models that can accurately map diverse referring expressions of GUI elements to their coordinates on the GUI across different platforms.
-
1 Aug 2024 1 repository listedThe recent success of large vision language models shows great potential in driving the agent system operating on user interfaces.
-
17 Jun 2024 1 repository listed Syntology ran 15 of 16 samples · 1 unverified · 16 pointer-only (licence)Utilizing Graphic User Interface (GUI) for human-computer interaction is essential for accessing a wide range of digital tools.
-
19 Apr 2024 1 repository listed Syntology ran 5 of 9 samples · 4 unverifiedWe introduce Groma, a Multimodal Large Language Model (MLLM) with grounded and fine-grained visual perception ability.
-
17 Jan 2024 1 repository listed Syntology ran 1 of 1 samples · 0 unverifiedIn our preliminary study, we have discovered a key challenge in developing visual GUI agents: GUI grounding -- the capacity to accurately locate screen elements based on instructions.
-
26 Feb 2023 1 repository listedIn this paper, we propose a method for improving the performance of natural language grounding in long videos by identifying and pruning out non-describable windows.
-
16 Sep 2022 1 repository listedIn this work, we focus on improving the captions generated by image-caption generation systems.
-
30 Mar 2022 1 repository listed Syntology ran 3 of 3 samples · 0 unverifiedWe consider the problem of localizing a spatio-temporal tube in a video corresponding to a given text query.
-
6 Dec 2021 1 repository listed Syntology ran 6 of 10 samples · 4 unverifiedWe show that a baseline model based on multi-context imitation learning performs poorly on CALVIN, suggesting that there is significant room for developing innovative agents that learn to relate human language to their…
-
10 Sep 2021 1 repository listedThis paper proposes Panoptic Narrative Grounding, a spatially fine and general formulation of the natural language visual grounding problem.
-
1 Jan 2021 1 repository listedThis paper proposes Panoptic Narrative Grounding, a spatially fine and general formulation of the natural language visual grounding problem.
-
7 Oct 2020 1 repository listedRecent models achieve promising results in visually grounded dialogues.
-
13 Feb 2020 1 repository listedTo address their limitations, this paper proposes a language-guided graph representation to capture the global context of grounding entities and their relations, and develop a cross-modal graph matching strategy for the…
-
3 Aug 2019 1 repository listedEspecially in ambiguous settings, humans prefer expressions (called relational referring expressions) that describe an object with respect to a distinguishing, unique object.
-
7 Apr 2019 1 repository listedComputer Vision applications often require a textual grounding module with precision, interpretability, and resilience to counterfactual inputs/queries.
-
8 Jan 2019 1 repository listedWe present a novel Dual Dynamic Attention Model (DUDA) to perform robust Change Captioning.
Syntology lines on 16 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections