Browse State-of-the-Art › Referring expression generation
Referring expression generation
26 papers with code · 2 benchmarks · 4 datasets archive 2025-07-28
Generate referring expressions
Description from the archive archive 2025-07-28.
Benchmarks archive 2025-07-28
2 leaderboard tables shown for this task, 2 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted.
| Dataset | Best model (first row in archive order) | Paper | Code | Syntology | Compare |
|---|---|---|---|---|---|
| ColonINST-v1 (Seen) (17 rows) | ColonGPT (w/ LoRA, w/o extra data) | Frontiers in Intelligent Colonoscopy | code | Syntology ran 0 of 6 samples · 6 unverified | Compare |
| ColonINST-v1 (Unseen) (17 rows) | ColonGPT (w/ LoRA, w/o extra data) | Frontiers in Intelligent Colonoscopy | code | Syntology ran 0 of 6 samples · 6 unverified | Compare |
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
4 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
No subtask under this task in the archive's task tree.
Most implemented papers archive 2025-07-28
26 shown of 26 papers with code (84 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
17 Apr 2023 13 repositories listed Syntology ran 16 of 51 samples · 35 unverifiedInstruction tuning large language models (LLMs) using machine-generated instruction-following data has improved zero-shot capabilities on new tasks, but the idea is less explored in the multimodal field.
-
5 Oct 2023 9 repositories listed Syntology ran 6 of 9 samples · 3 unverified · 8 pointer-only (licence)Large multimodal models (LMM) have recently shown encouraging progress with visual instruction tuning.
-
31 Jul 2016 4 repositories listed Syntology ran 1 of 4 samples · 3 unverifiedHumans refer to objects in their environments all the time, especially in dialogue with other people.
-
27 Mar 2024 2 repositories listed Syntology ran 7 of 8 samples · 1 unverifiedWe try to narrow the gap by mining the potential of VLMs for better performance and any-to-any workflow from three aspects, i.
-
14 Oct 2023 2 repositories listed Syntology ran 3 of 3 samples · 0 unverified · 1 pointer-only (licence)Motivated by this, we target to build a unified interface for completing many vision-language tasks including image description, visual question answering, and visual grounding, among others.
-
26 Jun 2023 2 repositories listedWe introduce Kosmos-2, a Multimodal Large Language Model (MLLM), enabling new capabilities of perceiving object descriptions (e.
-
22 Apr 2025 1 repository listedReferring Expression Generation (REG) is a core task for evaluating the pragmatic competence of vision-language systems, requiring not only accurate semantic grounding but also adherence to principles of cooperative…
-
22 Oct 2024 1 repository listed Syntology ran 0 of 6 samples · 6 unverifiedColonoscopy is currently one of the most sensitive screening methods for colorectal cancer.
-
4 Oct 2024 1 repository listed Syntology ran 2 of 2 samples · 0 unverified · 2 pointer-only (licence)We introduce a task and dataset for referring expression generation and comprehension in multi-agent embodied environments.
-
26 Sep 2024 1 repository listedTo mitigate the tug-of-war problem of multi-modal multi-task optimization in MLLMs, recent advances primarily focus on improving the LLM components, while neglecting the connector that bridges the gap between modalities.
-
9 Sep 2024 1 repository listedSecond, we propose the use of discourse-aware comprehension guiding as part of a generate-and-rerank strategy through which candidate REs generated with our REG model are reranked based on their discourse-dependent…
-
18 Apr 2024 1 repository listedScene context is well known to facilitate humans' perception of visible objects.
-
25 Mar 2024 1 repository listed Syntology ran 7 of 8 samples · 1 unverified · 8 pointer-only (licence)Multi-modal Large Language Models (MLLMs) have demonstrated their ability to perceive objects in still images, but their application in video-related tasks, such as object tracking, remains understudied.
-
5 Mar 2024 1 repository listed Syntology ran 4 of 7 samples · 3 unverifiedMultimodal Large Language Model (MLLMs) leverages Large Language Models as a cognitive framework for diverse visual-language tasks.
-
18 Feb 2024 1 repository listed Syntology ran 4 of 7 samples · 3 unverifiedMultimodal Large Language Models (MLLMs) have demonstrated notable capabilities in general visual understanding and reasoning tasks.
-
28 Dec 2023 1 repository listedWe present MobileVLM, a competent multimodal vision language model (MMVLM) targeted to run on mobile devices.
-
6 Nov 2023 1 repository listed Syntology ran 1 of 1 samples · 0 unverifiedIn this work, we present Grounding LMM (GLaMM), the first model that can generate natural language responses seamlessly intertwined with corresponding object segmentation masks.
-
10 Sep 2023 1 repository listedWe address these concerns by introducing a collaborative image ranking task, a grounded agreement game we call "A Game Of Sorts".
-
19 Aug 2023 1 repository listedReferring Expression Generation (REG) aims to generate unambiguous Referring Expressions (REs) for objects in a visual scene, with a dual task of Referring Expression Comprehension (REC) to locate the referred object.
-
1 Jun 2023 1 repository listedIn this paper, we propose a cost-efficient approach for training a vision-language conversational assistant that can answer open-ended research questions of biomedical images.
-
24 May 2023 1 repository listedNLP tasks are typically defined extensionally through datasets containing example instantiations (e.
-
1 Aug 2021 1 repository listedThis study introduces an enriched version of the E2E dataset, one of the most popular language resources for data-to-text NLG.
-
22 Sep 2019 1 repository listedWe follow the step-by-step approach to neural data-to-text generation we proposed in Moryossef et al (2019), in which the generation process is divided into a text-planning stage followed by a plan-realization stage.
-
4 Sep 2019 1 repository listedReferring Expression Generation (REG) is the task of generating contextually appropriate references to entities.
-
1 Nov 2018 1 repository listedThis paper describes the enrichment of WebNLG corpus (Gardent et al., 2017a, b), with the aim to further extend its usefulness as a resource for evaluating common NLG tasks, including Discourse Ordering, Lexicalization…
-
21 May 2018 1 repository listedTraditionally, Referring Expression Generation (REG) models first decide on the form and then on the content of references to discourse entities in text, typically relying on features such as salience and grammatical…
Syntology lines on 11 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections