Browse State-of-the-Art › Visual Prompting
Visual Prompting
60 papers with code · 0 benchmarks · 0 datasets archive 2025-07-28
Visual Prompting is the task of streamlining computer vision processes by harnessing the power of prompts, inspired by the breakthroughs of text prompting in NLP. This innovative approach involves using a few visual prompts to swiftly convert an unlabeled dataset into a deployed model, significantly reducing development time for both individual projects and enterprise solutions.
Description from the archive archive 2025-07-28.
Benchmarks archive 2025-07-28
No benchmark for this task in the archive.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
No dataset record in the archive lists this task.
Subtasks archive 2025-07-28
No subtask under this task in the archive's task tree.
Parent tasks archive 2025-07-28
Most implemented papers archive 2025-07-28
30 shown of 60 papers with code (127 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
5 Apr 2023 32 repositories listed Syntology ran 8 of 23 samples · 15 unverifiedWe introduce the Segment Anything (SA) project: a new task, model, and dataset for image segmentation.
-
1 Dec 2023 4 repositories listedFurthermore, we present ViP-Bench, a comprehensive benchmark to assess the capability of models in understanding visual prompts across multiple dimensions, enabling future research in this domain.
-
17 Oct 2023 4 repositories listed Syntology ran 0 of 1 samples · 1 unverifiedWe present Set-of-Mark (SoM), a new visual prompting method, to unleash the visual grounding abilities of large multimodal models (LMMs), such as GPT-4V.
-
22 Nov 2023 3 repositories listed Syntology ran 3 of 6 samples · 3 unverified · 6 pointer-only (licence)In-context prompting in large language models (LLMs) has become a prevalent approach to improve zero-shot capabilities, but this idea is less explored in the vision domain.
-
29 Aug 2023 2 repositories listed Syntology ran 2 of 2 samples · 0 unverified · 2 pointer-only (licence)This empirical investigation underscores the need for a nuanced understanding beyond mere accuracy in sparse and quantized settings, thereby paving the way for further exploration in Visual Prompting techniques tailored…
-
29 May 2023 2 repositories listedWe take inspiration from the widely-used pre-training and then prompt tuning protocols in NLP and propose a new visual prompting model, named Explicit Visual Prompting (EVP).
-
12 Oct 2022 2 repositories listed Syntology ran 1 of 1 samples · 0 unverifiedIn this work, we leverage visual prompting (VP) to improve adversarial robustness of a fixed, pre-trained model at testing time.
-
7 Jun 2025 1 repository listedMSD-CoT progressively disentangles image captions to eliminate semantic ambiguity, while RDVP injects spatial constraints into visual prompting and independently samples visual prompts for foreground and background…
-
7 May 2025 1 repository listed Syntology ran 1 of 2 samples · 1 unverified · 2 pointer-only (licence)Vision GNN (ViG) demonstrates superior performance by representing images as graph structures, providing a more natural way to capture irregular semantic patterns beyond traditional grid or sequence-based…
-
5 May 2025 1 repository listedExtensive experiments across various benchmarks demonstrate that TCPA significantly enhances the diversity and discriminative power of the extracted features.
-
25 Mar 2025 1 repository listedTo address the challenges of missing the useful future context, we develop a novel framework, named Online-MMSI-VLM, that leverages two complementary strategies: multi-party conversation forecasting and social-aware…
-
10 Mar 2025 1 repository listedLane topology extraction involves detecting lanes and traffic elements and determining their relationships, a key perception task for mapless autonomous driving.
-
8 Mar 2025 1 repository listedDepth ambiguity is a fundamental challenge in spatial scene understanding, especially in transparent scenes where single-depth estimates fail to capture full 3D structure.
-
2 Feb 2025 1 repository listedVisual prompting has gained popularity as a method for adapting pre-trained models to specific tasks, particularly in the realm of parameter-efficient tuning.
-
26 Jan 2025 1 repository listedThe stories and characters that captivate us as we grow up shape unique fantasy worlds, with images serving as the primary medium for visually experiencing these realms.
-
2 Jan 2025 1 repository listed Syntology ran 1 of 8 samples · 7 unverifiedTo address this, we introduce GPT4Scene, a novel visual prompting paradigm in VLM training and inference that helps build the global-local relationship, significantly improving the 3D spatial understanding of indoor…
-
12 Dec 2024 1 repository listed Syntology ran 1 of 1 samples · 0 unverified · 1 pointer-only (licence)To prevent the loss of discriminative information during state space propagation, SVP employs lightweight selective prompters for token-wise prompt generation, ensuring adaptive activation of the update and forget gates…
-
4 Dec 2024 1 repository listed Syntology ran 1 of 1 samples · 0 unverified · 1 pointer-only (licence)Building upon this pipeline, we proposed Inst-IT, a solution to enhance LMMs in Instance understanding via explicit visual prompt Instruction Tuning.
-
18 Nov 2024 1 repository listedGraphical User Interface (GUI) grounding plays a crucial role in enhancing the capabilities of Vision-Language Model (VLM) agents.
-
29 Oct 2024 1 repository listedAdditionally, we demonstrate that performance when using automated methods can be improved by up to 68% via a finetuning approach.
-
27 Sep 2024 1 repository listedPiVOT proposes a prompt generation network with the pre-trained foundation model CLIP to automatically generate and refine visual prompts, enabling the transfer of foundation model knowledge for tracking.
-
25 Sep 2024 1 repository listed Syntology ran 8 of 12 samples · 4 unverified · 2 pointer-only (licence)To fill this gap, in this work, we propose a new prompting technique named Attention Prompting on Image, which just simply overlays a text-query-guided attention heatmap on the original input image and effectively…
-
3 Sep 2024 1 repository listedAdapting pre-trained models to new tasks can exhibit varying effectiveness across datasets.
-
30 Aug 2024 1 repository listed Syntology ran 1 of 1 samples · 0 unverifiedVideo action localization aims to find the timings of specific actions from a long video.
-
6 Aug 2024 1 repository listedWith growing interest in recent years, medical visual question answering (Med-VQA) has rapidly evolved, with multimodal large language models (MLLMs) emerging as an alternative to classical model architectures.
-
18 Jul 2024 1 repository listedSpecifically, a shared visual encoding method is developed to establish the spatial pattern interpretation relationships between the multi-scale representations of input images and various visual prompts.
-
15 Jul 2024 1 repository listedWe design a visual prompt that directs MLLMs to utilize visualized sensor data alongside the target sensory task descriptions.
-
11 Jul 2024 1 repository listedWe hypothesize that automatic evaluation can be improved by collecting a targeted UI feedback dataset and then using this dataset to enhance the performance of general-purpose LLMs.
-
15 Jun 2024 1 repository listed Syntology ran 8 of 11 samples · 3 unverifiedContinual Test-Time Adaptation (CTTA) seeks to adapt a source pre-trained model to continually changing, unlabeled target domains.
-
12 Jun 2024 1 repository listed Syntology ran 12 of 16 samples · 4 unverifiedVision Transformers (ViTs) have demonstrated remarkable capabilities in learning representations, but their performance is compromised when applied to unseen domains.
Syntology lines on 13 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections