Browse State-of-the-Art › Visual Grounding › Papers, page 4
Visual Grounding
Papers archive 2025-07-28
archive papers tagged: 571 · with a code link: 299 · where Syntology ran a sample: 111 (95 with a run with no instrument failure, 16 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (111 of 571 tagged: 95 with a run with no instrument failure, 16 where every run was a failure of Syntology's instrument)
Page 4 of 6: papers 301 to 400 of 571, in archive order: by repositories listed in the archive (most first), then newest first, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers that list no repository come after every paper that lists one.
Papers without a page here are shown as plain text. A Syntology line reads “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the figure “where Syntology's instrument failed” counts failures of Syntology's instrument, not of the code. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code); hover it for the repositories the samples that ran came from. Abstracts are on each paper's page.
-
A Neural Representation Framework with LLM-Driven Spatial Reasoning for Open-Vocabulary 3D Visual Grounding9 Jul 2025 0 repositories listed
-
VisualTrap: A Stealthy Backdoor Attack on GUI Agents via Visual Grounding Manipulation9 Jul 2025 0 repositories listed
-
SPAZER: Spatial-Semantic Progressive Reasoning Agent for Zero-shot 3D Visual Grounding27 Jun 2025 0 repositories listed
-
GroundFlow: A Plug-in Module for Temporal Reasoning on 3D Point Cloud Sequential Grounding26 Jun 2025 0 repositories listed
-
HalluSegBench: Counterfactual Visual Reasoning for Segmentation Hallucination Evaluation26 Jun 2025 0 repositories listed
-
GEMeX-ThinkVG: Towards Thinking with Visual Grounding in Medical VQA via Reinforcement Learning22 Jun 2025 0 repositories listed
-
I Speak and You Find: Robust 3D Visual Grounding with Noisy and Ambiguous Speech Inputs17 Jun 2025 0 repositories listed
-
Unified Representation Space for 3D Visual Grounding17 Jun 2025 0 repositories listed
-
Semantic Localization Guiding Segment Anything Model For Reference Remote Sensing Image Segmentation12 Jun 2025 0 repositories listed
-
EconWebArena: Benchmarking Autonomous Agents on Economic Tasks in Realistic Web Environments9 Jun 2025 0 repositories listed
-
Does Your 3D Encoder Really Work? When Pretrain-SFT from 2D VLMs Meets 3D VLMs5 Jun 2025 0 repositories listed
-
From Objects to Anywhere: A Holistic Benchmark for Multi-level Visual Grounding in 3D Scenes5 Jun 2025 0 repositories listed
-
Perceptual Decoupling for Scalable Multi-modal Reasoning via Reward-Optimized Captioning5 Jun 2025 0 repositories listed
-
RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought4 Jun 2025 0 repositories listed
-
GUI-Actor: Coordinate-Free Visual Grounding for GUI Agents3 Jun 2025 0 repositories listed
-
MoDA: Modulation Adapter for Fine-Grained Visual Grounding in Instructional MLLMs2 Jun 2025 0 repositories listed
-
D2AF: A Dual-Driven Annotation and Filtering Framework for Visual Grounding30 May 2025 0 repositories listed
-
mRAG: Elucidating the Design Space of Multi-modal Retrieval-Augmented Generation29 May 2025 0 repositories listed
-
Zero-Shot 3D Visual Grounding from Vision-Language Models28 May 2025 0 repositories listed
-
Mitigating Hallucination in Large Vision-Language Models via Adaptive Attention Calibration27 May 2025 0 repositories listed
-
Two Causally Related Needles in a Video Haystack26 May 2025 0 repositories listed
-
Don't Look Only Once: Towards Multimodal Interactive Reasoning with Selective Visual Revisitation24 May 2025 0 repositories listed
-
More Thinking, Less Seeing? Assessing Amplified Hallucination in Multimodal Reasoning Models23 May 2025 0 repositories listed
-
Redemption Score: An Evaluation Framework to Rank Image Captions While Redeeming Image Semantics and Language Pragmatics22 May 2025 0 repositories listed
-
Training-Free Reasoning and Reflection in MLLMs22 May 2025 0 repositories listed
-
Seeing the Trees for the Forest: Rethinking Weakly-Supervised Medical Visual Grounding21 May 2025 0 repositories listed
-
UniVG-R1: Reasoning Guided Universal Visual Grounding with Reinforcement Learning20 May 2025 0 repositories listed
-
MedSG-Bench: A Benchmark for Medical Image Sequences Grounding17 May 2025 0 repositories listed
-
TinyRS-R1: Compact Multimodal Language Model for Remote Sensing17 May 2025 0 repositories listed
-
DenseGrounding: Improving Dense Language-Vision Semantics for Ego-Centric 3D Visual Grounding8 May 2025 0 repositories listed
-
3DWG: 3D Weakly Supervised Visual Grounding via Category and Instance-Level Alignment3 May 2025 0 repositories listed
-
VIST-GPT: Ushering in the Era of Visual Storytelling with LLMs?27 Apr 2025 0 repositories listed
-
Revisiting Data Auditing in Large Vision-Language Models25 Apr 2025 0 repositories listed
-
Visual Intention Grounding for Egocentric Assistants18 Apr 2025 0 repositories listed
-
COUNTS: Benchmarking Object Detectors and Multimodal Large Language Models under Distribution Shifts14 Apr 2025 0 repositories listed
-
DSM: Building A Diverse Semantic Map for 3D Visual Grounding11 Apr 2025 0 repositories listed
-
AerialVG: A Challenging Benchmark for Aerial Visual Grounding by Exploring Positional Relations10 Apr 2025 0 repositories listed
-
Towards Visual Text Grounding of Multimodal Large Language Model7 Apr 2025 0 repositories listed
-
Image Difference Grounding with Natural Language2 Apr 2025 0 repositories listed
-
Multimodal Reference Visual Grounding2 Apr 2025 0 repositories listed
-
Towards Unified Referring Expression Segmentation Across Omni-Level Visual Target Granularities2 Apr 2025 0 repositories listed
-
ReasonGrounder: LVLM-Guided Hierarchical Feature Splatting for Open-Vocabulary 3D Visual Grounding and Reasoning30 Mar 2025 0 repositories listed
-
Efficient Adaptation For Remote Sensing Visual Grounding29 Mar 2025 0 repositories listed
-
NuGrounding: A Multi-View 3D Visual Grounding Framework in Autonomous Driving28 Mar 2025 0 repositories listed
-
Beyond Object Categories: Multi-Attribute Reference Understanding for Visual Grounding25 Mar 2025 0 repositories listed
-
Seeing Speech and Sound: Distinguishing and Locating Audios in Visual Scenes24 Mar 2025 0 repositories listed
-
A Vision Centric Remote Sensing Benchmark20 Mar 2025 0 repositories listed
-
Your Large Vision-Language Model Only Needs A Few Attention Heads For Visual Grounding8 Mar 2025 0 repositories listed
-
Enhancing Abnormality Grounding for Vision Language Models with Knowledge Descriptions5 Mar 2025 0 repositories listed
-
Teaching Metric Distance to Autoregressive Multimodal Foundational Models4 Mar 2025 0 repositories listed
-
Structured Preference Optimization for Vision-Language Long-Horizon Task Planning28 Feb 2025 0 repositories listed
-
ProxyTransformation: Preshaping Point Cloud Manifold With Proxy Attention For 3D Visual Grounding26 Feb 2025 0 repositories listed
-
Programming with Pixels: Computer-Use Meets Software Engineering24 Feb 2025 0 repositories listed
-
19 Feb 2025 0 repositories listed
-
Leveraging Multimodal-LLMs Assisted by Instance Segmentation for Intelligent Traffic Monitoring16 Feb 2025 0 repositories listed
-
TRAVEL: Training-Free Retrieval and Alignment for Vision-and-Language Navigation11 Feb 2025 0 repositories listed
-
RLS3: RL-Based Synthetic Sample Selection to Enhance Spatial Reasoning in Vision-Language Models for Indoor Autonomous Perception31 Jan 2025 0 repositories listed
-
24 Jan 2025 0 repositories listed
-
FLORA: Formal Language Model Enables Robust Training-free Zero-shot Object Referring Analysis17 Jan 2025 0 repositories listed
-
AugRefer: Advancing 3D Visual Grounding via Cross-Modal Augmentation and Spatial Relation-based Referring16 Jan 2025 0 repositories listed
-
GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing12 Jan 2025 0 repositories listed
-
EAGLE: Enhanced Visual Grounding Minimizes Hallucinations in Instructional Multimodal Models6 Jan 2025 0 repositories listed
-
ViGiL3D: A Linguistically Diverse Dataset for 3D Visual Grounding2 Jan 2025 0 repositories listed
-
Beyond Human Perception: Understanding Multi-Object World from Monocular View1 Jan 2025 0 repositories listed
-
Seeing Speech and Sound: Distinguishing and Locating Audio Sources in Visual Scenes1 Jan 2025 0 repositories listed
-
Task-aware Cross-modal Feature Refinement Transformer with Large Language Models for Visual Grounding1 Jan 2025 0 repositories listed
-
VideoGLaMM : A Large Multimodal Model for Pixel-Level Visual Grounding in Videos1 Jan 2025 0 repositories listed
-
Referencing Where to Focus: Improving VisualGrounding with Referential Query26 Dec 2024 0 repositories listed
-
EarthDial: Turning Multi-sensory Earth Observations to Interactive Dialogues19 Dec 2024 0 repositories listed
-
FiVL: A Framework for Improved Vision-Language Alignment19 Dec 2024 0 repositories listed
-
GAGS: Granularity-Aware Feature Distillation for Language Gaussian Splatting18 Dec 2024 0 repositories listed
-
Barking Up The Syntactic Tree: Enhancing VLM Training with Syntactic Losses11 Dec 2024 0 repositories listed
-
3D Spatial Understanding in MLLMs: Disambiguation and Evaluation9 Dec 2024 0 repositories listed
-
SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding5 Dec 2024 0 repositories listed
-
Paint Outside the Box: Synthesizing and Selecting Training Data for Visual Grounding1 Dec 2024 0 repositories listed
-
3D Scene Graph Guided Vision-Language Pre-training27 Nov 2024 0 repositories listed
-
Motion-Grounded Video Reasoning: Understanding and Perceiving Motion at Pixel Level15 Nov 2024 0 repositories listed
-
LidaRefer: Outdoor 3D Visual Grounding for Autonomous Driving with Transformers7 Nov 2024 0 repositories listed
-
VideoGLaMM: A Large Multimodal Model for Pixel-Level Visual Grounding in Videos7 Nov 2024 0 repositories listed
-
Fine-Grained Spatial and Verbal Losses for 3D Visual Grounding5 Nov 2024 0 repositories listed
-
Parameter-Efficient Fine-Tuning Medical Multimodal Large Language Models for Medical Visual Grounding31 Oct 2024 0 repositories listed
-
Joint Top-Down and Bottom-Up Frameworks for 3D Visual Grounding21 Oct 2024 0 repositories listed
-
Learning to Ground VLMs without Forgetting14 Oct 2024 0 repositories listed
-
Neural Material Adaptor for Visual Grounding of Intrinsic Dynamics10 Oct 2024 0 repositories listed
-
GRAPPA: Generalizing and Adapting Robot Policies via Online Agentic Guidance9 Oct 2024 0 repositories listed
-
Context-Aware Command Understanding for Tabletop Scenarios8 Oct 2024 0 repositories listed
-
VLM2Vec: Training Vision-Language Models for Massive Multimodal Embedding Tasks7 Oct 2024 0 repositories listed
-
Adaptive Masking Enhances Visual Grounding4 Oct 2024 0 repositories listed
-
Individuation in Neural Models with and without Visual Grounding27 Sep 2024 0 repositories listed
-
ZALM3: Zero-Shot Enhancement of Vision-Language Alignment via In-Context Information in Multi-Turn Multimodal Medical Dialogue26 Sep 2024 0 repositories listed
-
Bayesian Self-Training for Semi-Supervised 3D Segmentation12 Sep 2024 0 repositories listed
-
Visual Prompting in Multimodal Large Language Models: A Survey5 Sep 2024 0 repositories listed
-
NanoMVG: USV-Centric Low-Power Multi-Task Visual Grounding based on Prompt-Guided Camera and 4D mmWave Radar30 Aug 2024 0 repositories listed
-
M4CXR: Exploring Multi-task Potentials of Multi-modal Large Language Models for Chest X-ray Interpretation29 Aug 2024 0 repositories listed
-
26 Aug 2024 0 repositories listed
-
Polaris: Open-ended Interactive Robotic Manipulation via Syn2Real Visual Grounding and Large Language Models15 Aug 2024 0 repositories listed
-
Task-oriented Sequential Grounding in 3D Scenes7 Aug 2024 0 repositories listed
-
UOUO: Uncontextualized Uncommon Objects for Measuring Knowledge Horizons of Vision Language Models25 Jul 2024 0 repositories listed
-
Unveiling and Mitigating Bias in Audio Visual Segmentation23 Jul 2024 0 repositories listed
-
PD-APE: A Parallel Decoding Framework with Adaptive Position Encoding for 3D Visual Grounding19 Jul 2024 0 repositories listed