Browse State-of-the-Art › Visual Grounding › Papers, page 5
Visual Grounding
Papers archive 2025-07-28
archive papers tagged: 571 · with a code link: 299 · where Syntology ran a sample: 111 (95 with a run with no instrument failure, 16 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (111 of 571 tagged: 95 with a run with no instrument failure, 16 where every run was a failure of Syntology's instrument)
Page 5 of 6: papers 401 to 500 of 571, in archive order: by repositories listed in the archive (most first), then newest first, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers that list no repository come after every paper that lists one.
Papers without a page here are shown as plain text. A Syntology line reads “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the figure “where Syntology's instrument failed” counts failures of Syntology's instrument, not of the code. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code); hover it for the repositories the samples that ran came from. Abstracts are on each paper's page.
-
Learning Visual Grounding from Generative Vision and Language Model18 Jul 2024 0 repositories listed
-
Open-Vocabulary 3D Semantic Segmentation with Text-to-Image Diffusion Models18 Jul 2024 0 repositories listed
-
VIMI: Grounding Video Generation through Multi-modal Instruction8 Jul 2024 0 repositories listed
-
Second Place Solution of WSDM2023 Toloka Visual Question Answering Challenge5 Jul 2024 0 repositories listed
-
ACTRESS: Active Retraining for Semi-supervised Visual Grounding3 Jul 2024 0 repositories listed
-
Visual Grounding with Attention-Driven Constraint Balancing3 Jul 2024 0 repositories listed
-
The Solution for the ICCV 2023 Perception Test Challenge 2023 -- Task 6 -- Grounded videoQA2 Jul 2024 0 repositories listed
-
ScanReason: Empowering 3D Visual Grounding with Reasoning Capabilities1 Jul 2024 0 repositories listed
-
From Local Concepts to Universals: Evaluating the Multicultural Understanding of Vision-Language Models28 Jun 2024 0 repositories listed
-
FlowVQA: Mapping Multimodal Logic in Visual Question Answering with Flowcharts27 Jun 2024 0 repositories listed
-
On the Role of Visual Grounding in VQA26 Jun 2024 0 repositories listed
-
Towards Open-World Grasping with Large Vision-Language Models26 Jun 2024 0 repositories listed
-
17 Jun 2024 0 repositories listed Syntology 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
Learning Language Structures through Grounding14 Jun 2024 0 repositories listed
-
Dual Attribute-Spatial Relation Alignment for 3D Visual Grounding13 Jun 2024 0 repositories listed
-
HPE-CogVLM: Advancing Vision Language Models with a Head Pose Grounding Task4 Jun 2024 0 repositories listed
-
HENASY: Learning to Assemble Scene-Entities for Egocentric Video-Language Model1 Jun 2024 0 repositories listed
-
Intent3D: 3D Object Detection in RGB-D Scans Based on Human Intention28 May 2024 0 repositories listed
-
LLM-Optic: Unveiling the Capabilities of Large Language Models for Universal Visual Grounding27 May 2024 0 repositories listed
-
Talk to Parallel LiDARs: A Human-LiDAR Interaction Method Based on 3D Visual Grounding24 May 2024 0 repositories listed
-
Visual grounding for desktop graphical user interfaces5 May 2024 0 repositories listed
-
Naturally Supervised 3D Visual Grounding with Language-Regularized Concept Learners30 Apr 2024 0 repositories listed
-
BlenderAlchemy: Editing 3D Graphics with Vision-Language Models26 Apr 2024 0 repositories listed
-
MedRG: Medical Report Grounding with Multi-modal Large Language Model10 Apr 2024 0 repositories listed
-
Data-Efficient 3D Visual Grounding via Order-Aware Referring25 Mar 2024 0 repositories listed
-
Surgical-LVLM: Learning to Adapt Large Vision-Language Model for Grounded Visual Question Answering in Robotic Surgery22 Mar 2024 0 repositories listed
-
VidLA: Video-Language Alignment at Scale21 Mar 2024 0 repositories listed
-
Learning from Synthetic Data for Visual Grounding20 Mar 2024 0 repositories listed
-
WaterVG: Waterway Visual Grounding based on Text-Guided Vision and mmWave Radar19 Mar 2024 0 repositories listed
-
Right Place, Right Time! Dynamizing Topological Graphs for Embodied Navigation14 Mar 2024 0 repositories listed
-
Detecting Concrete Visual Tokens for Multimodal Machine Translation5 Mar 2024 0 repositories listed
-
Adversarial Testing for Visual Grounding via Image-Aware Property Reduction2 Mar 2024 0 repositories listed
-
OmniACT: A Dataset and Benchmark for Enabling Multimodal Generalist Autonomous Agents for Desktop and Web27 Feb 2024 0 repositories listed
-
Neural Slot Interpreters: Grounding Object Semantics in Emergent Slot Representations2 Feb 2024 0 repositories listed
-
SCO-VIST: Social Interaction Commonsense Knowledge-based Visual Storytelling1 Feb 2024 0 repositories listed
-
LCV2: An Efficient Pretraining-Free Framework for Grounded Visual Question Answering29 Jan 2024 0 repositories listed
-
SceneVerse: Scaling 3D Vision-Language Learning for Grounded Scene Understanding17 Jan 2024 0 repositories listed
-
Towards Truly Zero-shot Compositional Visual Reasoning with LLMs as Programmers3 Jan 2024 0 repositories listed
-
G^3-LQ: Marrying Hyperbolic Alignment with Explicit Semantic-Geometric Modeling for 3D Visual Grounding1 Jan 2024 0 repositories listed
-
LQMFormer: Language-aware Query Mask Transformer for Referring Image Segmentation1 Jan 2024 0 repositories listed
-
Omni-Q: Omni-Directional Scene Understanding for Unsupervised Visual Grounding1 Jan 2024 0 repositories listed
-
Towards CLIP-driven Language-free 3D Visual Grounding via 2D-3D Relational Enhancement and Consistency1 Jan 2024 0 repositories listed
-
Viewpoint-Aware Visual Grounding in 3D Scenes1 Jan 2024 0 repositories listed
-
When Visual Grounding Meets Gigapixel-level Large-scale Scenes: Benchmark and Approach1 Jan 2024 0 repositories listed
-
Bridging Modality Gap for Visual Grounding with Effecitve Cross-modal Distillation29 Dec 2023 0 repositories listed
-
Cycle-Consistency Learning for Captioning and Grounding23 Dec 2023 0 repositories listed
-
Weakly-Supervised 3D Visual Grounding based on Visual Linguistic Alignment15 Dec 2023 0 repositories listed
-
Visual Grounding of Whole Radiology Reports for 3D CT Images8 Dec 2023 0 repositories listed
-
Improved Visual Grounding through Self-Consistent Explanations7 Dec 2023 0 repositories listed
-
Uni3DL: Unified Model for 3D and Language Understanding5 Dec 2023 0 repositories listed
-
Expand BERT Representation with Visual Information via Grounded Language Learning with Multimodal Partial Alignment4 Dec 2023 0 repositories listed
-
Context-Aware Indoor Point Cloud Object Generation through User Instructions26 Nov 2023 0 repositories listed
-
Enhancing Visual Grounding and Generalization: A Multi-Task Cycle Training Approach for Vision-Language Models21 Nov 2023 0 repositories listed
-
A Systematic Evaluation of GPT-4V's Multimodal Capability for Medical Image Analysis31 Oct 2023 0 repositories listed
-
Context Does Matter: End-to-end Panoptic Narrative Grounding with Deformable Attention Refined Matching Network25 Oct 2023 0 repositories listed
-
Lightweight In-Context Tuning for Multimodal Unified Models8 Oct 2023 0 repositories listed
-
Object2Scene: Putting Objects in Context for Open-Vocabulary 3D Detection18 Sep 2023 0 repositories listed
-
Four Ways to Improve Verbo-visual Fusion for Dense 3D Visual Grounding8 Sep 2023 0 repositories listed
-
Interpretable Visual Question Answering via Reasoning Supervision7 Sep 2023 0 repositories listed
-
FACET: Fairness in Computer Vision Evaluation Benchmark31 Aug 2023 0 repositories listed
-
WALL-E: Embodied Robotic WAiter Load Lifting with Large Language Model30 Aug 2023 0 repositories listed
-
3DRP-Net: 3D Relative Position-aware Network for 3D Visual Grounding25 Jul 2023 0 repositories listed
-
OG: Equip vision occupancy with instance segmentation and visual grounding12 Jul 2023 0 repositories listed
-
Learning with Difference Attention for Visually Grounded Self-supervised Representations26 Jun 2023 0 repositories listed
-
Extending CLIP's Image-Text Alignment to Referring Image Segmentation14 Jun 2023 0 repositories listed
-
Referring to Screen Texts with Voice Assistants10 Jun 2023 0 repositories listed
-
Benchmarking Diverse-Modal Entity Linking with Generative Models27 May 2023 0 repositories listed
-
Language-Guided 3D Object Detection in Point Cloud for Autonomous Driving25 May 2023 0 repositories listed
-
TreePrompt: Learning to Compose Tree Prompts for Explainable Visual Grounding19 May 2023 0 repositories listed
-
Vision-Language Pre-training with Object Contrastive Learning for 3D Scene Understanding18 May 2023 0 repositories listed
-
Sample-Specific Debiasing for Better Image-Text Models25 Apr 2023 0 repositories listed
-
Movie Box Office Prediction With Self-Supervised and Visually Grounded Pretraining20 Apr 2023 0 repositories listed
-
Medical Phrase Grounding with Region-Phrase Context Contrastive Alignment14 Mar 2023 0 repositories listed
-
Parallel Vertex Diffusion for Unified Visual Grounding13 Mar 2023 0 repositories listed
-
Focusing On Targets For Improving Weakly Supervised Visual Grounding22 Feb 2023 0 repositories listed
-
CoSign: Exploring Co-occurrence Signals in Skeleton-based Continuous Sign Language Recognition1 Jan 2023 0 repositories listed
-
Dynamic Inference With Grounding Based Vision and Language Models1 Jan 2023 0 repositories listed
-
GAFNet: A Global Fourier Self Attention Based Novel Network for multi-modal downstream tasks1 Jan 2023 0 repositories listed
-
ViewRefer: Grasp the Multi-view Knowledge for 3D Visual Grounding1 Jan 2023 0 repositories listed
-
Using Multiple Instance Learning to Build Multimodal Representations11 Dec 2022 0 repositories listed
-
UniT3D: A Unified Transformer for 3D Dense Captioning and Visual Grounding1 Dec 2022 0 repositories listed
-
MNER-QG: An End-to-End MRC framework for Multimodal Named Entity Recognition with Query Grounding27 Nov 2022 0 repositories listed
-
A survey on knowledge-enhanced multimodal learning19 Nov 2022 0 repositories listed
-
Are Current Decoding Strategies Capable of Facing the Challenges of Visual Dialogue?24 Oct 2022 0 repositories listed
-
A Visual Tour Of Current Challenges In Multimodal Language Models22 Oct 2022 0 repositories listed
-
Like a bilingual baby: The advantage of visually grounding a bilingual language model11 Oct 2022 0 repositories listed
-
YFACC: A Yorùbá speech-image dataset for cross-lingual keyword localisation through visual grounding10 Oct 2022 0 repositories listed
-
MAMO: Masked Multimodal Modeling for Fine-Grained Vision-Language Representation Learning9 Oct 2022 0 repositories listed
-
Differentiable Parsing and Visual Grounding of Natural Language Instructions for Object Placement1 Oct 2022 0 repositories listed
-
Dynamic MDETR: A Dynamic Multimodal Transformer Decoder for Visual Grounding28 Sep 2022 0 repositories listed
-
Visual Grounding of Inter-lingual Word-Embeddings8 Sep 2022 0 repositories listed
-
VLMAE: Vision-Language Masked Autoencoder19 Aug 2022 0 repositories listed
-
Toward Explainable and Fine-Grained 3D Grounding through Referring Textual Phrases5 Jul 2022 0 repositories listed
-
How direct is the link between words and images?30 Jun 2022 0 repositories listed
-
Tell Me the Evidence? Dual Visual-Linguistic Interaction for Answer Grounding21 Jun 2022 0 repositories listed
-
Bear the Query in Mind: Visual Grounding with Query-conditioned Convolution18 Jun 2022 0 repositories listed
-
Guiding Visual Question Answering with Attention Priors25 May 2022 0 repositories listed
-
Sim-To-Real Transfer of Visual Grounding for Human-Aided Ambiguity Resolution24 May 2022 0 repositories listed
-
Weakly-supervised segmentation of referring expressions10 May 2022 0 repositories listed
-
FindIt: Generalized Localization with Natural Language Queries31 Mar 2022 0 repositories listed
Syntology lines on 1 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced; each line links to that paper's sample list. Syntology's record for this page has not changed since , the first build that kept a record date for it; when this build read Syntology's graph is in the build record. For agents: get_harvested_code_for_paper(arxiv_id) lists each paper's samples; how to connect.