Browse State-of-the-Art › Scene Understanding › Papers, page 9
Scene Understanding
Papers archive 2025-07-28
archive papers tagged: 1,723 · with a code link: 720 · where Syntology ran a sample: 208 (182 with a run with no instrument failure, 26 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (208 of 1,723 tagged: 182 with a run with no instrument failure, 26 where every run was a failure of Syntology's instrument)
Page 9 of 18: papers 801 to 900 of 1,723, in archive order: by repositories listed in the archive (most first), then newest first, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers that list no repository come after every paper that lists one.
Papers without a page here are shown as plain text. A Syntology line reads “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the figure “where Syntology's instrument failed” counts failures of Syntology's instrument, not of the code. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code); hover it for the repositories the samples that ran came from. Abstracts are on each paper's page.
-
Single-Input Multi-Output Model Merging: Leveraging Foundation Models for Dense Multi-Task Learning15 Apr 2025 0 repositories listed
-
Foundation Models for Remote Sensing: An Analysis of MLLMs for Object Localization14 Apr 2025 0 repositories listed
-
DSM: Building A Diverse Semantic Map for 3D Visual Grounding11 Apr 2025 0 repositories listed
-
FindAnything: Open-Vocabulary and Object-Centric Mapping for Robot Exploration in Any Environment11 Apr 2025 0 repositories listed
-
FMLGS: Fast Multilevel Language Embedded Gaussians for Part-level Interactive Agents11 Apr 2025 0 repositories listed
-
DGOcc: Depth-aware Global Query-based Network for Monocular 3D Occupancy Prediction10 Apr 2025 0 repositories listed
-
Attributes-aware Visual Emotion Representation Learning9 Apr 2025 0 repositories listed
-
Audio-visual Event Localization on Portrait Mode Short Videos9 Apr 2025 0 repositories listed
-
MovSAM: A Single-image Moving Object Segmentation Framework Based on Deep Thinking9 Apr 2025 0 repositories listed
-
RayFronts: Open-Set Semantic Ray Frontiers for Online Scene Understanding and Exploration9 Apr 2025 0 repositories listed
-
PRIMEDrive-CoT: A Precognitive Chain-of-Thought Framework for Uncertainty-Aware Object Interaction in Driving Scene Scenario8 Apr 2025 0 repositories listed
-
RS-RAG: Bridging Remote Sensing Imagery and Comprehensive Knowledge with a Multi-Modal Dataset and Retrieval-Augmented Generation Model7 Apr 2025 0 repositories listed
-
CoMatcher: Multi-View Collaborative Feature Matching2 Apr 2025 0 repositories listed
-
Overlap-Aware Feature Learning for Robust Unsupervised Domain Adaptation for 3D Semantic Segmentation2 Apr 2025 0 repositories listed
-
Ross3D: Reconstructive Visual Instruction Tuning with 3D-Awareness2 Apr 2025 0 repositories listed
-
TransforMerger: Transformer-based Voice-Gesture Fusion for Robust Human-Robot Communication2 Apr 2025 0 repositories listed
-
Context-Aware Human Behavior Prediction Using Multimodal Large Language Models: Challenges and Insights1 Apr 2025 0 repositories listed
-
Zero-Shot 4D Lidar Panoptic Segmentation1 Apr 2025 0 repositories listed
-
PhysPose: Refining 6D Object Poses with Physical Constraints30 Mar 2025 0 repositories listed
-
Can DeepSeek Reason Like a Surgeon? An Empirical Evaluation for Vision-Language Understanding in Robotic-Assisted Surgery29 Mar 2025 0 repositories listed
-
Empowering Large Language Models with 3D Situation Awareness29 Mar 2025 0 repositories listed
-
Open-Vocabulary Semantic Segmentation with Uncertainty Alignment for Robotic Scene Understanding in Indoor Building Environments29 Mar 2025 0 repositories listed
-
A Dataset for Semantic Segmentation in the Presence of Unknowns28 Mar 2025 0 repositories listed
-
Endo-TTAP: Robust Endoscopic Tissue Tracking via Multi-Facet Guided Attention and Hybrid Flow-point Supervision28 Mar 2025 0 repositories listed
-
Evaluating Multimodal Language Models as Visual Assistants for Visually Impaired Users28 Mar 2025 0 repositories listed
-
Next-Best-Trajectory Planning of Robot Manipulators for Effective Observation and Exploration28 Mar 2025 0 repositories listed
-
NuGrounding: A Multi-View 3D Visual Grounding Framework in Autonomous Driving28 Mar 2025 0 repositories listed
-
Visual Jenga: Discovering Object Dependencies via Counterfactual Inpainting27 Mar 2025 0 repositories listed
-
DINeMo: Learning Neural Mesh Models with no 3D Annotations26 Mar 2025 0 repositories listed
-
OpenLex3D: A New Evaluation Benchmark for Open-Vocabulary 3D Scene Representations25 Mar 2025 0 repositories listed
-
Predicting the Road Ahead: A Knowledge Graph based Foundation Model for Scene Understanding in Autonomous Driving24 Mar 2025 0 repositories listed
-
Geometric Constrained Non-Line-of-Sight Imaging23 Mar 2025 0 repositories listed
-
MLLM-For3D: Adapting Multimodal Large Language Model for 3D Reasoning Segmentation23 Mar 2025 0 repositories listed
-
PanoGS: Gaussian-based Panoptic Segmentation for 3D Open Vocabulary Scene Understanding23 Mar 2025 0 repositories listed
-
PanopticSplatting: End-to-End Panoptic Gaussian Splatting23 Mar 2025 0 repositories listed
-
ClaraVid: A Holistic Scene Reconstruction Benchmark From Aerial Perspective With Delentropy-Based Complexity Profiling22 Mar 2025 0 repositories listed
-
ExCap3D: Expressive 3D Scene Understanding via Object Captioning with Varying Detail21 Mar 2025 0 repositories listed
-
From Monocular Vision to Autonomous Action: Guiding Tumor Resection via 3D Reconstruction20 Mar 2025 0 repositories listed
-
SemanticFlow: A Self-Supervised Framework for Joint Scene Flow Prediction and Instance Segmentation in Dynamic Environments19 Mar 2025 0 repositories listed
-
ChatBEV: A Visual Language Model that Understands BEV Maps18 Mar 2025 0 repositories listed
-
18 Mar 2025 0 repositories listed Syntology 9 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 2 violated, 5 with no contract checked; 2 where Syntology's instrument failed) · 1 unverified (of 10 harvested samples) · 4 pointer-only (licence)
-
These Magic Moments: Differentiable Uncertainty Quantification of Radiance Field Models18 Mar 2025 0 repositories listed
-
HIS-GPT: Towards 3D Human-In-Scene Multimodal Understanding17 Mar 2025 0 repositories listed
-
Learning-based 3D Reconstruction in Autonomous Driving: A Comprehensive Survey17 Mar 2025 0 repositories listed
-
EgoSplat: Open-Vocabulary Egocentric Scene Understanding with Language Embedded 3D Gaussian Splatting14 Mar 2025 0 repositories listed
-
Road Rage Reasoning with Vision-language Models (VLMs): Task Definition and Evaluation Dataset14 Mar 2025 0 repositories listed
-
Graph-Grounded LLMs: Leveraging Graphical Function Calling to Minimize LLM Hallucinations13 Mar 2025 0 repositories listed
-
TARS: Traffic-Aware Radar Scene Flow Estimation13 Mar 2025 0 repositories listed
-
TGP: Two-modal occupancy prediction with 3D Gaussian and sparse points for 3D Environment Awareness13 Mar 2025 0 repositories listed
-
Object-Aware DINO (Oh-A-Dino): Enhancing Self-Supervised Representations for Multi-Object Instance Retrieval12 Mar 2025 0 repositories listed
-
DIV-FF: Dynamic Image-Video Feature Fields For Environment Understanding in Egocentric Videos11 Mar 2025 0 repositories listed
-
General-Purpose Aerial Intelligent Agents Empowered by Large Language Models11 Mar 2025 0 repositories listed
-
Generating Robot Constitutions & Benchmarks for Semantic Safety11 Mar 2025 0 repositories listed
-
MaskAttn-UNet: A Mask Attention-Driven Framework for Universal Low-Resolution Image Segmentation11 Mar 2025 0 repositories listed
-
CoT-Drive: Efficient Motion Forecasting for Autonomous Driving with LLMs and Chain-of-Thought Prompting10 Mar 2025 0 repositories listed
-
LLaFEA: Frame-Event Complementary Fusion for Fine-Grained Spatiotemporal Understanding in LMMs10 Mar 2025 0 repositories listed
-
Feature-EndoGaussian: Feature Distilled Gaussian Splatting in Surgical Deformable Scene Reconstruction8 Mar 2025 0 repositories listed
-
Segment Anything, Even Occluded8 Mar 2025 0 repositories listed
-
SplatTalk: 3D VQA with Gaussian Splatting8 Mar 2025 0 repositories listed
-
EvidMTL: Evidential Multi-Task Learning for Uncertainty-Aware Semantic Surface Mapping from Monocular RGB Images6 Mar 2025 0 repositories listed
-
Improving 6D Object Pose Estimation of metallic Household and Industry Objects5 Mar 2025 0 repositories listed
-
SurgiSAM2: Fine-tuning a foundational model for surgical video anatomy segmentation and detection5 Mar 2025 0 repositories listed
-
Vision-Language Models Struggle to Align Entities across Modalities5 Mar 2025 0 repositories listed
-
Label-Efficient LiDAR Panoptic Segmentation4 Mar 2025 0 repositories listed
-
SSNet: Saliency Prior and State Space Model-based Network for Salient Object Detection in RGB-D Images4 Mar 2025 0 repositories listed
-
Every SAM Drop Counts: Embracing Semantic Priors for Multi-Modality Image Fusion and Beyond3 Mar 2025 0 repositories listed
-
OpenGS-SLAM: Open-Set Dense Semantic SLAM with 3D Gaussian Splatting for Object-Level Scene Understanding3 Mar 2025 0 repositories listed
-
vS-Graphs: Integrating Visual SLAM and Situational Graphs through Multi-level Scene Understanding3 Mar 2025 0 repositories listed
-
Floorplan-SLAM: A Real-Time, High-Accuracy, and Long-Term Multi-Session Point-Plane SLAM for Efficient Floorplan Reconstruction1 Mar 2025 0 repositories listed
-
VLM-E2E: Enhancing End-to-End Autonomous Driving with Multimodal Driver Attention Fusion25 Feb 2025 0 repositories listed
-
AAD-LLM: Neural Attention-Driven Auditory Scene Understanding24 Feb 2025 0 repositories listed
-
Dr. Splat: Directly Referring 3D Gaussian Splatting via Direct Language Embedding Registration23 Feb 2025 0 repositories listed
-
AVD2: Accident Video Diffusion for Accident Video Description20 Feb 2025 0 repositories listed
-
Sce2DriveX: A Generalized MLLM Framework for Scene-to-Drive Learning19 Feb 2025 0 repositories listed
-
Understanding and Evaluating Hallucinations in 3D Visual Language Models18 Feb 2025 0 repositories listed
-
Surgical Scene Understanding in the Era of Foundation AI Models: A Comprehensive Review16 Feb 2025 0 repositories listed
-
3D-Grounded Vision-Language Framework for Robotic Task Planning: Automated Prompt Synthesis and Supervised Reasoning13 Feb 2025 0 repositories listed
-
FLARES: Fast and Accurate LiDAR Multi-Range Semantic Segmentation13 Feb 2025 0 repositories listed
-
sshELF: Single-Shot Hierarchical Extrapolation of Latent Features for 3D Reconstruction from Sparse-Views6 Feb 2025 0 repositories listed
-
Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation4 Feb 2025 0 repositories listed
-
AquaticCLIP: A Vision-Language Foundation Model for Underwater Scene Analysis3 Feb 2025 0 repositories listed
-
Integrating LMM Planners and 3D Skill Policies for Generalizable Manipulation30 Jan 2025 0 repositories listed
-
Efficient Interactive 3D Multi-Object Removal29 Jan 2025 0 repositories listed
-
Contextual Self-paced Learning for Weakly Supervised Spatio-Temporal Video Grounding28 Jan 2025 0 repositories listed
-
PhysBench: Benchmarking and Enhancing Vision-Language Models for Physical World Understanding27 Jan 2025 0 repositories listed
-
Unveiling the Potential of iMarkers: Invisible Fiducial Markers for Advanced Robotics26 Jan 2025 0 repositories listed
-
Scene Understanding Enabled Semantic Communication with Open Channel Coding24 Jan 2025 0 repositories listed
-
GeomGS: LiDAR-Guided Geometry-Aware Gaussian Splatting for Robot Localization23 Jan 2025 0 repositories listed
-
Neural Radiance Fields for the Real World: A Survey22 Jan 2025 0 repositories listed
-
Separated Inter/Intra-Modal Fusion Prompts for Compositional Zero-Shot Learning22 Jan 2025 0 repositories listed
-
20 Jan 2025 0 repositories listed
-
A Vision-Language Framework for Multispectral Scene Representation Using Language-Grounded Features17 Jan 2025 0 repositories listed
-
YETI (YET to Intervene) Proactive Interventions by Multimodal AI Agents in Augmented Reality Tasks16 Jan 2025 0 repositories listed
-
Embodied Scene Understanding for Vision Language Models via MetaVQA15 Jan 2025 0 repositories listed
-
Zero-Shot Scene Understanding for Automatic Target Recognition Using Large Vision-Language Models13 Jan 2025 0 repositories listed
-
Application of Vision-Language Model to Pedestrians Behavior and Scene Understanding in Autonomous Driving12 Jan 2025 0 repositories listed
-
UniQ: Unified Decoder with Task-specific Queries for Efficient Scene Graph Generation10 Jan 2025 0 repositories listed
-
A Systematic Literature Review on Deep Learning-based Depth Estimation in Computer Vision9 Jan 2025 0 repositories listed
-
Vision-Language Models for Autonomous Driving: CLIP-Based Dynamic Scene Understanding9 Jan 2025 0 repositories listed
-
NextStop: An Improved Tracker For Panoptic LIDAR Segmentation Data8 Jan 2025 0 repositories listed
Syntology lines on 1 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced; each line links to that paper's sample list. Syntology's record for this page has not changed since , the first build that kept a record date for it; when this build read Syntology's graph is in the build record. For agents: get_harvested_code_for_paper(arxiv_id) lists each paper's samples; how to connect.