Browse State-of-the-Art › Multimodal Large Language Model › Papers, page 3
Multimodal Large Language Model
Papers archive 2025-07-28
archive papers tagged: 347 · with a code link: 160 · where Syntology ran a sample: 59 (51 with a run with no instrument failure, 8 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (59 of 347 tagged: 51 with a run with no instrument failure, 8 where every run was a failure of Syntology's instrument)
Page 3 of 4: papers 201 to 300 of 347, in archive order: by repositories listed in the archive (most first), then newest first, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers that list no repository come after every paper that lists one.
Papers without a page here are shown as plain text. A Syntology line reads “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the figure “where Syntology's instrument failed” counts failures of Syntology's instrument, not of the code. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code); hover it for the repositories the samples that ran came from. Abstracts are on each paper's page.
-
CleanMAP: Distilling Multimodal LLMs for Confidence-Driven Crowdsourced HD Map Updates14 Apr 2025 0 repositories listed
-
Mavors: Multi-granularity Video Representation for Multimodal Large Language Model14 Apr 2025 0 repositories listed
-
Marmot: Multi-Agent Reasoning for Multi-Object Self-Correcting in Improving Image-Text Alignment10 Apr 2025 0 repositories listed
-
Face-LLaVA: Facial Expression and Attribute Understanding through Instruction Tuning9 Apr 2025 0 repositories listed
-
MovSAM: A Single-image Moving Object Segmentation Framework Based on Deep Thinking9 Apr 2025 0 repositories listed
-
Q-Agent: Quality-Driven Chain-of-Thought Image Restoration Agent through Robust Multimodal Large Language Model9 Apr 2025 0 repositories listed
-
Towards Visual Text Grounding of Multimodal Large Language Model7 Apr 2025 0 repositories listed
-
Universal Item Tokenization for Transferable Generative Recommendation6 Apr 2025 0 repositories listed
-
Towards Unified Referring Expression Segmentation Across Omni-Level Visual Target Granularities2 Apr 2025 0 repositories listed
-
Open-Qwen2VL: Compute-Efficient Pre-Training of Fully-Open Multimodal LLMs on Academic Resources1 Apr 2025 0 repositories listed
-
Orchestrate Multimodal Data with Batch Post-Balancing to Accelerate Multimodal Large Language Model Training31 Mar 2025 0 repositories listed
-
Dynamic Pyramid Network for Efficient Multimodal Large Language Model26 Mar 2025 0 repositories listed
-
MLLM-For3D: Adapting Multimodal Large Language Model for 3D Reasoning Segmentation23 Mar 2025 0 repositories listed
-
LEGION: Learning to Ground and Explain for Synthetic Image Detection19 Mar 2025 0 repositories listed
-
UPME: An Unsupervised Peer Review Framework for Multimodal Large Language Model Evaluation19 Mar 2025 0 repositories listed
-
SpaceVLLM: Endowing Multimodal Large Language Model with Spatio-Temporal Video Grounding Capability18 Mar 2025 0 repositories listed
-
HiDe-LLaVA: Hierarchical Decoupling for Continual Instruction Tuning of Multimodal Large Language Model17 Mar 2025 0 repositories listed
-
GeoRSMLLM: A Multimodal Large Language Model for Vision-Language Tasks in Geoscience and Remote Sensing16 Mar 2025 0 repositories listed
-
When neural implant meets multimodal LLM: A dual-loop system for neuromodulation and naturalistic neuralbehavioral research16 Mar 2025 0 repositories listed
-
OmniDiff: A Comprehensive Benchmark for Fine-grained Image Difference Captioning14 Mar 2025 0 repositories listed
-
CINEMA: Coherent Multi-Subject Video Generation via MLLM-Based Guidance13 Mar 2025 0 repositories listed
-
Hybrid Agents for Image Restoration13 Mar 2025 0 repositories listed
-
Lightweight Multimodal Artificial Intelligence Framework for Maritime Multi-Scene Recognition10 Mar 2025 0 repositories listed
-
CL-MoE: Enhancing Multimodal Large Language Model with Dual Momentum Mixture-of-Experts for Continual Visual Question Answering1 Mar 2025 0 repositories listed
-
Optimus-2: Multimodal Minecraft Agent with Goal-Observation-Action Conditioned Policy27 Feb 2025 0 repositories listed
-
Gesture-Aware Zero-Shot Speech Recognition for Patients with Language Disorders18 Feb 2025 0 repositories listed
-
MMRC: A Large-Scale Benchmark for Understanding Multimodal Large Language Model in Real-World Conversation17 Feb 2025 0 repositories listed
-
Leveraging Multimodal-LLMs Assisted by Instance Segmentation for Intelligent Traffic Monitoring16 Feb 2025 0 repositories listed
-
Distraction is All You Need for Multimodal Large Language Model Jailbreaking15 Feb 2025 0 repositories listed
-
On Fairness of Unified Multimodal Large Language Model for Image Generation5 Feb 2025 0 repositories listed
-
MPIC: Position-Independent Multimodal Context Caching System for Efficient MLLM Serving4 Feb 2025 0 repositories listed
-
Learning Free Token Reduction for Multi-Modal Large Language Models29 Jan 2025 0 repositories listed
-
HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding25 Jan 2025 0 repositories listed
-
EventVL: Understand Event Streams via Multimodal Large Language Model23 Jan 2025 0 repositories listed
-
Interpretable Droplet Digital PCR Assay for Trustworthy Molecular Diagnostics16 Jan 2025 0 repositories listed
-
Omni-RGPT: Unifying Image and Video Region-level Understanding via Token Marks14 Jan 2025 0 repositories listed
-
MinMo: A Multimodal Large Language Model for Seamless Voice Interaction10 Jan 2025 0 repositories listed
-
LLaVA-Octopus: Unlocking Instruction-Driven Adaptive Projector Fusion for Video Understanding9 Jan 2025 0 repositories listed
-
Interpretable Face Anti-Spoofing: Enhancing Generalization with Multimodal Large Language Models3 Jan 2025 0 repositories listed
-
Beyond Text: Implementing Multimodal Large Language Model-Powered Multi-Agent Systems Using a No-Code Platform1 Jan 2025 0 repositories listed
-
GroundingFace: Fine-grained Face Understanding via Pixel Grounding Multimodal Large Language Model1 Jan 2025 0 repositories listed
-
S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Model with Spatio-Temporal Visual Representation1 Jan 2025 0 repositories listed
-
ST³: Accelerating Multimodal Large Language Model by Spatial-Temporal Visual Token Trimming28 Dec 2024 0 repositories listed
-
A Large-scale Interpretable Multi-modality Benchmark for Facial Image Forgery Localization27 Dec 2024 0 repositories listed
-
SubstationAI: Multimodal Large Model-Based Approaches for Analyzing Substation Equipment Faults22 Dec 2024 0 repositories listed
-
J-EDI QA: Benchmark for deep-sea organism-specific multimodal LLM20 Dec 2024 0 repositories listed
-
Make Imagination Clearer! Stable Diffusion-based Visual Imagination for Multimodal Machine Translation17 Dec 2024 0 repositories listed
-
A Survey of Mathematical Reasoning in the Era of Multimodal Large Language Model: Benchmark, Method & Challenges16 Dec 2024 0 repositories listed
-
MERaLiON-SpeechEncoder: Towards a Speech Foundation Model for Singapore and Beyond16 Dec 2024 0 repositories listed
-
EasyRef: Omni-Generalized Group Image Reference for Diffusion Models via Multimodal LLM12 Dec 2024 0 repositories listed
-
COEF-VQ: Cost-Efficient Video Quality Understanding through a Cascaded Multimodal LLM Framework11 Dec 2024 0 repositories listed
-
DiffSensei: Bridging Multi-Modal LLMs and Diffusion Models for Customized Manga Generation10 Dec 2024 0 repositories listed
-
9 Dec 2024 0 repositories listed
-
EditScout: Locating Forged Regions from Diffusion-based Edited Images with Multimodal LLM5 Dec 2024 0 repositories listed
-
EgoPlan-Bench2: A Benchmark for Multimodal Large Language Model Planning in Real-World Scenarios5 Dec 2024 0 repositories listed
-
DynamicControl: Adaptive Condition Selection for Improved Text-to-Image Generation4 Dec 2024 0 repositories listed
-
ObjectFinder: An Open-Vocabulary Assistive System for Interactive Object Search by Blind People4 Dec 2024 0 repositories listed
-
WSI-LLaVA: A Multimodal Large Language Model for Whole Slide Image3 Dec 2024 0 repositories listed
-
MoTrans: Customized Motion Transfer with Text-driven Video Diffusion Models2 Dec 2024 0 repositories listed
-
SeqAfford: Sequential 3D Affordance Reasoning via Multimodal Large Language Model2 Dec 2024 0 repositories listed
-
Realistic Corner Case Generation for Autonomous Vehicles with Multimodal Large Language Model29 Nov 2024 0 repositories listed
-
Multimodal large language model for wheat breeding: a new exploration of smart breeding20 Nov 2024 0 repositories listed
-
CUE-M: Contextual Understanding and Enhanced Search with Multimodal Large Language Model19 Nov 2024 0 repositories listed
-
Med-2E3: A 2D-Enhanced 3D Medical Multimodal Large Language Model19 Nov 2024 0 repositories listed
-
StreetviewLLM: Extracting Geographic Information Using a Chain-of-Thought Multimodal Large Language Model19 Nov 2024 0 repositories listed
-
Mitigating Hallucination in Multimodal Large Language Model via Hallucination-targeted Direct Preference Optimization15 Nov 2024 0 repositories listed
-
11 Nov 2024 0 repositories listed
-
TourSynbio-Search: A Large Language Model Driven Agent Framework for Unified Search Method for Protein Engineering9 Nov 2024 0 repositories listed
-
ChatTracker: Enhancing Visual Tracking Performance via Chatting with Multimodal Large Language Model4 Nov 2024 0 repositories listed
-
Can Multimodal Large Language Model Think Analogically?2 Nov 2024 0 repositories listed
-
Web-Scale Visual Entity Recognition: An LLM-Driven Data Approach31 Oct 2024 0 repositories listed
-
Ferret-UI 2: Mastering Universal User Interface Understanding Across Platforms24 Oct 2024 0 repositories listed
-
Interpretable Bilingual Multimodal Large Language Model for Diverse Biomedical Tasks24 Oct 2024 0 repositories listed
-
Towards Real Zero-Shot Camouflaged Object Segmentation without Camouflaged Annotations22 Oct 2024 0 repositories listed
-
LLaVA-Ultra: Large Chinese Language and Vision Assistant for Ultrasound19 Oct 2024 0 repositories listed
-
MoChat: Joints-Grouped Spatio-Temporal Grounding LLM for Multi-Turn Motion Comprehension and Description15 Oct 2024 0 repositories listed
-
ForgeryGPT: Multimodal Large Language Model For Explainable Image Forgery Detection and Localization14 Oct 2024 0 repositories listed
-
ViT3D Alignment of LLaMA3: 3D Medical Image Report Generation11 Oct 2024 0 repositories listed
-
RespLLM: Unifying Audio and Text with Multimodal LLMs for Generalized Respiratory Health Prediction7 Oct 2024 0 repositories listed
-
OCC-MLLM:Empowering Multimodal Large Language Model For the Understanding of Occluded Objects2 Oct 2024 0 repositories listed
-
VMAD: Visual-enhanced Multimodal Large Language Model for Zero-Shot Anomaly Detection30 Sep 2024 0 repositories listed
-
CadVLM: Bridging Language and Vision in the Generation of Parametric CAD Sketches26 Sep 2024 0 repositories listed
-
EAGLE: Egocentric AGgregated Language-video Engine26 Sep 2024 0 repositories listed
-
CLSP: High-Fidelity Contrastive Language-State Pre-training for Agent State Representation24 Sep 2024 0 repositories listed
-
Decoding Style: Efficient Fine-Tuning of LLMs for Image-Guided Outfit Recommendation with Preference18 Sep 2024 0 repositories listed
-
MIP-GAF: A MLLM-annotated Benchmark for Most Important Person Localization and Group Context Understanding10 Sep 2024 0 repositories listed
-
Multimodal Large Language Model Driven Scenario Testing for Autonomous Vehicles10 Sep 2024 0 repositories listed
-
MLLM-LLaVA-FL: Multimodal Large Language Model Assisted Federated Learning9 Sep 2024 0 repositories listed
-
A Medical Multimodal Large Language Model for Pediatric Pneumonia4 Sep 2024 0 repositories listed
-
Balancing Performance and Efficiency: A Multimodal Large Language Model Pruning Method based Image Text Interaction2 Sep 2024 0 repositories listed
-
DPDEdit: Detail-Preserved Diffusion Models for Multimodal Fashion Image Editing2 Sep 2024 0 repositories listed
-
Multimodal Multi-turn Conversation Stance Detection: A Challenge Dataset and Effective Model1 Sep 2024 0 repositories listed
-
OrthoDoc: Multimodal Large Language Model for Assisting Diagnosis in Computed Tomography30 Aug 2024 0 repositories listed
-
22 Aug 2024 0 repositories listed
-
Vintern-1B: An Efficient Multimodal Large Language Model for Vietnamese22 Aug 2024 0 repositories listed
-
CaRDiff: Video Salient Object Ranking Chain of Thought Reasoning for Saliency Prediction with Diffusion21 Aug 2024 0 repositories listed
-
EE-MLLM: A Data-Efficient and Compute-Efficient Multimodal Large Language Model21 Aug 2024 0 repositories listed
-
Video Emotion Open-vocabulary Recognition Based on Multimodal Large Language Model21 Aug 2024 0 repositories listed
-
PanoSent: A Panoptic Sextuple Extraction Benchmark for Multimodal Conversational Aspect-based Sentiment Analysis18 Aug 2024 0 repositories listed
-
ChatGPT Meets Iris Biometrics9 Aug 2024 0 repositories listed