Browse State-of-the-Art › Multimodal Large Language Model
Multimodal Large Language Model
160 papers with code · 0 benchmarks · 1 dataset archive 2025-07-28
Benchmarks archive 2025-07-28
No benchmark for this task in the archive.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
1 dataset whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
No subtask under this task in the archive's task tree.
Most implemented papers archive 2025-07-28
30 shown of 160 papers with code (347 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
23 Jun 2023 4 repositories listedMultimodal Large Language Model (MLLM) relies on the powerful LLM to perform multimodal tasks, showing amazing emergent abilities in recent studies, such as writing poems based on an image.
-
27 Feb 2024 3 repositories listed Syntology ran 9 of 17 samples · 8 unverifiedThis paper presents ShapeLLM, the first 3D Multimodal Large Language Model (LLM) designed for embodied interaction, exploring a universal 3D object understanding with 3D point clouds and languages.
-
21 Jan 2025 2 repositories listedWe present VARGPT, a novel multimodal large language model (MLLM) that unifies visual understanding and generation within a single autoregressive framework.
-
3 Dec 2024 2 repositories listedThis survey fills a critical gap in the literature by providing an integrated overview of RSTVLM, offering a foundation for further advancements in remote sensing temporal image understanding.
-
17 Nov 2024 2 repositories listedTo balance the trade-off between generalization and specialization, we propose measuring the parameter importance for both pre-trained and fine-tuning distributions, based on frozen pre-trained weight magnitude and…
-
11 Oct 2024 2 repositories listedThe salient multimodal capabilities and interactive experience of GPT-4o highlight its critical role in practical applications, yet it lacks a high-performing open-source counterpart.
-
31 May 2024 2 repositories listedHowever, the misalignment between two embedding strategies in MLLMs -- the structural textual embeddings based on an embedding look-up table and the continuous embeddings generated directly by the vision encoder --…
-
4 Apr 2024 2 repositories listed Syntology ran 2 of 2 samples · 0 unverifiedThis paper introduces MiniGPT4-Video, a multimodal Large Language Model (LLM) designed specifically for video understanding.
-
4 Feb 2024 2 repositories listed Syntology ran 4 of 4 samples · 0 unverified · 3 pointer-only (licence)This paper focuses on jailbreaking attacks against multi-modal large language models (MLLMs), seeking to elicit MLLMs to generate objectionable responses to harmful user queries.
-
19 Jan 2024 2 repositories listed Syntology ran 8 of 11 samples · 3 unverifiedRecently, the astonishing performance of large language models (LLMs) in natural language comprehension and generation tasks triggered lots of exploration of using them as central controllers to build agent systems.
-
28 Dec 2023 2 repositories listed Syntology ran 7 of 7 samples · 0 unverifiedIn recent years, multimodal large language models (MLLMs) such as GPT-4V have demonstrated remarkable advancements, excelling in a variety of vision-language tasks.
-
4 Dec 2023 2 repositories listed Syntology ran 7 of 11 samples · 4 unverifiedThis work proposes TimeChat, a time-sensitive multimodal large language model specifically designed for long video understanding.
-
11 Oct 2023 2 repositories listed Syntology ran 7 of 8 samples · 1 unverified · 8 pointer-only (licence)We introduce Ferret, a new Multimodal Large Language Model (MLLM) capable of understanding spatial referring of any shape or granularity within an image and accurately grounding open-vocabulary descriptions.
-
26 Jun 2023 2 repositories listedWe introduce Kosmos-2, a Multimodal Large Language Model (MLLM), enabling new capabilities of perceiving object descriptions (e.
-
15 Jul 2025 1 repository listedSmoke is the first visible indicator of a wildfire.With the advancement of deep learning, image-based smoke detection has become a crucial method for detecting and preventing forest fires.
-
26 Jun 2025 1 repository listedWhile end-to-end video-to-audio generation has greatly improved, producing high-fidelity audio that authentically captures the nuances of visual content remains challenging.
-
26 Jun 2025 1 repository listedAs one of the earliest ancient languages, Oracle Bone Script (OBS) encapsulates the cultural records and intellectual expressions of ancient civilizations.
-
23 Jun 2025 1 repository listedAccurate and interpretable multi-disease diagnosis remains a critical challenge in medical research, particularly when leveraging heterogeneous multimodal medical data.
-
22 Jun 2025 1 repository listed Syntology ran 3 of 3 samples · 0 unverified · 3 pointer-only (licence)Recent advances in multimodal generative models have unlocked photorealistic, instruction-aligned image generation, yet leading systems like GPT-4o-Image remain proprietary and inaccessible.
-
19 Jun 2025 1 repository listedThis paper explores the relationship between the condition number of a neural network's weight tensor and the extent of information encoded by the associated processing unit, viewed through the lens of information…
-
16 Jun 2025 1 repository listed Syntology ran 0 of 6 samples · 6 unverified · 6 pointer-only (licence)Data visualization generation using Large Language Models (LLMs) has shown promising results but often produces suboptimal visualizations that require human intervention for improvement.
-
30 May 2025 1 repository listedPeriodic or quasi-periodic phenomena reveal intrinsic characteristics in various natural processes, such as weather patterns, movement behaviors, traffic flows, and biological signals.
-
30 May 2025 1 repository listedCompared to discriminative models like CLIP, generative models are better at capturing image details because they are trained to learn the data distribution of images.
-
28 May 2025 1 repository listedText-to-image generation increasingly demands access to domain-specific, fine-grained, and rapidly evolving knowledge that pretrained models cannot fully capture.
-
27 May 2025 1 repository listedTo address data scarcity, we introduce SuperRS-VQA (avg.
-
26 May 2025 1 repository listedDiffusion models have become the mainstream architecture for text-to-image generation, achieving remarkable progress in visual quality and prompt controllability.
-
26 May 2025 1 repository listed Syntology ran 3 of 3 samples · 0 unverified · 3 pointer-only (licence)Multimodal Large Language Models (MLLMs) have demonstrated impressive capabilities across tasks, yet they often exhibit difficulty in distinguishing task-relevant from irrelevant signals, particularly in tasks like…
-
26 May 2025 1 repository listed Syntology ran 5 of 5 samples · 0 unverified · 5 pointer-only (licence)While foundation models update slowly due to resource-intensive training requirements, domain-specific models evolve between updates.
-
22 May 2025 1 repository listedTo fill this gap, in this paper, we propose ChemMLLM, a unified chemical multimodal large language model for molecule understanding and generation.
-
22 May 2025 1 repository listed Syntology ran 1 of 1 samples · 0 unverified · 1 pointer-only (licence)We observe that training with a purely discrete diffusion approach leads to significant training instability, suboptimal performance, and severe length bias issues.
Syntology lines on 12 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections