| OralAgent: Integrating Reasoning, Tools, and Knowledge for Interactive Dental Image Analysis added by Syntology |
2026-05 (from id) |
isjinghao/OralAgent/oralagent/llava/mm_utils.py c3ee9d07c900dd55 |
ran
|
Apache-2.0 (permissive) |
| CLIP Tricks You: Training-free Token Pruning for Efficient Pixel Grounding in Large Vision-Language Models added by Syntology |
2026-05 (from id) |
sejong-rcv/LiteLVLM/model/llava/mm_utils.py 3aada9ef083faccd |
ran · honoured contract
|
Apache-2.0 (permissive) |
| Uncertainty-Aware Knowledge Distillation for Multimodal Large Language Models added by Syntology |
2026-03 (from id) |
Jingchensun/beta-kd/mobilevlm/utils.py 3aada9ef083faccd |
ran · honoured contract
|
no licence file found · pointer only |
| Beyond Next-Token Alignment: Distilling Multimodal Large Language Models via Token Interactions added by Syntology |
2026-02 (from id) |
lchen1019/Align-TI/alignti/mm_utils.py c3ee9d07c900dd55 |
ran
|
Apache-2.0 (permissive) |
| Benchmarking Direct Preference Optimization for Medical Large Vision-Language Models added by Syntology |
2026-01 (from id) |
dmis-lab/med-vlm-dpo/inference/LLaVA-Med/llava/mm_utils.py c3ee9d07c900dd55 |
ran
|
no licence file found · pointer only |
| Uni-X: Mitigating Modality Conflict with a Two-End-Separated Architecture for Unified Multimodal Models added by Syntology |
2025-09 (from id) |
CURRENTF/Uni-X/uni_arch/mm_utils.py c3ee9d07c900dd55 |
ran
|
MIT (permissive) |
| arXiv:2507.18300 |
2025-07 (from id) |
360CVGroup/LMM-Det/llava/mm_utils.py c3ee9d07c900dd55 |
ran
|
Apache-2.0 (permissive) |
| arXiv:2507.04976 |
2025-07 (from id) |
EsYoon7/UVQA/llava/mm_utils.py c3ee9d07c900dd55 |
ran
|
no licence file found · pointer only |
| Hunyuan3D 2.5: Towards High-Fidelity 3D Assets Generation with Ultimate Details |
19 Jun 2025 |
tencent/hunyuan3d-2/api_server.py 0dde2e782959c0bd |
unverified |
licence not identified · pointer only |
| Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models |
27 May 2025 |
jefferyzhan/griffon/griffon/mm_utils.py c3ee9d07c900dd55 |
ran
|
Apache-2.0 (permissive) |
| SegEarth-R1: Geospatial Pixel Reasoning via Large Language Model |
13 Apr 2025 |
earth-insights/segearth-r1/segearth_r1/mm_utils.py c3ee9d07c900dd55 |
ran
|
Apache-2.0 (permissive) |
| EVEv2: Improved Baselines for Encoder-Free Vision-Language Models |
10 Feb 2025 |
baaivision/EVE/EVEv1/eve/mm_utils.py 0dde2e782959c0bd |
unverified |
MIT (permissive) |
| MedRAX: Medical Reasoning Agent for Chest X-ray |
4 Feb 2025 |
bowang-lab/medrax/medrax/llava/mm_utils.py c3ee9d07c900dd55 |
ran
|
Apache-2.0 (permissive) |
| VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding |
22 Jan 2025 |
damo-nlp-sg/videollama3/videollama3/mm_utils.py 0dde2e782959c0bd |
unverified |
Apache-2.0 (permissive) |
| InstructSeg: Unifying Instructed Visual Segmentation with Multi-modal Large Language Models |
18 Dec 2024 |
congvvc/instructseg/instructseg/model/mipha/mm_utils.py c3ee9d07c900dd55 |
ran
|
Apache-2.0 (permissive) |
| WiseAD: Knowledge Augmented End-to-End Autonomous Driving with Vision-Language Model |
13 Dec 2024 |
wyddmw/WiseAD/mobilevlm/utils.py 3aada9ef083faccd |
ran · honoured contract
|
no licence file found · pointer only |
| LinVT: Empower Your Image-level Large Language Model to Understand Videos |
6 Dec 2024 |
gls0425/linvt/streamlit_demo/model_worker.py 0dde2e782959c0bd |
unverified |
no licence file found · pointer only |
| FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression |
5 Dec 2024 |
codefanw/flashsloth/flashsloth/mm_utils.py c3ee9d07c900dd55 |
ran
|
no licence file found · pointer only |
| LongVALE: Vision-Audio-Language-Event Benchmark Towards Time-Aware Omni-Modal Perception of Long Videos |
29 Nov 2024 |
ttgeng233/LongVALE/longvalellm/mm_utils.py c3ee9d07c900dd55 |
ran
|
MIT (permissive) |
| HyperSeg: Towards Universal Visual Segmentation with Large Language Model |
26 Nov 2024 |
congvvc/HyperSeg/hyperseg/model/mipha/mm_utils.py c3ee9d07c900dd55 |
ran
|
Apache-2.0 (permissive) |
| Steering Away from Harm: An Adaptive Approach to Defending Vision Language Model Against Jailbreaks |
23 Nov 2024 |
ASTRAL-Group/ASTRA/utility_eval/minigpt_mmbench.py 3aada9ef083faccd |
ran · honoured contract
|
no licence file found · pointer only |
| V-DPO: Mitigating Hallucination in Large Vision Language Models via Vision-Guided Direct Preference Optimization |
5 Nov 2024 |
yuxixie/v-dpo/llava_dpo/mm_utils.py c3ee9d07c900dd55 |
ran
|
Apache-2.0 (permissive) |
| VidEgoThink: Assessing Egocentric Video Understanding Capabilities for Embodied AI |
15 Oct 2024 |
adacheng/egothink/models/lego/mm_utils.py c3ee9d07c900dd55 |
ran
|
Apache-2.0 (permissive) |
| Mitigating Modality Prior-Induced Hallucinations in Multimodal Large Language Models via Deciphering Attention Causality |
7 Oct 2024 |
The-Martyr/CausalMM/llava-1.5/experiments/llava/mm_utils.py c3ee9d07c900dd55 |
ran
|
MIT (permissive) |
| CDChat: A Large Multimodal Model for Remote Sensing Change Description |
24 Sep 2024 |
techmn/cdchat/cdchat/mm_utils.py c3ee9d07c900dd55 |
ran
|
no licence file found · pointer only |
| M$^2$PT: Multimodal Prompt Tuning for Zero-shot Instruction Learning |
24 Sep 2024 |
william-wang618/mmpt-emnlp2024/M2PT/mm_utils.py c3ee9d07c900dd55 |
ran
|
no licence file found · pointer only |
| Fit and Prune: Fast and Training-free Visual Token Pruning for Multi-modal Large Language Models |
16 Sep 2024 |
ywh187/fitprune/LLaVA_1.5/llava/mm_utils.py c3ee9d07c900dd55 |
ran
|
no licence file found · pointer only |
| ChangeChat: An Interactive Model for Remote Sensing Change Analysis via Multimodal Instruction Tuning |
13 Sep 2024 |
hanlinwu/changechat/changechat/mm_utils.py c3ee9d07c900dd55 |
ran
|
no licence file found · pointer only |
| VILA-U: a Unified Foundation Model Integrating Visual Understanding and Generation |
6 Sep 2024 |
mit-han-lab/vila-u/vila_u/mm_utils.py 0dde2e782959c0bd |
unverified |
MIT (permissive) |
| Recoverable Compression: A Multimodal Vision Token Recovery Mechanism Guided by Text Information |
2 Sep 2024 |
banjiuyufen/Recoverable-Compression/llava/mm_utils.py c3ee9d07c900dd55 |
ran
|
Apache-2.0 (permissive) |
| IAA: Inner-Adaptor Architecture Empowers Frozen Large Language Model with Multimodal Capabilities |
23 Aug 2024 |
360cvgroup/inner-adaptor-architecture/iaa/mm_utils.py c3ee9d07c900dd55 |
ran
|
Apache-2.0 (permissive) |
| ECG-Chat: A Large ECG-Language Model for Cardiac Disease Diagnosis |
16 Aug 2024 |
YubaoZhao/ECG-Chat/llava/llava/mm_utils.py c3ee9d07c900dd55 |
ran
|
no licence file found · pointer only |
| Paying More Attention to Image: A Training-Free Method for Alleviating Hallucination in LVLMs |
31 Jul 2024 |
hasanar1f/llava-hallunication-fix/modPAI/llava/mm_utils.py c3ee9d07c900dd55 |
ran
|
Apache-2.0 (permissive) |
| DenseFusion-1M: Merging Vision Experts for Comprehensive Multimodal Perception |
11 Jul 2024 |
baaivision/DenseFusion/densefusion/mm_utils.py 0dde2e782959c0bd |
unverified |
no licence file found · pointer only |
| Web2Code: A Large-scale Webpage-to-Code Dataset and Evaluation Framework for Multimodal LLMs |
28 Jun 2024 |
MBZUAI-LLM/web2code/web2code/llava/mm_utils.py c3ee9d07c900dd55 |
ran
|
no licence file found · pointer only |
| Biomedical Visual Instruction Tuning with Clinician Preference Alignment |
19 Jun 2024 |
mao1207/BioMed-VITAL/backbone/mm_utils.py c3ee9d07c900dd55 |
ran
|
no licence file found · pointer only |
| Holmes-VAD: Towards Unbiased and Explainable Video Anomaly Detection via Multi-modal LLM |
18 Jun 2024 |
pipixin321/holmesvad/videollava/mm_utils.py c3ee9d07c900dd55 |
ran
|
MIT (permissive) |
| mDPO: Conditional Preference Optimization for Multimodal Large Language Models |
17 Jun 2024 |
luka-group/mDPO/bunny/bunny_utils/util/mm_utils.py 3aada9ef083faccd |
ran · honoured contract
|
no licence file found · pointer only |
| DevBench: A multimodal developmental benchmark for language learning |
14 Jun 2024 |
alvinwmtan/dev-bench/model_classes/modeling_tinyllava_phi.py c3ee9d07c900dd55 |
ran
|
no licence file found · pointer only |
| VideoGPT+: Integrating Image and Video Encoders for Enhanced Video Understanding |
13 Jun 2024 |
mbzuai-oryx/videogpt-plus/videogpt_plus/mm_utils.py c3ee9d07c900dd55 |
ran
|
CC-BY-4.0 · pointer only |
| Multimodal Table Understanding |
12 Jun 2024 |
spursgozmy/table-llava/llava/mm_utils.py c3ee9d07c900dd55 |
ran
|
Apache-2.0 (permissive) |
| 3D-GRAND: A Million-Scale Dataset for 3D-LLMs with Better Grounding and Less Hallucination |
7 Jun 2024 |
sled-group/3D-GRAND/demo/llava/mm_utils.py c3ee9d07c900dd55 |
ran
|
no licence file found · pointer only |
| PosterLLaVa: Constructing a Unified Multi-modal Layout Generator with LLM |
5 Jun 2024 |
posterllava/posterllava/llava/mm_utils.py c3ee9d07c900dd55 |
ran
|
no licence file found · pointer only |
| Adaptive Image Quality Assessment via Teaching Large Multimodal Model to Compare |
29 May 2024 |
Q-Future/Compare2Score/q_align/mm_utils.py c3ee9d07c900dd55 |
ran
|
MIT (permissive) |
| LOVA3: Learning to Visual Question Answering, Asking and Assessment |
23 May 2024 |
showlab/LOVA3/llava/mm_utils.py c3ee9d07c900dd55 |
ran
|
no licence file found · pointer only |
| Imp: Highly Capable Large Multimodal Models for Mobile Devices |
20 May 2024 |
milvlg/imp/imp_llava/mm_utils.py c3ee9d07c900dd55 |
ran
|
Apache-2.0 (permissive) |
| Automated Multi-level Preference for MLLMs |
18 May 2024 |
takomc/amp/llava/mm_utils.py c3ee9d07c900dd55 |
ran
|
no licence file found · pointer only |
| Detecting and Mitigating Hallucination in Large Vision Language Models via Fine-Grained AI Feedback |
22 Apr 2024 |
Mr-Loevan/HSA-DPO/hsa_dpo/models/llava-v1_5/llava/mm_utils.py c3ee9d07c900dd55 |
ran
|
no licence file found · pointer only |
| Self-Supervised Visual Preference Alignment |
16 Apr 2024 |
Kevinz-code/SeVa/seva/llava/mm_utils.py c3ee9d07c900dd55 |
ran
|
GPL-3.0 (copyleft) · pointer only |
| LaSagnA: Language-based Segmentation Assistant for Complex Queries |
12 Apr 2024 |
congvvc/lasagna/model/llava/mm_utils.py 0dde2e782959c0bd |
unverified |
Apache-2.0 (permissive) |
| Unsolvable Problem Detection: Evaluating Trustworthiness of Vision Language Models |
29 Mar 2024 |
atsumiyai/upd/vlms/cogvlm/cogvlm_vqa_updbench.py c3ee9d07c900dd55 |
ran
|
Apache-2.0 (permissive) |
| VHM: Versatile and Honest Vision Language Model for Remote Sensing Image Analysis |
29 Mar 2024 |
opendatalab/h2rsvlm/vhm/mm_utils.py 0dde2e782959c0bd |
unverified |
Apache-2.0 (permissive) |
| Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models |
27 Mar 2024 |
dvlab-research/minigemini/mgm/mm_utils.py c3ee9d07c900dd55 |
ran
|
Apache-2.0 (permissive) |
| PSALM: Pixelwise SegmentAtion with Large Multi-Modal Model |
21 Mar 2024 |
zamling/PSALM/psalm/mm_utils.py c3ee9d07c900dd55 |
ran
|
Apache-2.0 (permissive) |
| Towards a clinically accessible radiology foundation model: open-access and lightweight, with automated evaluation |
12 Mar 2024 |
microsoft/llava-rad/llava/mm_utils.py c3ee9d07c900dd55 |
ran
|
no licence file found · pointer only |
| CAT: Enhancing Multimodal Large Language Model to Answer Questions in Dynamic Audio-Visual Scenarios |
7 Mar 2024 |
rikeilong/bay-cat/ADPO_CAT/mm_utils.py c3ee9d07c900dd55 |
ran
|
Apache-2.0 (permissive) |
| ImgTrojan: Jailbreaking Vision-Language Models with ONE Image |
5 Mar 2024 |
xijia-tao/imgtrojan/finetune/llava/mm_utils.py c3ee9d07c900dd55 |
ran
|
no licence file found · pointer only |
| CoCo-Agent: A Comprehensive Cognitive MLLM Agent for Smartphone GUI Automation |
19 Feb 2024 |
xbmxb/coco-agent/llava/mm_utils.py c3ee9d07c900dd55 |
ran
|
no licence file found · pointer only |
| ScreenAgent: A Vision Language Model-driven Computer Control Agent |
9 Feb 2024 |
niuzaisheng/ScreenAgent/train/cogagent_model_worker.py 6493522dff8f8fa5 |
ran
|
licence not identified · pointer only |
| MobileVLM V2: Faster and Stronger Baseline for Vision Language Model |
6 Feb 2024 |
meituan-automl/mobilevlm/mobilevlm/utils.py 3aada9ef083faccd |
ran · honoured contract
|
Apache-2.0 (permissive) |
| SciMMIR: Benchmarking Scientific Multi-modal Information Retrieval |
24 Jan 2024 |
wusiwei0410/scimmir/mm_utils.py c3ee9d07c900dd55 |
ran
|
no licence file found · pointer only |
| Prometheus-Vision: Vision-Language Model as a Judge for Fine-Grained Evaluation |
12 Jan 2024 |
kaistai/prometheus-vision/llava/mm_utils.py c3ee9d07c900dd55 |
ran
|
Apache-2.0 (permissive) |
| Eyes Wide Shut? Exploring the Visual Shortcomings of Multimodal LLMs |
11 Jan 2024 |
tsb0601/MMVP/LLaVA/llava/mm_utils.py c3ee9d07c900dd55 |
ran
|
no licence file found · pointer only |
| GroundingGPT:Language Enhanced Multi-modal Grounding Model |
11 Jan 2024 |
lzw-lzw/groundinggpt/lego/mm_utils.py c3ee9d07c900dd55 |
ran
|
Apache-2.0 (permissive) |
| MLLM-Protector: Ensuring MLLM's Safety without Hurting Performance |
5 Jan 2024 |
pipilurj/mllm-protector/llava/mm_utils.py c3ee9d07c900dd55 |
ran
|
Apache-2.0 (permissive) |
| LLaVA-Phi: Efficient Multi-Modal Assistant with Small Language Model |
4 Jan 2024 |
zhuyiche/llava-phi/llava_phi/mm_utils.py c3ee9d07c900dd55 |
ran
|
no licence file found · pointer only |
| LISA++: An Improved Baseline for Reasoning Segmentation with Large Language Model |
28 Dec 2023 |
dvlab-research/lisa/model/llava/mm_utils.py 0dde2e782959c0bd |
unverified |
Apache-2.0 (permissive) |
| VCoder: Versatile Vision Encoders for Multimodal Large Language Models |
21 Dec 2023 |
shi-labs/vcoder/vcoder_llava/mm_utils.py c3ee9d07c900dd55 |
ran
|
Apache-2.0 (permissive) |
| GSVA: Generalized Segmentation via Multimodal Large Language Models |
15 Dec 2023 |
leaplabthu/gsva/model/llava/mm_utils.py 0dde2e782959c0bd |
unverified |
Apache-2.0 (permissive) |
| Osprey: Pixel Understanding with Visual Instruction Tuning |
15 Dec 2023 |
circleradon/osprey/osprey/mm_utils.py c3ee9d07c900dd55 |
ran
|
Apache-2.0 (permissive) |
| Genixer: Empowering Multimodal Large Language Models as a Powerful Data Generator |
11 Dec 2023 |
zhaohengyuan1/genixer/Genixer_LLaVA/llava/mm_utils.py c3ee9d07c900dd55 |
ran
|
no licence file found · pointer only |
| Quilt-LLaVA: Visual Instruction Tuning by Extracting Localized Narratives from Open-Source Histopathology Videos |
7 Dec 2023 |
aldraus/quilt-llava/llava/mm_utils.py c3ee9d07c900dd55 |
ran
|
MIT (permissive) |
| LLaVA-Grounding: Grounded Visual Chat with Large Multimodal Models |
5 Dec 2023 |
ux-decoder/llava-grounding/llava/mm_utils.py c3ee9d07c900dd55 |
ran
|
Apache-2.0 (permissive) |
| VTimeLLM: Empower LLM to Grasp Video Moments |
30 Nov 2023 |
huangb23/vtimellm/vtimellm/mm_utils.py c3ee9d07c900dd55 |
ran
|
no licence file found · pointer only |
| Safe-CLIP: Removing NSFW Concepts from Vision-and-Language Models |
27 Nov 2023 |
aimagelab/safe-clip/LLaVA_generation/llava/mm_utils.py c3ee9d07c900dd55 |
ran
|
no licence file found · pointer only |
| InstructMol: Multi-Modal Integration for Building a Versatile and Reliable Molecular Assistant in Drug Discovery |
27 Nov 2023 |
idea-xl/instructmol/llava/mm_utils.py c3ee9d07c900dd55 |
ran
|
Apache-2.0 (permissive) |
| GeoChat: Grounded Large Vision-Language Model for Remote Sensing |
24 Nov 2023 |
mbzuai-oryx/geochat/geochat/mm_utils.py c3ee9d07c900dd55 |
ran
|
no licence file found · pointer only |
| Qilin-Med-VL: Towards Chinese Large Vision-Language Model for General Healthcare |
27 Oct 2023 |
williamliujl/qilin-med-vl/llava/mm_utils.py c3ee9d07c900dd55 |
ran
|
no licence file found · pointer only |
| HallE-Control: Controlling Object Hallucination in Large Multimodal Models |
3 Oct 2023 |
bronyayang/HallE_Switch/llava/mm_utils.py c3ee9d07c900dd55 |
ran
|
no licence file found · pointer only |
| Sight Beyond Text: Multi-Modal Training Enhances LLMs in Truthfulness and Ethics |
13 Sep 2023 |
ucsc-vlaa/sight-beyond-text/llava/mm_utils.py c3ee9d07c900dd55 |
ran
|
Apache-2.0 (permissive) |
| StableLLaVA: Enhanced Visual Instruction Tuning with Synthesized Image-Dialogue Data |
20 Aug 2023 |
icoz69/stablellava/llava/mm_utils.py c3ee9d07c900dd55 |
ran
|
Apache-2.0 (permissive) |
| Self-attention Does Not Need $O(n^2)$ Memory |
10 Dec 2021 |
X-iZhang/Libra/libra/mm_utils.py c3ee9d07c900dd55 |
ran
|
Apache-2.0 (permissive) |
| arXiv:Xia_GSVA_Generalized_Segmentation_via_Multimodal_Large_Language_Models_CVPR_2024_paper |
|
LeapLabTHU/GSVA/model/llava/mm_utils.py 0dde2e782959c0bd |
unverified |
Apache-2.0 (permissive) |
| arXiv:2025.naacl-long.579 |
|
DAMO-NLP-SG/VideoLLaMA2/videollama2/mm_utils.py 0dde2e782959c0bd |
unverified |
Apache-2.0 (permissive) |
| arXiv:2024.findings-naacl.226 |
|
nguyennm1024/OSCaR/llava/mm_utils.py c3ee9d07c900dd55 |
ran
|
Apache-2.0 (permissive) |
| arXiv:2024.findings-emnlp.775 |
|
YuxiXie/V-DPO/llava_dpo/mm_utils.py c3ee9d07c900dd55 |
ran
|
Apache-2.0 (permissive) |
| arXiv:2024.findings-emnlp.268 |
|
HZQ950419/Math-LLaVA/llava/mm_utils.py c3ee9d07c900dd55 |
ran
|
Apache-2.0 (permissive) |