| OralAgent: Integrating Reasoning, Tools, and Knowledge for Interactive Dental Image Analysis added by Syntology |
2026-05 (from id) |
isjinghao/OralAgent/oralagent/llava/mm_utils.py f0f8c9c5a77c9b5b |
ran
|
Apache-2.0 (permissive) |
| CLIP Tricks You: Training-free Token Pruning for Efficient Pixel Grounding in Large Vision-Language Models added by Syntology |
2026-05 (from id) |
sejong-rcv/LiteLVLM/model/llava/mm_utils.py 1df990c375318896 |
ran
|
Apache-2.0 (permissive) |
| Logit-Attention Divergence: Mitigating Position Bias in Multi-Image Retrieval via Attention-Guided Calibration added by Syntology |
2026-05 (from id) |
brightXian/LAD/src/agd.py f53971fbc29934cc |
ran · honoured contract
|
no licence file found · pointer only |
| Uncertainty-Aware Knowledge Distillation for Multimodal Large Language Models added by Syntology |
2026-03 (from id) |
Jingchensun/beta-kd/mobilevlm/utils.py e8f3a2a1be36a4d7 |
ran · our draft was wrong
|
no licence file found · pointer only |
| Beyond Next-Token Alignment: Distilling Multimodal Large Language Models via Token Interactions added by Syntology |
2026-02 (from id) |
lchen1019/Align-TI/alignti/mm_utils.py 344dff4791fd1381 |
unverified |
Apache-2.0 (permissive) |
| Benchmarking Direct Preference Optimization for Medical Large Vision-Language Models added by Syntology |
2026-01 (from id) |
dmis-lab/med-vlm-dpo/inference/LLaVA-Med/llava/mm_utils.py f0f8c9c5a77c9b5b |
ran
|
no licence file found · pointer only |
| Uni-X: Mitigating Modality Conflict with a Two-End-Separated Architecture for Unified Multimodal Models added by Syntology |
2025-09 (from id) |
CURRENTF/Uni-X/uni_arch/mm_utils.py d6153bc4456b4b0c |
unverified |
MIT (permissive) |
| Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models |
27 May 2025 |
jefferyzhan/griffon/griffon/mm_utils.py 29489f5332508b33 |
unverified |
Apache-2.0 (permissive) |
| SegEarth-R1: Geospatial Pixel Reasoning via Large Language Model |
13 Apr 2025 |
earth-insights/segearth-r1/segearth_r1/mm_utils.py 1df990c375318896 |
ran
|
Apache-2.0 (permissive) |
| On Large Multimodal Models as Open-World Image Classifiers |
27 Mar 2025 |
altndrr/lmms-owc/src/models/_instructblip.py 24156f8626f7fa5d |
unverified |
MIT (permissive) |
| EVEv2: Improved Baselines for Encoder-Free Vision-Language Models |
10 Feb 2025 |
baaivision/EVE/EVEv1/eve/mm_utils.py a87d7637b728e681 |
unverified |
MIT (permissive) |
| MedRAX: Medical Reasoning Agent for Chest X-ray |
4 Feb 2025 |
bowang-lab/medrax/medrax/llava/mm_utils.py f0f8c9c5a77c9b5b |
ran
|
Apache-2.0 (permissive) |
| Dispider: Enabling Video LLMs with Active Real-Time Interaction via Disentangled Perception, Decision, and Reaction |
6 Jan 2025 |
Mark12Ding/Dispider/dispider/mm_utils.py 18015a195a7c0e68 |
unverified |
Apache-2.0 (permissive) |
| InstructSeg: Unifying Instructed Visual Segmentation with Multi-modal Large Language Models |
18 Dec 2024 |
congvvc/instructseg/instructseg/model/mipha/mm_utils.py 344dff4791fd1381 |
unverified |
Apache-2.0 (permissive) |
| WiseAD: Knowledge Augmented End-to-End Autonomous Driving with Vision-Language Model |
13 Dec 2024 |
wyddmw/WiseAD/mobilevlm/utils.py e8f3a2a1be36a4d7 |
ran · our draft was wrong
|
no licence file found · pointer only |
| LongVALE: Vision-Audio-Language-Event Benchmark Towards Time-Aware Omni-Modal Perception of Long Videos |
29 Nov 2024 |
ttgeng233/LongVALE/longvalellm/mm_utils.py 1df990c375318896 |
ran
|
MIT (permissive) |
| HyperSeg: Towards Universal Visual Segmentation with Large Language Model |
26 Nov 2024 |
congvvc/HyperSeg/hyperseg/model/mipha/mm_utils.py 344dff4791fd1381 |
unverified |
Apache-2.0 (permissive) |
| V-DPO: Mitigating Hallucination in Large Vision Language Models via Vision-Guided Direct Preference Optimization |
5 Nov 2024 |
yuxixie/v-dpo/llava_dpo/mm_utils.py 344dff4791fd1381 |
unverified |
Apache-2.0 (permissive) |
| Mitigating Modality Prior-Induced Hallucinations in Multimodal Large Language Models via Deciphering Attention Causality |
7 Oct 2024 |
The-Martyr/CausalMM/llava-1.5/experiments/llava/mm_utils.py 344dff4791fd1381 |
unverified |
MIT (permissive) |
| Attention Prompting on Image for Large Vision-Language Models |
25 Sep 2024 |
yu-rp/apiprompting/API/API_LLaVA/functions.py e8f3a2a1be36a4d7 |
ran · our draft was wrong
|
MIT (permissive) |
| CDChat: A Large Multimodal Model for Remote Sensing Change Description |
24 Sep 2024 |
techmn/cdchat/cdchat/mm_utils.py ab382f4b6886dee2 |
ran
|
no licence file found · pointer only |
| M$^2$PT: Multimodal Prompt Tuning for Zero-shot Instruction Learning |
24 Sep 2024 |
william-wang618/mmpt-emnlp2024/M2PT/mm_utils.py 1df990c375318896 |
ran
|
no licence file found · pointer only |
| Fit and Prune: Fast and Training-free Visual Token Pruning for Multi-modal Large Language Models |
16 Sep 2024 |
ywh187/fitprune/LLaVA_1.5/llava/mm_utils.py 344dff4791fd1381 |
unverified |
no licence file found · pointer only |
| ChangeChat: An Interactive Model for Remote Sensing Change Analysis via Multimodal Instruction Tuning |
13 Sep 2024 |
hanlinwu/changechat/changechat/mm_utils.py bec15ff45f271129 |
ran
|
no licence file found · pointer only |
| Recoverable Compression: A Multimodal Vision Token Recovery Mechanism Guided by Text Information |
2 Sep 2024 |
banjiuyufen/Recoverable-Compression/llava/mm_utils.py 344dff4791fd1381 |
unverified |
Apache-2.0 (permissive) |
| FlowRetrieval: Flow-Guided Data Retrieval for Few-Shot Imitation Learning |
29 Aug 2024 |
lihenglin/bridge_training_code/data_processing/bridgedata_raw_to_numpy.py d2537349eee21015 |
unverified |
MIT (permissive) |
| CoGen: Learning from Feedback with Coupled Comprehension and Generation |
28 Aug 2024 |
lil-lab/cogen/models/joint_inference.py 5f621459c8e0f985 |
ran · honoured contract
|
no licence file found · pointer only |
| IAA: Inner-Adaptor Architecture Empowers Frozen Large Language Model with Multimodal Capabilities |
23 Aug 2024 |
360cvgroup/inner-adaptor-architecture/iaa/mm_utils.py 344dff4791fd1381 |
unverified |
Apache-2.0 (permissive) |
| Paying More Attention to Image: A Training-Free Method for Alleviating Hallucination in LVLMs |
31 Jul 2024 |
hasanar1f/llava-hallunication-fix/modPAI/llava/mm_utils.py 344dff4791fd1381 |
unverified |
Apache-2.0 (permissive) |
| DenseFusion-1M: Merging Vision Experts for Comprehensive Multimodal Perception |
11 Jul 2024 |
baaivision/DenseFusion/densefusion/mm_utils.py e8f3a2a1be36a4d7 |
ran · our draft was wrong
|
no licence file found · pointer only |
| Web2Code: A Large-scale Webpage-to-Code Dataset and Evaluation Framework for Multimodal LLMs |
28 Jun 2024 |
MBZUAI-LLM/web2code/web2code/llava/mm_utils.py 344dff4791fd1381 |
unverified |
no licence file found · pointer only |
| Biomedical Visual Instruction Tuning with Clinician Preference Alignment |
19 Jun 2024 |
mao1207/BioMed-VITAL/backbone/mm_utils.py f0f8c9c5a77c9b5b |
ran
|
no licence file found · pointer only |
| Holmes-VAD: Towards Unbiased and Explainable Video Anomaly Detection via Multi-modal LLM |
18 Jun 2024 |
pipixin321/holmesvad/videollava/mm_utils.py 344dff4791fd1381 |
unverified |
MIT (permissive) |
| mDPO: Conditional Preference Optimization for Multimodal Large Language Models |
17 Jun 2024 |
luka-group/mDPO/bunny/bunny_utils/util/mm_utils.py e8f3a2a1be36a4d7 |
ran · our draft was wrong
|
no licence file found · pointer only |
| DevBench: A multimodal developmental benchmark for language learning |
14 Jun 2024 |
alvinwmtan/dev-bench/model_classes/modeling_tinyllava_phi.py 58a932ffe36876f1 |
ran
|
no licence file found · pointer only |
| VideoGPT+: Integrating Image and Video Encoders for Enhanced Video Understanding |
13 Jun 2024 |
mbzuai-oryx/videogpt-plus/videogpt_plus/mm_utils.py 1df990c375318896 |
ran
|
CC-BY-4.0 · pointer only |
| Multimodal Table Understanding |
12 Jun 2024 |
spursgozmy/table-llava/llava/mm_utils.py 344dff4791fd1381 |
unverified |
Apache-2.0 (permissive) |
| PosterLLaVa: Constructing a Unified Multi-modal Layout Generator with LLM |
5 Jun 2024 |
posterllava/posterllava/llava/mm_utils.py 344dff4791fd1381 |
unverified |
no licence file found · pointer only |
| Adaptive Image Quality Assessment via Teaching Large Multimodal Model to Compare |
29 May 2024 |
Q-Future/Compare2Score/q_align/mm_utils.py af0990e5a0f15a4d |
ran
|
MIT (permissive) |
| LOVA3: Learning to Visual Question Answering, Asking and Assessment |
23 May 2024 |
showlab/LOVA3/llava/mm_utils.py 344dff4791fd1381 |
unverified |
no licence file found · pointer only |
| Imp: Highly Capable Large Multimodal Models for Mobile Devices |
20 May 2024 |
milvlg/imp/imp_llava/mm_utils.py 344dff4791fd1381 |
unverified |
Apache-2.0 (permissive) |
| Automated Multi-level Preference for MLLMs |
18 May 2024 |
takomc/amp/llava/mm_utils.py 344dff4791fd1381 |
unverified |
no licence file found · pointer only |
| Detecting and Mitigating Hallucination in Large Vision Language Models via Fine-Grained AI Feedback |
22 Apr 2024 |
Mr-Loevan/HSA-DPO/hsa_dpo/models/llava-v1_5/llava/mm_utils.py 344dff4791fd1381 |
unverified |
no licence file found · pointer only |
| Self-Supervised Visual Preference Alignment |
16 Apr 2024 |
Kevinz-code/SeVa/seva/llava/mm_utils.py 344dff4791fd1381 |
unverified |
GPL-3.0 (copyleft) · pointer only |
| LaSagnA: Language-based Segmentation Assistant for Complex Queries |
12 Apr 2024 |
congvvc/lasagna/model/llava/mm_utils.py 1df990c375318896 |
ran
|
Apache-2.0 (permissive) |
| VHM: Versatile and Honest Vision Language Model for Remote Sensing Image Analysis |
29 Mar 2024 |
opendatalab/h2rsvlm/vhm/mm_utils.py e8f3a2a1be36a4d7 |
ran · our draft was wrong
|
Apache-2.0 (permissive) |
| Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models |
27 Mar 2024 |
dvlab-research/minigemini/mgm/mm_utils.py d6153bc4456b4b0c |
unverified |
Apache-2.0 (permissive) |
| PSALM: Pixelwise SegmentAtion with Large Multi-Modal Model |
21 Mar 2024 |
zamling/PSALM/psalm/mm_utils.py 1df990c375318896 |
ran
|
Apache-2.0 (permissive) |
| Towards a clinically accessible radiology foundation model: open-access and lightweight, with automated evaluation |
12 Mar 2024 |
microsoft/llava-rad/llava/mm_utils.py 344dff4791fd1381 |
unverified |
no licence file found · pointer only |
| CAT: Enhancing Multimodal Large Language Model to Answer Questions in Dynamic Audio-Visual Scenarios |
7 Mar 2024 |
rikeilong/bay-cat/ADPO_CAT/mm_utils.py 1df990c375318896 |
ran
|
Apache-2.0 (permissive) |
| ImgTrojan: Jailbreaking Vision-Language Models with ONE Image |
5 Mar 2024 |
xijia-tao/imgtrojan/finetune/llava/mm_utils.py 344dff4791fd1381 |
unverified |
no licence file found · pointer only |
| CoCo-Agent: A Comprehensive Cognitive MLLM Agent for Smartphone GUI Automation |
19 Feb 2024 |
xbmxb/coco-agent/llava/mm_utils.py 1df990c375318896 |
ran
|
no licence file found · pointer only |
| MobileVLM V2: Faster and Stronger Baseline for Vision Language Model |
6 Feb 2024 |
meituan-automl/mobilevlm/mobilevlm/utils.py e8f3a2a1be36a4d7 |
ran · our draft was wrong
|
Apache-2.0 (permissive) |
| SciMMIR: Benchmarking Scientific Multi-modal Information Retrieval |
24 Jan 2024 |
wusiwei0410/scimmir/mm_utils.py af0990e5a0f15a4d |
ran
|
no licence file found · pointer only |
| SciMMIR: Benchmarking Scientific Multi-modal Information Retrieval |
24 Jan 2024 |
wusiwei0410/scimmir/LLMs_Embedding.py 7278092ebcd87065 |
ran
|
no licence file found · pointer only |
| Prometheus-Vision: Vision-Language Model as a Judge for Fine-Grained Evaluation |
12 Jan 2024 |
kaistai/prometheus-vision/llava/mm_utils.py 344dff4791fd1381 |
unverified |
Apache-2.0 (permissive) |
| Eyes Wide Shut? Exploring the Visual Shortcomings of Multimodal LLMs |
11 Jan 2024 |
tsb0601/MMVP/LLaVA/llava/mm_utils.py 344dff4791fd1381 |
unverified |
no licence file found · pointer only |
| MLLM-Protector: Ensuring MLLM's Safety without Hurting Performance |
5 Jan 2024 |
pipilurj/mllm-protector/llava/mm_utils.py 344dff4791fd1381 |
unverified |
Apache-2.0 (permissive) |
| LLaVA-Phi: Efficient Multi-Modal Assistant with Small Language Model |
4 Jan 2024 |
zhuyiche/llava-phi/llava_phi/mm_utils.py 344dff4791fd1381 |
unverified |
no licence file found · pointer only |
| LISA++: An Improved Baseline for Reasoning Segmentation with Large Language Model |
28 Dec 2023 |
dvlab-research/lisa/model/llava/mm_utils.py 1df990c375318896 |
ran
|
Apache-2.0 (permissive) |
| VCoder: Versatile Vision Encoders for Multimodal Large Language Models |
21 Dec 2023 |
shi-labs/vcoder/vcoder_llava/mm_utils.py 344dff4791fd1381 |
unverified |
Apache-2.0 (permissive) |
| GSVA: Generalized Segmentation via Multimodal Large Language Models |
15 Dec 2023 |
leaplabthu/gsva/model/llava/mm_utils.py 1df990c375318896 |
ran
|
Apache-2.0 (permissive) |
| Osprey: Pixel Understanding with Visual Instruction Tuning |
15 Dec 2023 |
circleradon/osprey/osprey/mm_utils.py 344dff4791fd1381 |
unverified |
Apache-2.0 (permissive) |
| Genixer: Empowering Multimodal Large Language Models as a Powerful Data Generator |
11 Dec 2023 |
zhaohengyuan1/genixer/Genixer_LLaVA/llava/mm_utils.py 344dff4791fd1381 |
unverified |
no licence file found · pointer only |
| Quilt-LLaVA: Visual Instruction Tuning by Extracting Localized Narratives from Open-Source Histopathology Videos |
7 Dec 2023 |
aldraus/quilt-llava/llava/mm_utils.py 344dff4791fd1381 |
unverified |
MIT (permissive) |
| LLaVA-Grounding: Grounded Visual Chat with Large Multimodal Models |
5 Dec 2023 |
ux-decoder/llava-grounding/llava/mm_utils.py 1df990c375318896 |
ran
|
Apache-2.0 (permissive) |
| VTimeLLM: Empower LLM to Grasp Video Moments |
30 Nov 2023 |
huangb23/vtimellm/vtimellm/mm_utils.py 1df990c375318896 |
ran
|
no licence file found · pointer only |
| Safe-CLIP: Removing NSFW Concepts from Vision-and-Language Models |
27 Nov 2023 |
aimagelab/safe-clip/LLaVA_generation/llava/mm_utils.py 344dff4791fd1381 |
unverified |
no licence file found · pointer only |
| InstructMol: Multi-Modal Integration for Building a Versatile and Reliable Molecular Assistant in Drug Discovery |
27 Nov 2023 |
idea-xl/instructmol/llava/mm_utils.py 1df990c375318896 |
ran
|
Apache-2.0 (permissive) |
| GeoChat: Grounded Large Vision-Language Model for Remote Sensing |
24 Nov 2023 |
mbzuai-oryx/geochat/geochat/mm_utils.py bec15ff45f271129 |
ran
|
no licence file found · pointer only |
| Qilin-Med-VL: Towards Chinese Large Vision-Language Model for General Healthcare |
27 Oct 2023 |
williamliujl/qilin-med-vl/llava/mm_utils.py 344dff4791fd1381 |
unverified |
no licence file found · pointer only |
| HallE-Control: Controlling Object Hallucination in Large Multimodal Models |
3 Oct 2023 |
bronyayang/HallE_Switch/llava/mm_utils.py 344dff4791fd1381 |
unverified |
no licence file found · pointer only |
| Sight Beyond Text: Multi-Modal Training Enhances LLMs in Truthfulness and Ethics |
13 Sep 2023 |
ucsc-vlaa/sight-beyond-text/llava/mm_utils.py 1df990c375318896 |
ran
|
Apache-2.0 (permissive) |
| BridgeData V2: A Dataset for Robot Learning at Scale |
24 Aug 2023 |
rail-berkeley/BridgeData-V2/data_processing/bridgedata_raw_to_numpy.py d2537349eee21015 |
unverified |
MIT (permissive) |
| StableLLaVA: Enhanced Visual Instruction Tuning with Synthesized Image-Dialogue Data |
20 Aug 2023 |
icoz69/stablellava/llava/mm_utils.py 344dff4791fd1381 |
unverified |
Apache-2.0 (permissive) |
| Self-attention Does Not Need $O(n^2)$ Memory |
10 Dec 2021 |
jihaonew/mm-instruct/llava/mm_utils.py 344dff4791fd1381 |
unverified |
Apache-2.0 (permissive) |
| Self-attention Does Not Need $O(n^2)$ Memory |
10 Dec 2021 |
X-iZhang/Libra/libra/mm_utils.py 012fae59ee4131ba |
unverified |
Apache-2.0 (permissive) |
| arXiv:Xia_GSVA_Generalized_Segmentation_via_Multimodal_Large_Language_Models_CVPR_2024_paper |
|
LeapLabTHU/GSVA/model/llava/mm_utils.py 1df990c375318896 |
ran
|
Apache-2.0 (permissive) |
| arXiv:2024.findings-naacl.226 |
|
nguyennm1024/OSCaR/llava/mm_utils.py 344dff4791fd1381 |
unverified |
Apache-2.0 (permissive) |
| arXiv:2024.findings-emnlp.775 |
|
YuxiXie/V-DPO/llava_dpo/mm_utils.py 344dff4791fd1381 |
unverified |
Apache-2.0 (permissive) |
| arXiv:2024.findings-emnlp.268 |
|
HZQ950419/Math-LLaVA/llava/mm_utils.py 344dff4791fd1381 |
unverified |
Apache-2.0 (permissive) |