| OralAgent: Integrating Reasoning, Tools, and Knowledge for Interactive Dental Image Analysis added by Syntology |
2026-05 (from id) |
isjinghao/OralAgent/oralagent/llava/mm_utils.py 908505ed68ff4871 |
ran
|
Apache-2.0 (permissive) |
| Uncertainty-Aware Knowledge Distillation for Multimodal Large Language Models added by Syntology |
2026-03 (from id) |
Jingchensun/beta-kd/mobilevlm/utils.py 592b3c1a88f93d7c |
ran · our draft was wrong
|
no licence file found · pointer only |
| Beyond Next-Token Alignment: Distilling Multimodal Large Language Models via Token Interactions added by Syntology |
2026-02 (from id) |
lchen1019/Align-TI/alignti/mm_utils.py 592b3c1a88f93d7c |
ran · our draft was wrong
|
Apache-2.0 (permissive) |
| Benchmarking Direct Preference Optimization for Medical Large Vision-Language Models added by Syntology |
2026-01 (from id) |
dmis-lab/med-vlm-dpo/inference/LLaVA-Med/llava/mm_utils.py 908505ed68ff4871 |
ran
|
no licence file found · pointer only |
| UniX: Unifying Autoregression and Diffusion for Chest X-Ray Understanding and Generation added by Syntology |
2026-01 (from id) |
ZrH42/UniX/modeling/unix_vlm/models/image_processing_vlm.py 592b3c1a88f93d7c |
ran · our draft was wrong
|
MIT (permissive) |
| Uni-X: Mitigating Modality Conflict with a Two-End-Separated Architecture for Unified Multimodal Models added by Syntology |
2025-09 (from id) |
CURRENTF/Uni-X/uni_arch/mm_utils.py 592b3c1a88f93d7c |
ran · our draft was wrong
|
MIT (permissive) |
| arXiv:2507.18300 |
2025-07 (from id) |
360CVGroup/LMM-Det/llava/mm_utils.py 592b3c1a88f93d7c |
ran · our draft was wrong
|
Apache-2.0 (permissive) |
| arXiv:2507.04976 |
2025-07 (from id) |
EsYoon7/UVQA/llava/mm_utils.py 592b3c1a88f93d7c |
ran · our draft was wrong
|
no licence file found · pointer only |
| ReasonGen-R1: CoT for Autoregressive Image generation models through SFT and RL |
30 May 2025 |
Franklin-Zhang0/ReasonGen-R1/Janus/janus/models/image_processing_vlm.py 592b3c1a88f93d7c |
ran · our draft was wrong
|
Apache-2.0 (permissive) |
| Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models |
27 May 2025 |
jefferyzhan/griffon/griffon/mm_utils.py 592b3c1a88f93d7c |
ran · our draft was wrong
|
Apache-2.0 (permissive) |
| EVEv2: Improved Baselines for Encoder-Free Vision-Language Models |
10 Feb 2025 |
baaivision/EVE/EVEv1/eve/mm_utils.py 592b3c1a88f93d7c |
ran · our draft was wrong
|
MIT (permissive) |
| MedRAX: Medical Reasoning Agent for Chest X-ray |
4 Feb 2025 |
bowang-lab/medrax/medrax/llava/mm_utils.py 908505ed68ff4871 |
ran
|
Apache-2.0 (permissive) |
| VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding |
22 Jan 2025 |
damo-nlp-sg/videollama3/videollama3/mm_utils.py 592b3c1a88f93d7c |
ran · our draft was wrong
|
Apache-2.0 (permissive) |
| InstructSeg: Unifying Instructed Visual Segmentation with Multi-modal Large Language Models |
18 Dec 2024 |
congvvc/instructseg/instructseg/model/mipha/mm_utils.py 592b3c1a88f93d7c |
ran · our draft was wrong
|
Apache-2.0 (permissive) |
| WiseAD: Knowledge Augmented End-to-End Autonomous Driving with Vision-Language Model |
13 Dec 2024 |
wyddmw/WiseAD/mobilevlm/utils.py 592b3c1a88f93d7c |
ran · our draft was wrong
|
no licence file found · pointer only |
| Liquid: Language Models are Scalable Multi-modal Generators |
5 Dec 2024 |
foundationvision/liquid/evaluation/inference_i2t.py 592b3c1a88f93d7c |
ran · our draft was wrong
|
MIT (permissive) |
| FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression |
5 Dec 2024 |
codefanw/flashsloth/flashsloth/mm_utils.py 592b3c1a88f93d7c |
ran · our draft was wrong
|
no licence file found · pointer only |
| HyperSeg: Towards Universal Visual Segmentation with Large Language Model |
26 Nov 2024 |
congvvc/HyperSeg/hyperseg/model/mipha/mm_utils.py 592b3c1a88f93d7c |
ran · our draft was wrong
|
Apache-2.0 (permissive) |
| Devils in Middle Layers of Large Vision-Language Models: Interpreting, Detecting and Mitigating Object Hallucinations via Attention Lens |
23 Nov 2024 |
zhangqijiang07/middle_layers_indicating_hallucinations/eval_data_loader.py 592b3c1a88f93d7c |
ran · our draft was wrong
|
no licence file found · pointer only |
| V-DPO: Mitigating Hallucination in Large Vision Language Models via Vision-Guided Direct Preference Optimization |
5 Nov 2024 |
yuxixie/v-dpo/llava_dpo/mm_utils.py 592b3c1a88f93d7c |
ran · our draft was wrong
|
Apache-2.0 (permissive) |
| Mitigating Modality Prior-Induced Hallucinations in Multimodal Large Language Models via Deciphering Attention Causality |
7 Oct 2024 |
The-Martyr/CausalMM/llava-1.5/experiments/llava/mm_utils.py 592b3c1a88f93d7c |
ran · our draft was wrong
|
MIT (permissive) |
| Attention Prompting on Image for Large Vision-Language Models |
25 Sep 2024 |
identical code first harvested elsewhere 592b3c1a88f93d7c |
ran · our draft was wrong
|
licence of this copy not recorded |
| CDChat: A Large Multimodal Model for Remote Sensing Change Description |
24 Sep 2024 |
techmn/cdchat/cdchat/mm_utils.py 592b3c1a88f93d7c |
ran · our draft was wrong
|
no licence file found · pointer only |
| Fit and Prune: Fast and Training-free Visual Token Pruning for Multi-modal Large Language Models |
16 Sep 2024 |
ywh187/fitprune/LLaVA_1.5/llava/mm_utils.py 592b3c1a88f93d7c |
ran · our draft was wrong
|
no licence file found · pointer only |
| ChangeChat: An Interactive Model for Remote Sensing Change Analysis via Multimodal Instruction Tuning |
13 Sep 2024 |
hanlinwu/changechat/changechat/mm_utils.py 592b3c1a88f93d7c |
ran · our draft was wrong
|
no licence file found · pointer only |
| Recoverable Compression: A Multimodal Vision Token Recovery Mechanism Guided by Text Information |
2 Sep 2024 |
banjiuyufen/Recoverable-Compression/llava/mm_utils.py 592b3c1a88f93d7c |
ran · our draft was wrong
|
Apache-2.0 (permissive) |
| IAA: Inner-Adaptor Architecture Empowers Frozen Large Language Model with Multimodal Capabilities |
23 Aug 2024 |
360cvgroup/inner-adaptor-architecture/iaa/mm_utils.py 592b3c1a88f93d7c |
ran · our draft was wrong
|
Apache-2.0 (permissive) |
| Math-PUMA: Progressive Upward Multimodal Alignment to Enhance Mathematical Reasoning |
16 Aug 2024 |
wwzhuang01/math-puma/models/deepseek_math/image_processing_vlm.py 592b3c1a88f93d7c |
ran · our draft was wrong
|
GPL-3.0 (copyleft) · pointer only |
| Paying More Attention to Image: A Training-Free Method for Alleviating Hallucination in LVLMs |
31 Jul 2024 |
hasanar1f/llava-hallunication-fix/modPAI/llava/mm_utils.py 592b3c1a88f93d7c |
ran · our draft was wrong
|
Apache-2.0 (permissive) |
| DenseFusion-1M: Merging Vision Experts for Comprehensive Multimodal Perception |
11 Jul 2024 |
baaivision/DenseFusion/densefusion/mm_utils.py 592b3c1a88f93d7c |
ran · our draft was wrong
|
no licence file found · pointer only |
| Web2Code: A Large-scale Webpage-to-Code Dataset and Evaluation Framework for Multimodal LLMs |
28 Jun 2024 |
MBZUAI-LLM/web2code/web2code/llava/mm_utils.py 592b3c1a88f93d7c |
ran · our draft was wrong
|
no licence file found · pointer only |
| Biomedical Visual Instruction Tuning with Clinician Preference Alignment |
19 Jun 2024 |
mao1207/BioMed-VITAL/backbone/mm_utils.py 908505ed68ff4871 |
ran
|
no licence file found · pointer only |
| Holmes-VAD: Towards Unbiased and Explainable Video Anomaly Detection via Multi-modal LLM |
18 Jun 2024 |
pipixin321/holmesvad/videollava/mm_utils.py 592b3c1a88f93d7c |
ran · our draft was wrong
|
MIT (permissive) |
| mDPO: Conditional Preference Optimization for Multimodal Large Language Models |
17 Jun 2024 |
luka-group/mDPO/bunny/bunny_utils/util/mm_utils.py 592b3c1a88f93d7c |
ran · our draft was wrong
|
no licence file found · pointer only |
| DevBench: A multimodal developmental benchmark for language learning |
14 Jun 2024 |
alvinwmtan/dev-bench/model_classes/modeling_tinyllava_phi.py 592b3c1a88f93d7c |
ran · our draft was wrong
|
no licence file found · pointer only |
| Multimodal Table Understanding |
12 Jun 2024 |
spursgozmy/table-llava/llava/mm_utils.py 592b3c1a88f93d7c |
ran · our draft was wrong
|
Apache-2.0 (permissive) |
| PosterLLaVa: Constructing a Unified Multi-modal Layout Generator with LLM |
5 Jun 2024 |
posterllava/posterllava/llava/mm_utils.py 592b3c1a88f93d7c |
ran · our draft was wrong
|
no licence file found · pointer only |
| Adaptive Image Quality Assessment via Teaching Large Multimodal Model to Compare |
29 May 2024 |
Q-Future/Compare2Score/q_align/mm_utils.py 592b3c1a88f93d7c |
ran · our draft was wrong
|
MIT (permissive) |
| Adaptive Image Quality Assessment via Teaching Large Multimodal Model to Compare |
29 May 2024 |
Q-Future/Compare2Score/q_align/model/modeling_mplug_owl2.py 0d475d995befa10b |
ran
|
MIT (permissive) |
| ConvLLaVA: Hierarchical Backbones as Visual Encoder for Large Multimodal Models |
24 May 2024 |
alibaba/conv-llava/llava/eval/evaluate_grounding.py 592b3c1a88f93d7c |
ran · our draft was wrong
|
Apache-2.0 (permissive) |
| LOVA3: Learning to Visual Question Answering, Asking and Assessment |
23 May 2024 |
showlab/LOVA3/llava/mm_utils.py 592b3c1a88f93d7c |
ran · our draft was wrong
|
no licence file found · pointer only |
| Imp: Highly Capable Large Multimodal Models for Mobile Devices |
20 May 2024 |
milvlg/imp/imp_llava/mm_utils.py 592b3c1a88f93d7c |
ran · our draft was wrong
|
Apache-2.0 (permissive) |
| Automated Multi-level Preference for MLLMs |
18 May 2024 |
takomc/amp/llava/mm_utils.py 592b3c1a88f93d7c |
ran · our draft was wrong
|
no licence file found · pointer only |
| GSCo: Towards Generalizable AI in Medicine via Generalist-Specialist Collaboration |
23 Apr 2024 |
sunanhe/meddr/src/dataset/transforms.py 592b3c1a88f93d7c |
ran · our draft was wrong
|
MIT (permissive) |
| Detecting and Mitigating Hallucination in Large Vision Language Models via Fine-Grained AI Feedback |
22 Apr 2024 |
Mr-Loevan/HSA-DPO/hsa_dpo/models/llava-v1_5/llava/mm_utils.py 592b3c1a88f93d7c |
ran · our draft was wrong
|
no licence file found · pointer only |
| Self-Supervised Visual Preference Alignment |
16 Apr 2024 |
Kevinz-code/SeVa/seva/llava/mm_utils.py 592b3c1a88f93d7c |
ran · our draft was wrong
|
GPL-3.0 (copyleft) · pointer only |
| VHM: Versatile and Honest Vision Language Model for Remote Sensing Image Analysis |
29 Mar 2024 |
opendatalab/h2rsvlm/vhm/mm_utils.py 592b3c1a88f93d7c |
ran · our draft was wrong
|
Apache-2.0 (permissive) |
| Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models |
27 Mar 2024 |
dvlab-research/minigemini/mgm/mm_utils.py 592b3c1a88f93d7c |
ran · our draft was wrong
|
Apache-2.0 (permissive) |
| Towards a clinically accessible radiology foundation model: open-access and lightweight, with automated evaluation |
12 Mar 2024 |
microsoft/llava-rad/llava/mm_utils.py 592b3c1a88f93d7c |
ran · our draft was wrong
|
no licence file found · pointer only |
| DeepSeek-VL: Towards Real-World Vision-Language Understanding |
8 Mar 2024 |
deepseek-ai/deepseek-vl/deepseek_vl/models/image_processing_vlm.py 592b3c1a88f93d7c |
ran · our draft was wrong
|
MIT (permissive) |
| ImgTrojan: Jailbreaking Vision-Language Models with ONE Image |
5 Mar 2024 |
xijia-tao/imgtrojan/finetune/llava/mm_utils.py 592b3c1a88f93d7c |
ran · our draft was wrong
|
no licence file found · pointer only |
| MobileVLM V2: Faster and Stronger Baseline for Vision Language Model |
6 Feb 2024 |
meituan-automl/mobilevlm/mobilevlm/utils.py 592b3c1a88f93d7c |
ran · our draft was wrong
|
Apache-2.0 (permissive) |
| SciMMIR: Benchmarking Scientific Multi-modal Information Retrieval |
24 Jan 2024 |
wusiwei0410/scimmir/mm_utils.py 592b3c1a88f93d7c |
ran · our draft was wrong
|
no licence file found · pointer only |
| Prometheus-Vision: Vision-Language Model as a Judge for Fine-Grained Evaluation |
12 Jan 2024 |
kaistai/prometheus-vision/llava/mm_utils.py 592b3c1a88f93d7c |
ran · our draft was wrong
|
Apache-2.0 (permissive) |
| Eyes Wide Shut? Exploring the Visual Shortcomings of Multimodal LLMs |
11 Jan 2024 |
tsb0601/MMVP/LLaVA/llava/mm_utils.py 592b3c1a88f93d7c |
ran · our draft was wrong
|
no licence file found · pointer only |
| MLLM-Protector: Ensuring MLLM's Safety without Hurting Performance |
5 Jan 2024 |
pipilurj/mllm-protector/llava/mm_utils.py 592b3c1a88f93d7c |
ran · our draft was wrong
|
Apache-2.0 (permissive) |
| LLaVA-Phi: Efficient Multi-Modal Assistant with Small Language Model |
4 Jan 2024 |
zhuyiche/llava-phi/llava_phi/mm_utils.py 592b3c1a88f93d7c |
ran · our draft was wrong
|
no licence file found · pointer only |
| VCoder: Versatile Vision Encoders for Multimodal Large Language Models |
21 Dec 2023 |
shi-labs/vcoder/vcoder_llava/mm_utils.py 592b3c1a88f93d7c |
ran · our draft was wrong
|
Apache-2.0 (permissive) |
| V*: Guided Visual Search as a Core Mechanism in Multimodal LLMs |
21 Dec 2023 |
penghao-wu/vstar/vstar_bench_eval.py e02a1c976f061bb2 |
ran · our draft was wrong
|
MIT (permissive) |
| Osprey: Pixel Understanding with Visual Instruction Tuning |
15 Dec 2023 |
circleradon/osprey/osprey/mm_utils.py 592b3c1a88f93d7c |
ran · our draft was wrong
|
Apache-2.0 (permissive) |
| Genixer: Empowering Multimodal Large Language Models as a Powerful Data Generator |
11 Dec 2023 |
zhaohengyuan1/genixer/Genixer_LLaVA/llava/mm_utils.py 592b3c1a88f93d7c |
ran · our draft was wrong
|
no licence file found · pointer only |
| Quilt-LLaVA: Visual Instruction Tuning by Extracting Localized Narratives from Open-Source Histopathology Videos |
7 Dec 2023 |
aldraus/quilt-llava/llava/mm_utils.py 592b3c1a88f93d7c |
ran · our draft was wrong
|
MIT (permissive) |
| Safe-CLIP: Removing NSFW Concepts from Vision-and-Language Models |
27 Nov 2023 |
aimagelab/safe-clip/LLaVA_generation/llava/mm_utils.py 592b3c1a88f93d7c |
ran · our draft was wrong
|
no licence file found · pointer only |
| GeoChat: Grounded Large Vision-Language Model for Remote Sensing |
24 Nov 2023 |
mbzuai-oryx/geochat/geochat/mm_utils.py 592b3c1a88f93d7c |
ran · our draft was wrong
|
no licence file found · pointer only |
| Towards Improving Document Understanding: An Exploration on Text-Grounding via MLLMs |
22 Nov 2023 |
harrytea/tgdoc/tgdoc/mm_utils.py 8c5ec74733471314 |
ran
|
no licence file found · pointer only |
| Qilin-Med-VL: Towards Chinese Large Vision-Language Model for General Healthcare |
27 Oct 2023 |
williamliujl/qilin-med-vl/llava/mm_utils.py 592b3c1a88f93d7c |
ran · our draft was wrong
|
no licence file found · pointer only |
| Zero123++: a Single Image to Consistent Multi-view Diffusion Base Model |
23 Oct 2023 |
sudo-ai-3d/zero123plus/gradio_app.py 592b3c1a88f93d7c |
ran · our draft was wrong
|
Apache-2.0 (permissive) |
| Vision-by-Language for Training-Free Compositional Image Retrieval |
13 Oct 2023 |
explainableml/vision_by_language/src/datasets.py e75146036d2d1504 |
ran
|
MIT (permissive) |
| HallE-Control: Controlling Object Hallucination in Large Multimodal Models |
3 Oct 2023 |
bronyayang/HallE_Switch/llava/mm_utils.py 592b3c1a88f93d7c |
ran · our draft was wrong
|
no licence file found · pointer only |
| Pink: Unveiling the Power of Referential Comprehension for Multi-modal LLMs |
1 Oct 2023 |
sy-xuan/pink/pink/eval/model_gqa.py ee1020522a070904 |
ran
|
no licence file found · pointer only |
| StableLLaVA: Enhanced Visual Instruction Tuning with Synthesized Image-Dialogue Data |
20 Aug 2023 |
icoz69/stablellava/llava/mm_utils.py 592b3c1a88f93d7c |
ran · our draft was wrong
|
Apache-2.0 (permissive) |
| TOPIQ: A Top-down Approach from Semantics to Distortions for Image Quality Assessment |
6 Aug 2023 |
chaofengc/iqa-pytorch/pyiqa/archs/compare2score_arch.py 66c27cd25f8c24da |
unverified |
licence not identified · pointer only |
| End-to-End Zero-Shot HOI Detection via Vision and Language Knowledge Distillation |
1 Apr 2022 |
mrwu-mac/EoID/models/EoID.py 592b3c1a88f93d7c |
ran · our draft was wrong
|
Apache-2.0 (permissive) |
| Self-attention Does Not Need $O(n^2)$ Memory |
10 Dec 2021 |
jihaonew/mm-instruct/llava/mm_utils.py 592b3c1a88f93d7c |
ran · our draft was wrong
|
Apache-2.0 (permissive) |
| Self-attention Does Not Need $O(n^2)$ Memory |
10 Dec 2021 |
X-iZhang/Libra/libra/mm_utils.py 1fe910f8046380d4 |
unverified |
Apache-2.0 (permissive) |
| arXiv:Zhang_Learning_Rain_Location_Prior_for_Nighttime_Deraining_ICCV_2023_paper |
|
zkawfanx/RLP/rlp/utils.py 184826b4b0deed67 |
unverified |
MIT (permissive) |
| arXiv:2025.naacl-long.579 |
|
DAMO-NLP-SG/VideoLLaMA2/videollama2/mm_utils.py 592b3c1a88f93d7c |
ran · our draft was wrong
|
Apache-2.0 (permissive) |
| arXiv:2024.findings-naacl.226 |
|
nguyennm1024/OSCaR/llava/mm_utils.py 592b3c1a88f93d7c |
ran · our draft was wrong
|
Apache-2.0 (permissive) |
| arXiv:2024.findings-emnlp.775 |
|
YuxiXie/V-DPO/llava_dpo/mm_utils.py 592b3c1a88f93d7c |
ran · our draft was wrong
|
Apache-2.0 (permissive) |
| arXiv:2024.findings-emnlp.268 |
|
HZQ950419/Math-LLaVA/llava/mm_utils.py 592b3c1a88f93d7c |
ran · our draft was wrong
|
Apache-2.0 (permissive) |