| Evaluating the Diagnostic Classification Ability of Multimodal Large Language Models: Insights from the Osteoarthritis Initiative added by Syntology |
2026-01 (from id) |
wanglihx/LLaVA-OA/llava_trainer_weighted.py bb35e3ac741bb2c9 |
unverified |
no licence file found · pointer only |
| COIDO: Efficient Data Selection for Visual Instruction Tuning via Coupled Importance-Diversity Optimization added by Syntology |
2025-10 (from id) |
SuDIS-ZJU/CoIDO/coido_scorer/coido_trainer.py 77a16f78394b3620 |
unverified |
AGPL-3.0 (copyleft) · pointer only |
| arXiv:2507.15504 |
2025-07 (from id) |
PKU-YuanGroup/Video-LLaVA/videollava/train/llava_trainer.py bb35e3ac741bb2c9 |
unverified |
Apache-2.0 (permissive) |
| SECOND: Mitigating Perceptual Hallucination in Vision-Language Models via Selective and Contrastive Decoding |
10 Jun 2025 |
identical code first harvested elsewhere bb35e3ac741bb2c9 |
unverified |
licence of this copy not recorded |
| Daily-Omni: Towards Audio-Visual Reasoning with Temporal Alignment across Modalities |
23 May 2025 |
identical code first harvested elsewhere bb35e3ac741bb2c9 |
unverified |
licence of this copy not recorded |
| BLIP3-o: A Family of Fully Open Unified Multimodal Models-Architecture, Training and Dataset |
14 May 2025 |
identical code first harvested elsewhere bb35e3ac741bb2c9 |
unverified |
licence of this copy not recorded |
| T2I-R1: Reinforcing Image Generation with Collaborative Semantic-level and Token-level CoT |
1 May 2025 |
identical code first harvested elsewhere bb35e3ac741bb2c9 |
unverified |
licence of this copy not recorded |
| Safety Mirage: How Spurious Correlations Undermine VLM Safety Fine-tuning |
14 Mar 2025 |
OPTML-Group/VLM-Safety-Unlearn/llava/train/llava_unlearn_full_trainer.py bb35e3ac741bb2c9 |
unverified |
MIT (permissive) |
| R1-Omni: Explainable Omni-Multimodal Emotion Recognition with Reinforcement Learning |
7 Mar 2025 |
humanmllm/r1-omni/humanomni/humanomni_trainer.py 14fb23ef45b146d1 |
unverified |
no licence file found · pointer only |
| Magma: A Foundation Model for Multimodal AI Agents |
18 Feb 2025 |
microsoft/Magma/trainer/trainer.py bb35e3ac741bb2c9 |
unverified |
MIT (permissive) |
| VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding |
22 Jan 2025 |
damo-nlp-sg/videollama3/videollama3/videollama3_trainer.py 14fb23ef45b146d1 |
unverified |
Apache-2.0 (permissive) |
| Mitigating Hallucinations in Large Vision-Language Models via DPO: On-Policy Data Hold the Key |
16 Jan 2025 |
zhyang2226/opa-dpo/opadpo/opa_models/opa_trainer.py bb35e3ac741bb2c9 |
unverified |
MIT recorded; this copy not marked cleared · pointer only |
| MMedPO: Aligning Medical Vision-Language Models with Clinical-Aware Multimodal Preference Optimization |
9 Dec 2024 |
aiming-lab/mmedpo/train/dpo/llava_trainer_weighted.py bb35e3ac741bb2c9 |
unverified |
Apache-2.0 (permissive) |
| Insight-V: Exploring Long-Chain Visual Reasoning with Multimodal Large Language Models |
21 Nov 2024 |
identical code first harvested elsewhere bb35e3ac741bb2c9 |
unverified |
licence of this copy not recorded |
| PPLLaVA: Varied Video Sequence Understanding With Prompt Guidance |
4 Nov 2024 |
identical code first harvested elsewhere bb35e3ac741bb2c9 |
unverified |
licence of this copy not recorded |
| Senna: Bridging Large Vision-Language Models and End-to-End Autonomous Driving |
29 Oct 2024 |
hustvl/senna/llava/senna/senna_llava_trainer.py bb35e3ac741bb2c9 |
unverified |
Apache-2.0 (permissive) |
| LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding |
22 Oct 2024 |
Vision-CAIR/LongVU/longvu/mm_datautils.py 2ed115fa11e791e8 |
unverified |
Apache-2.0 (permissive) |
| MMed-RAG: Versatile Multimodal RAG System for Medical Vision Language Models |
16 Oct 2024 |
richard-peng-xia/MMed-RAG/train/dpo/llava_trainer_2stages.py bb35e3ac741bb2c9 |
unverified |
MIT (permissive) |
| Reconstructive Visual Instruction Tuning |
12 Oct 2024 |
haochen-wang409/ross/ross/ross_trainer.py bb35e3ac741bb2c9 |
unverified |
Apache-2.0 (permissive) |
| TRACE: Temporal Grounding Video LLM via Causal Event Modeling |
8 Oct 2024 |
gyxxyg/TRACE/trace/trace_trainer.py bb35e3ac741bb2c9 |
unverified |
Apache-2.0 (permissive) |
| FakeShield: Explainable Image Forgery Detection and Localization via Multi-modal Large Language Models |
3 Oct 2024 |
identical code first harvested elsewhere bb35e3ac741bb2c9 |
unverified |
licence of this copy not recorded |
| Ferret: Federated Full-Parameter Tuning at Scale for Large Language Models |
10 Sep 2024 |
identical code first harvested elsewhere bb35e3ac741bb2c9 |
unverified |
licence of this copy not recorded |
| Recoverable Compression: A Multimodal Vision Token Recovery Mechanism Guided by Text Information |
2 Sep 2024 |
identical code first harvested elsewhere bb35e3ac741bb2c9 |
unverified |
licence of this copy not recorded |
| LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models |
10 Jul 2024 |
identical code first harvested elsewhere bb35e3ac741bb2c9 |
unverified |
licence of this copy not recorded |
| STLLaVA-Med: Self-Training Large Language and Vision Assistant for Medical Question-Answering |
28 Jun 2024 |
heliossun/stllava-med/llava_dpo_trainer.py bb35e3ac741bb2c9 |
unverified |
MIT (permissive) |
| VoCo-LLaMA: Towards Vision Compression with Large Language Models |
18 Jun 2024 |
identical code first harvested elsewhere bb35e3ac741bb2c9 |
unverified |
licence of this copy not recorded |
| MMDU: A Multi-Turn Multi-Image Dialog Understanding Benchmark and Instruction-Tuning Dataset for LVLMs |
17 Jun 2024 |
identical code first harvested elsewhere bb35e3ac741bb2c9 |
unverified |
licence of this copy not recorded |
| Unveiling Encoder-Free Vision-Language Models |
17 Jun 2024 |
identical code first harvested elsewhere bb35e3ac741bb2c9 |
unverified |
licence of this copy not recorded |
| Parrot: Multilingual Visual Instruction Tuning |
4 Jun 2024 |
AIDC-AI/Parrot/parrot/train/parrot_trainer.py bb35e3ac741bb2c9 |
unverified |
Apache-2.0 (permissive) |
| Data-augmented phrase-level alignment for mitigating object hallucination |
28 May 2024 |
pritamqu/HALVA/llava/train/halva_trainer.py bb35e3ac741bb2c9 |
unverified |
no licence file found · pointer only |
| MoVA: Adapting Mixture of Vision Experts to Multimodal Context |
19 Apr 2024 |
identical code first harvested elsewhere bb35e3ac741bb2c9 |
unverified |
licence of this copy not recorded |
| ST-LLM: Large Language Models Are Effective Temporal Learners |
30 Mar 2024 |
identical code first harvested elsewhere bb35e3ac741bb2c9 |
unverified |
licence of this copy not recorded |
| SQ-LLaVA: Self-Questioning for Large Vision-Language Assistant |
17 Mar 2024 |
heliossun/sq-llava/llava_trainer.py bb35e3ac741bb2c9 |
unverified |
MIT (permissive) |
| Images are Achilles' Heel of Alignment: Exploiting Visual Vulnerabilities for Jailbreaking Multimodal Large Language Models |
14 Mar 2024 |
identical code first harvested elsewhere bb35e3ac741bb2c9 |
unverified |
licence of this copy not recorded |
| Your Vision-Language Model Itself Is a Strong Filter: Towards High-Quality Instruction Tuning with Data Selection |
19 Feb 2024 |
rayruibochen/self-filter/self_filter/self_filter_trainer.py bb35e3ac741bb2c9 |
unverified |
AGPL-3.0 (copyleft) · pointer only |
| LLaGA: Large Language and Graph Assistant |
13 Feb 2024 |
chenrunjin/llaga/train/llaga_trainer.py bb35e3ac741bb2c9 |
unverified |
Apache-2.0 (permissive) |
| G-LLaVA: Solving Geometric Problem with Multi-Modal Large Language Model |
18 Dec 2023 |
pipilurj/g-llava/gllava/train/llava_trainer.py bb35e3ac741bb2c9 |
unverified |
no licence file found · pointer only |
| Osprey: Pixel Understanding with Visual Instruction Tuning |
15 Dec 2023 |
identical code first harvested elsewhere bb35e3ac741bb2c9 |
unverified |
licence of this copy not recorded |
| Hallucination Augmented Contrastive Learning for Multimodal Large Language Model |
12 Dec 2023 |
identical code first harvested elsewhere bb35e3ac741bb2c9 |
unverified |
licence of this copy not recorded |
| GeoChat: Grounded Large Vision-Language Model for Remote Sensing |
24 Nov 2023 |
identical code first harvested elsewhere bb35e3ac741bb2c9 |
unverified |
licence of this copy not recorded |
| mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration |
7 Nov 2023 |
identical code first harvested elsewhere bb35e3ac741bb2c9 |
unverified |
licence of this copy not recorded |
| Improved Baselines with Visual Instruction Tuning |
5 Oct 2023 |
identical code first harvested elsewhere bb35e3ac741bb2c9 |
unverified |
licence of this copy not recorded |
| Visual Instruction Tuning |
17 Apr 2023 |
haotian-liu/LLaVA/llava/train/llava_trainer.py bb35e3ac741bb2c9 |
unverified |
Apache-2.0 (permissive) |
| Unicom: Universal and Compact Representation Learning for Image Retrieval |
12 Apr 2023 |
identical code first harvested elsewhere bb35e3ac741bb2c9 |
unverified |
licence of this copy not recorded |