| Training-Free Hashing-Based Attention via Binary Principal Components added by Syntology |
2026-08 (from id) |
identical code first harvested elsewhere 616ffbdc154ed2d8 |
unverified |
licence of this copy not recorded |
| CP-MoE: Consistency-Preserving Mixture-of-Experts for Continual Learning added by Syntology |
2026-05 (from id) |
YangLiu-Lewis/CP-MoE/llava/train/train_MOE.py 616ffbdc154ed2d8 |
unverified |
MIT (permissive) |
| Reasoning Portability: Guiding Continual Learning for MLLMs in the RLVR Era added by Syntology |
2026-05 (from id) |
lluosi/RDB-CL/ETrain/Train/Base_trainer.py 616ffbdc154ed2d8 |
unverified |
no licence file found · pointer only |
| Out of the Memory Barrier: A Highly Memory Efficient Training System for LLMs with Million-Token Contexts added by Syntology |
2026-02 (from id) |
wenhaoli-xmu/OOMB/chunkoptim/modifier.py 616ffbdc154ed2d8 |
unverified |
no licence file found · pointer only |
| Evaluating the Diagnostic Classification Ability of Multimodal Large Language Models: Insights from the Osteoarthritis Initiative added by Syntology |
2026-01 (from id) |
wanglihx/LLaVA-OA/llava_trainer_weighted.py 735025744c1ab0cf |
unverified |
no licence file found · pointer only |
| Evaluating the Diagnostic Classification Ability of Multimodal Large Language Models: Insights from the Osteoarthritis Initiative added by Syntology |
2026-01 (from id) |
wanglihx/LLaVA-OA/train_weighted.py 616ffbdc154ed2d8 |
unverified |
no licence file found · pointer only |
| COIDO: Efficient Data Selection for Visual Instruction Tuning via Coupled Importance-Diversity Optimization added by Syntology |
2025-10 (from id) |
SuDIS-ZJU/CoIDO/coido_scorer/coido_trainer.py 64c18975ebd3ed40 |
unverified |
AGPL-3.0 (copyleft) · pointer only |
| COIDO: Efficient Data Selection for Visual Instruction Tuning via Coupled Importance-Diversity Optimization added by Syntology |
2025-10 (from id) |
SuDIS-ZJU/CoIDO/coido_scorer/stage1.py 46e09ebb15e4afb6 |
unverified |
AGPL-3.0 (copyleft) · pointer only |
| arXiv:2507.15504 |
2025-07 (from id) |
PKU-YuanGroup/Video-LLaVA/videollava/train/llava_trainer.py 735025744c1ab0cf |
unverified |
Apache-2.0 (permissive) |
| SECOND: Mitigating Perceptual Hallucination in Vision-Language Models via Selective and Contrastive Decoding |
10 Jun 2025 |
identical code first harvested elsewhere 735025744c1ab0cf |
unverified |
licence of this copy not recorded |
| ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge |
28 May 2025 |
tutujingyugang1/ChatVLA_public/qwen2_vla/model_load_utils.py 616ffbdc154ed2d8 |
unverified |
MIT (permissive) |
| Daily-Omni: Towards Audio-Visual Reasoning with Temporal Alignment across Modalities |
23 May 2025 |
identical code first harvested elsewhere 735025744c1ab0cf |
unverified |
licence of this copy not recorded |
| BLIP3-o: A Family of Fully Open Unified Multimodal Models-Architecture, Training and Dataset |
14 May 2025 |
identical code first harvested elsewhere 735025744c1ab0cf |
unverified |
licence of this copy not recorded |
| T2I-R1: Reinforcing Image Generation with Collaborative Semantic-level and Token-level CoT |
1 May 2025 |
identical code first harvested elsewhere 735025744c1ab0cf |
unverified |
licence of this copy not recorded |
| Med3DVLM: An Efficient Vision-Language Model for 3D Medical Image Analysis |
25 Mar 2025 |
mirthai/med3dvlm/src/train/train_vlm.py 616ffbdc154ed2d8 |
unverified |
MIT (permissive) |
| Safety Mirage: How Spurious Correlations Undermine VLM Safety Fine-tuning |
14 Mar 2025 |
OPTML-Group/VLM-Safety-Unlearn/llava/train/llava_unlearn_full_trainer.py 735025744c1ab0cf |
unverified |
MIT (permissive) |
| Safety Mirage: How Spurious Correlations Undermine VLM Safety Fine-tuning |
14 Mar 2025 |
OPTML-Group/VLM-Safety-Unlearn/llava/train/train_unlearn.py 616ffbdc154ed2d8 |
unverified |
MIT (permissive) |
| R1-Omni: Explainable Omni-Multimodal Emotion Recognition with Reinforcement Learning |
7 Mar 2025 |
identical code first harvested elsewhere 616ffbdc154ed2d8 |
unverified |
licence of this copy not recorded |
| Re-Imagining Multimodal Instruction Tuning: A Representation View |
2 Mar 2025 |
identical code first harvested elsewhere 616ffbdc154ed2d8 |
unverified |
licence of this copy not recorded |
| Magma: A Foundation Model for Multimodal AI Agents |
18 Feb 2025 |
microsoft/Magma/trainer/trainer.py 735025744c1ab0cf |
unverified |
MIT (permissive) |
| X-Boundary: Establishing Exact Safety Boundary to Shield LLMs from Multi-Turn Jailbreaks without Compromising Usability |
14 Feb 2025 |
ai45lab/x-boundary/src/lorra_x_boundary.py 16f29eca39ef4fff |
ran · our draft was wrong
fingerprinted |
no licence file found · pointer only |
| Ola: Pushing the Frontiers of Omni-Modal Language Model |
6 Feb 2025 |
ola-omni/ola/ola/utils.py 720e6dbf9d7eb553 |
unverified |
Apache-2.0 (permissive) |
| VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding |
22 Jan 2025 |
damo-nlp-sg/videollama3/videollama3/videollama3_trainer.py 616ffbdc154ed2d8 |
unverified |
Apache-2.0 (permissive) |
| Mitigating Hallucinations in Large Vision-Language Models via DPO: On-Policy Data Hold the Key |
16 Jan 2025 |
zhyang2226/opa-dpo/opadpo/opa_models/opa_trainer.py 735025744c1ab0cf |
unverified |
MIT recorded; this copy not marked cleared · pointer only |
| Mitigating Hallucinations in Large Vision-Language Models via DPO: On-Policy Data Hold the Key |
16 Jan 2025 |
zhyang2226/opa-dpo/opadpo/opa_train.py 616ffbdc154ed2d8 |
unverified |
MIT recorded; this copy not marked cleared · pointer only |
| PruneVid: Visual Token Pruning for Efficient Video Large Language Models |
20 Dec 2024 |
identical code first harvested elsewhere 735025744c1ab0cf |
unverified |
licence of this copy not recorded |
| MMedPO: Aligning Medical Vision-Language Models with Clinical-Aware Multimodal Preference Optimization |
9 Dec 2024 |
aiming-lab/mmedpo/train/dpo/llava_trainer_weighted.py 735025744c1ab0cf |
unverified |
Apache-2.0 (permissive) |
| Insight-V: Exploring Long-Chain Visual Reasoning with Multimodal Large Language Models |
21 Nov 2024 |
identical code first harvested elsewhere 735025744c1ab0cf |
unverified |
licence of this copy not recorded |
| PPLLaVA: Varied Video Sequence Understanding With Prompt Guidance |
4 Nov 2024 |
identical code first harvested elsewhere 735025744c1ab0cf |
unverified |
licence of this copy not recorded |
| Senna: Bridging Large Vision-Language Models and End-to-End Autonomous Driving |
29 Oct 2024 |
hustvl/senna/llava/senna/senna_llava_trainer.py 735025744c1ab0cf |
unverified |
Apache-2.0 (permissive) |
| Senna: Bridging Large Vision-Language Models and End-to-End Autonomous Driving |
29 Oct 2024 |
hustvl/senna/llava/senna/train_senna_llava_laion_pretrain.py 616ffbdc154ed2d8 |
unverified |
Apache-2.0 (permissive) |
| LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding |
22 Oct 2024 |
Vision-CAIR/LongVU/longvu/mm_datautils.py 25aaa1a93cc25f4d |
unverified |
Apache-2.0 (permissive) |
| Improve Vision Language Model Chain-of-thought Reasoning |
21 Oct 2024 |
riflezhang/llava-hound-dpo/llava_hound_dpo/dpo_scripts/run_dpo.py 616ffbdc154ed2d8 |
unverified |
no licence file found · pointer only |
| RAP: Retrieval-Augmented Personalization for Multimodal Large Language Models |
17 Oct 2024 |
hoar012/rap-mllm/llava/train/rap_train.py 616ffbdc154ed2d8 |
unverified |
no licence file found · pointer only |
| MMed-RAG: Versatile Multimodal RAG System for Medical Vision Language Models |
16 Oct 2024 |
richard-peng-xia/MMed-RAG/train/dpo/llava_trainer_2stages.py 735025744c1ab0cf |
unverified |
MIT (permissive) |
| Reconstructive Visual Instruction Tuning |
12 Oct 2024 |
haochen-wang409/ross/ross/ross_trainer.py 735025744c1ab0cf |
unverified |
Apache-2.0 (permissive) |
| TRACE: Temporal Grounding Video LLM via Causal Event Modeling |
8 Oct 2024 |
gyxxyg/TRACE/trace/trace_trainer.py 735025744c1ab0cf |
unverified |
Apache-2.0 (permissive) |
| TRACE: Temporal Grounding Video LLM via Causal Event Modeling |
8 Oct 2024 |
gyxxyg/trace/trace/train_mt.py 616ffbdc154ed2d8 |
unverified |
Apache-2.0 (permissive) |
| FakeShield: Explainable Image Forgery Detection and Localization via Multi-modal Large Language Models |
3 Oct 2024 |
identical code first harvested elsewhere 735025744c1ab0cf |
unverified |
licence of this copy not recorded |
| Ferret: Federated Full-Parameter Tuning at Scale for Large Language Models |
10 Sep 2024 |
identical code first harvested elsewhere 735025744c1ab0cf |
unverified |
licence of this copy not recorded |
| Recoverable Compression: A Multimodal Vision Token Recovery Mechanism Guided by Text Information |
2 Sep 2024 |
identical code first harvested elsewhere 735025744c1ab0cf |
unverified |
licence of this copy not recorded |
| UniFashion: A Unified Vision-Language Model for Multimodal Fashion Retrieval and Generation |
21 Aug 2024 |
xiangyu-mm/UniFashion/src/blip_fine_tune_2.py 7890d7ff378dc2fc |
unverified |
no licence file found · pointer only |
| LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models |
10 Jul 2024 |
identical code first harvested elsewhere 735025744c1ab0cf |
unverified |
licence of this copy not recorded |
| STLLaVA-Med: Self-Training Large Language and Vision Assistant for Medical Question-Answering |
28 Jun 2024 |
heliossun/stllava-med/llava_dpo_trainer.py 735025744c1ab0cf |
unverified |
MIT (permissive) |
| STLLaVA-Med: Self-Training Large Language and Vision Assistant for Medical Question-Answering |
28 Jun 2024 |
heliossun/stllava-med/train_dpo.py 616ffbdc154ed2d8 |
unverified |
MIT (permissive) |
| PRESTO: Progressive Pretraining Enhances Synthetic Chemistry Outcomes |
19 Jun 2024 |
idea-xl/presto/presto/model_utils.py 616ffbdc154ed2d8 |
unverified |
Apache-2.0 (permissive) |
| VoCo-LLaMA: Towards Vision Compression with Large Language Models |
18 Jun 2024 |
identical code first harvested elsewhere 735025744c1ab0cf |
unverified |
licence of this copy not recorded |
| mDPO: Conditional Preference Optimization for Multimodal Large Language Models |
17 Jun 2024 |
luka-group/mDPO/bunny/run_mdpo_bunny.py 616ffbdc154ed2d8 |
unverified |
no licence file found · pointer only |
| MMDU: A Multi-Turn Multi-Image Dialog Understanding Benchmark and Instruction-Tuning Dataset for LVLMs |
17 Jun 2024 |
identical code first harvested elsewhere 735025744c1ab0cf |
unverified |
licence of this copy not recorded |
| Unveiling Encoder-Free Vision-Language Models |
17 Jun 2024 |
identical code first harvested elsewhere 735025744c1ab0cf |
unverified |
licence of this copy not recorded |
| ISR-DPO: Aligning Large Multimodal Models for Videos by Iterative Self-Retrospective DPO |
17 Jun 2024 |
yonseivnl/vlm-rlaif/RLAIF/finetune_policy_init.py 616ffbdc154ed2d8 |
unverified |
Apache-2.0 (permissive) |
| Vript: A Video Is Worth Thousands of Words |
10 Jun 2024 |
mutonix/Vript/vriptor/train_hf.py 616ffbdc154ed2d8 |
unverified |
no licence file found · pointer only |
| Parrot: Multilingual Visual Instruction Tuning |
4 Jun 2024 |
AIDC-AI/Parrot/parrot/train/parrot_trainer.py 735025744c1ab0cf |
unverified |
Apache-2.0 (permissive) |
| Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis |
31 May 2024 |
PhysGame/PhysGame/train_dpo.py 616ffbdc154ed2d8 |
unverified |
Apache-2.0 (permissive) |
| Data-augmented phrase-level alignment for mitigating object hallucination |
28 May 2024 |
pritamqu/HALVA/llava/train/halva_trainer.py 735025744c1ab0cf |
unverified |
no licence file found · pointer only |
| Data-augmented phrase-level alignment for mitigating object hallucination |
28 May 2024 |
pritamqu/HALVA/llava/train/train_halva.py 616ffbdc154ed2d8 |
unverified |
no licence file found · pointer only |
| VoCoT: Unleashing Visually Grounded Multi-Step Reasoning in Large Multi-Modal Models |
27 May 2024 |
rupertluo/vocot/train_volcano.py d10512f432705d71 |
unverified |
no licence file found · pointer only |
| PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning |
25 Apr 2024 |
magic-research/PLLaVA/tasks/train/train_pllava_nframe_accel.py 735025744c1ab0cf |
unverified |
no licence file found · pointer only |
| Detecting and Mitigating Hallucination in Large Vision Language Models via Fine-Grained AI Feedback |
22 Apr 2024 |
Mr-Loevan/HSA-DPO/hsa_dpo/models/llava-v1_5/train_dpo.py 616ffbdc154ed2d8 |
unverified |
no licence file found · pointer only |
| MoVA: Adapting Mixture of Vision Experts to Multimodal Context |
19 Apr 2024 |
identical code first harvested elsewhere 735025744c1ab0cf |
unverified |
licence of this copy not recorded |
| Self-Supervised Visual Preference Alignment |
16 Apr 2024 |
Kevinz-code/SeVa/seva/train_dpo_ours.py 616ffbdc154ed2d8 |
unverified |
GPL-3.0 (copyleft) · pointer only |
| ST-LLM: Large Language Models Are Effective Temporal Learners |
30 Mar 2024 |
identical code first harvested elsewhere 735025744c1ab0cf |
unverified |
licence of this copy not recorded |
| SQ-LLaVA: Self-Questioning for Large Vision-Language Assistant |
17 Mar 2024 |
heliossun/sq-llava/llava_trainer.py 735025744c1ab0cf |
unverified |
MIT (permissive) |
| Images are Achilles' Heel of Alignment: Exploiting Visual Vulnerabilities for Jailbreaking Multimodal Large Language Models |
14 Mar 2024 |
identical code first harvested elsewhere 735025744c1ab0cf |
unverified |
licence of this copy not recorded |
| CAT: Enhancing Multimodal Large Language Model to Answer Questions in Dynamic Audio-Visual Scenarios |
7 Mar 2024 |
rikeilong/bay-cat/ADPO_CAT/train_dpo.py 616ffbdc154ed2d8 |
unverified |
Apache-2.0 (permissive) |
| Grounding Language Models for Visual Entity Recognition |
28 Feb 2024 |
mrzilinxiao/autover/train_oven.py 616ffbdc154ed2d8 |
unverified |
Apache-2.0 (permissive) |
| Your Vision-Language Model Itself Is a Strong Filter: Towards High-Quality Instruction Tuning with Data Selection |
19 Feb 2024 |
rayruibochen/self-filter/self_filter/self_filter_trainer.py 735025744c1ab0cf |
unverified |
AGPL-3.0 (copyleft) · pointer only |
| Your Vision-Language Model Itself Is a Strong Filter: Towards High-Quality Instruction Tuning with Data Selection |
19 Feb 2024 |
rayruibochen/self-filter/self_filter/stage1.py 616ffbdc154ed2d8 |
unverified |
AGPL-3.0 (copyleft) · pointer only |
| DoRA: Weight-Decomposed Low-Rank Adaptation |
14 Feb 2024 |
NVlabs/DoRA/visual_instruction_tuning/llava/train/train_dora.py 616ffbdc154ed2d8 |
unverified |
no licence file found · pointer only |
| LLaGA: Large Language and Graph Assistant |
13 Feb 2024 |
chenrunjin/llaga/train/llaga_trainer.py 735025744c1ab0cf |
unverified |
Apache-2.0 (permissive) |
| LLaVA-Phi: Efficient Multi-Modal Assistant with Small Language Model |
4 Jan 2024 |
zhuyiche/llava-phi/llava_phi/train/convert_model2base_llava_phi.py 616ffbdc154ed2d8 |
unverified |
no licence file found · pointer only |
| G-LLaVA: Solving Geometric Problem with Multi-Modal Large Language Model |
18 Dec 2023 |
pipilurj/g-llava/gllava/train/llava_trainer.py 735025744c1ab0cf |
unverified |
no licence file found · pointer only |
| Osprey: Pixel Understanding with Visual Instruction Tuning |
15 Dec 2023 |
identical code first harvested elsewhere 735025744c1ab0cf |
unverified |
licence of this copy not recorded |
| Hallucination Augmented Contrastive Learning for Multimodal Large Language Model |
12 Dec 2023 |
identical code first harvested elsewhere 735025744c1ab0cf |
unverified |
licence of this copy not recorded |
| LLaVA-Grounding: Grounded Visual Chat with Large Multimodal Models |
5 Dec 2023 |
ux-decoder/llava-grounding/llava/train/train_grounding_1st.py 616ffbdc154ed2d8 |
unverified |
Apache-2.0 (permissive) |
| Beyond Hallucinations: Enhancing LVLMs through Hallucination-Aware Direct Preference Optimization |
28 Nov 2023 |
identical code first harvested elsewhere 616ffbdc154ed2d8 |
unverified |
licence of this copy not recorded |
| GeoChat: Grounded Large Vision-Language Model for Remote Sensing |
24 Nov 2023 |
identical code first harvested elsewhere 735025744c1ab0cf |
unverified |
licence of this copy not recorded |
| mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration |
7 Nov 2023 |
identical code first harvested elsewhere 735025744c1ab0cf |
unverified |
licence of this copy not recorded |
| Improved Baselines with Visual Instruction Tuning |
5 Oct 2023 |
identical code first harvested elsewhere 735025744c1ab0cf |
unverified |
licence of this copy not recorded |
| HallE-Control: Controlling Object Hallucination in Large Multimodal Models |
3 Oct 2023 |
bronyayang/HallE_Switch/llava/train/train_switch.py 616ffbdc154ed2d8 |
unverified |
no licence file found · pointer only |
| Visual Instruction Tuning |
17 Apr 2023 |
haotian-liu/LLaVA/llava/train/llava_trainer.py 735025744c1ab0cf |
unverified |
Apache-2.0 (permissive) |
| Unicom: Universal and Compact Representation Learning for Image Retrieval |
12 Apr 2023 |
identical code first harvested elsewhere 735025744c1ab0cf |
unverified |
licence of this copy not recorded |
| arXiv:Wang_SMoLoRA_Exploring_and_Defying_Dual_Catastrophic_Forgetting_in_Continual_Visual_ICCV_2025_paper |
|
Minato-Zackie/SMoLoRA/llava/train/train_SMoLoRA.py 616ffbdc154ed2d8 |
unverified |
Apache-2.0 (permissive) |