| Beyond Language Priors: Diagnosing and Fixing Visual-Origin Hallucinations in Multimodal LLM added by Syntology |
2026-09 (from id) |
zxp555/ACFT_MM26/ACFT/llava/mm_utils.py 3999ff487573f32c |
ran · fixture could not drive it
|
no licence file found · pointer only |
| MLLM-DataEngine: Closing the Loop of Multimodal Instruction Tuning Data Generation added by Syntology |
2026-07 (from id) |
opendatalab/MLLM-DataEngine/LLaVA/llava/mm_utils.py 3999ff487573f32c |
ran · fixture could not drive it
|
Apache-2.0 (permissive) |
| Do Text Edits Generalize to Visual Generation? Benchmarking Cross-Modal Knowledge Editing in UMMs added by Syntology |
2026-06 (from id) |
gxx27/UniKE/BLIP3o/mm_utils.py 3999ff487573f32c |
ran · fixture could not drive it
|
MIT (permissive) |
| HiDrop: Hierarchical Vision Token Reduction in MLLMs via Late Injection, Concave Pyramid Pruning, and Early Exit added by Syntology |
2026-02 (from id) |
EIT-NLP/HiDrop/llava/mm_utils.py 3999ff487573f32c |
ran · fixture could not drive it
|
Apache-2.0 (permissive) |
| Natural Language Instructions for Scene-Responsive Human-in-the-Loop Motion Planning in Autonomous Driving using Vision-Language-Action Models added by Syntology |
2026-02 (from id) |
Mi3-Lab/doScenes-VLM-Planning/src/llava/mm_utils.py 3999ff487573f32c |
ran · fixture could not drive it
|
AGPL-3.0 (copyleft) · pointer only |
| Natural Language Instructions for Scene-Responsive Human-in-the-Loop Motion Planning in Autonomous Driving using Vision-Language-Action Models added by Syntology |
2026-02 (from id) |
Mi3-Lab/doScenes-VLM-Planning/src/llava/mm_utils_differentiable.py 3c876f5095840dc7 |
unverified |
AGPL-3.0 (copyleft) · pointer only |
| Mitigating Object Hallucinations via Sentence-Level Early Intervention |
16 Jul 2025 |
pspdada/SENTINEL/llava/mm_utils.py 3999ff487573f32c |
ran · fixture could not drive it
|
Apache-2.0 (permissive) |
| LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs |
1 Jul 2025 |
CnFaker/LLaVA-SP/llava/mm_utils.py 3999ff487573f32c |
ran · fixture could not drive it
|
Apache-2.0 (permissive) |
| UniFork: Exploring Modality Alignment for Unified Multimodal Understanding and Generation |
20 Jun 2025 |
tliby/unifork/unifork/mm_utils.py 3999ff487573f32c |
ran · fixture could not drive it
|
no licence file found · pointer only |
| Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs |
12 Jun 2025 |
theia-4869/cdpruner/llava/mm_utils.py 3ee0f92602576a06 |
unverified |
Apache-2.0 (permissive) |
| ILLUME+: Illuminating Unified MLLM with Dual Visual Tokenization and Diffusion Refinement |
2 Apr 2025 |
illume-unified-mllm/ILLUME_plus/ILLUME/illume/mm_utils.py 29d67c6717548a53 |
unverified |
Apache-2.0 (permissive) |
| arXiv:2504.00502 |
2025-04 (from id) |
icip-cas/ShortV/llava/mm_utils.py 3999ff487573f32c |
ran · fixture could not drive it
|
Apache-2.0 (permissive) |
| GraspCoT: Integrating Physical Property Reasoning for 6-DoF Grasping under Flexible Language Instructions |
20 Mar 2025 |
cxmomo/GraspCoT/llava/mm_utils.py 3999ff487573f32c |
ran · fixture could not drive it
|
MIT (permissive) |
| LLaVA-MORE: A Comparative Study of LLMs and Visual Backbones for Enhanced Visual Instruction Tuning |
19 Mar 2025 |
aimagelab/LLaVA-MORE/src/llava/mm_utils.py 3999ff487573f32c |
ran · fixture could not drive it
|
Apache-2.0 (permissive) |
| Safety Mirage: How Spurious Correlations Undermine VLM Safety Fine-tuning |
14 Mar 2025 |
optml-group/vlm-safety-mu/llava/mm_utils.py 3999ff487573f32c |
ran · fixture could not drive it
|
MIT (permissive) |
| FastVID: Dynamic Density Pruning for Fast Video Large Language Models |
14 Mar 2025 |
cokeshao/holitom/holitom/llava_arch.py 3999ff487573f32c |
ran · fixture could not drive it
|
Apache-2.0 (permissive) |
| Magma: A Foundation Model for Multimodal AI Agents |
18 Feb 2025 |
microsoft/Magma/magma/image_processing_magma.py 3999ff487573f32c |
ran · fixture could not drive it
|
MIT (permissive) |
| DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding |
13 Dec 2024 |
deepseek-ai/deepseek-vl2/deepseek_vl2/models/processing_deepseek_vl_v2.py 6408ebfc6065bf19 |
unverified |
MIT (permissive) |
| [CLS] Token Tells Everything Needed for Training-free Efficient MLLMs |
8 Dec 2024 |
thu-mig/vtc-cls/llava/mm_utils.py 3999ff487573f32c |
ran · fixture could not drive it
|
Apache-2.0 (permissive) |
| TeamCraft: A Benchmark for Multi-Modal Multi-Agent Systems in Minecraft |
6 Dec 2024 |
teamcraft-bench/teamcraft/llava_teamcraft/llava/mm_utils.py 3999ff487573f32c |
ran · fixture could not drive it
|
MIT (permissive) |
| p-MoD: Building Mixture-of-Depths MLLMs via Progressive Ratio Decay |
5 Dec 2024 |
mcg-nju/p-mod/llava/mm_utils.py 3999ff487573f32c |
ran · fixture could not drive it
|
Apache-2.0 (permissive) |
| Agri-LLaVA: Knowledge-Infused Large Multimodal Assistant on Agricultural Pests and Diseases |
3 Dec 2024 |
kki2eve/agri-llava/agri_llava/mm_utils.py 3999ff487573f32c |
ran · fixture could not drive it
|
Apache-2.0 (permissive) |
| Beyond Text-Visual Attention: Exploiting Visual Cues for Effective Token Pruning in VLMs |
2 Dec 2024 |
theia-4869/fastervlm/llava/mm_utils.py 30113c28bc9b982c |
unverified |
Apache-2.0 (permissive) |
| Beyond Text-Visual Attention: Exploiting Visual Cues for Effective Token Pruning in VLMs |
2 Dec 2024 |
theia-4869/vispruner/llava/mm_utils.py 3ee0f92602576a06 |
unverified |
Apache-2.0 (permissive) |
| Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering |
25 Nov 2024 |
aimagelab/reflectiva/llava/mm_utils.py 3999ff487573f32c |
ran · fixture could not drive it
|
Apache-2.0 (permissive) |
| TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models |
17 Nov 2024 |
tingyu215/ts-llava/llava/mm_utils.py 3999ff487573f32c |
ran · fixture could not drive it
|
Apache-2.0 (permissive) |
| Multi-Stage Vision Token Dropping: Towards Efficient Multimodal Large Language Model |
16 Nov 2024 |
liuting20/mustdrop/llava/mm_utils.py 3999ff487573f32c |
ran · fixture could not drive it
|
no licence file found · pointer only |
| LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding |
22 Oct 2024 |
Vision-CAIR/LongVU/longvu/mm_utils.py 3999ff487573f32c |
ran · fixture could not drive it
|
Apache-2.0 (permissive) |
| PyramidDrop: Accelerating Your Large Vision-Language Models via Pyramid Visual Redundancy Reduction |
22 Oct 2024 |
cooperx521/pyramiddrop/llava/mm_utils.py 3999ff487573f32c |
ran · fixture could not drive it
|
MIT (permissive) |
| Improve Vision Language Model Chain-of-thought Reasoning |
21 Oct 2024 |
riflezhang/llava-reasoner-dpo/llava_reasoner/llava/mm_utils.py 3999ff487573f32c |
ran · fixture could not drive it
|
no licence file found · pointer only |
| RAP: Retrieval-Augmented Personalization for Multimodal Large Language Models |
17 Oct 2024 |
hoar012/rap-mllm/llava/mm_utils.py 3999ff487573f32c |
ran · fixture could not drive it
|
no licence file found · pointer only |
| Reconstructive Visual Instruction Tuning |
12 Oct 2024 |
haochen-wang409/ross/ross/mm_utils.py 3999ff487573f32c |
ran · fixture could not drive it
|
Apache-2.0 (permissive) |
| Q-VLM: Post-training Quantization for Large Vision-Language Models |
10 Oct 2024 |
changyuanwang17/qvlm/llava/mm_utils.py 3999ff487573f32c |
ran · fixture could not drive it
|
Apache-2.0 (permissive) |
| Deciphering Cross-Modal Alignment in Large Vision-Language Models with Modality Integration Rate |
9 Oct 2024 |
shikiw/modality-integration-rate/llava/mm_utils.py 3999ff487573f32c |
ran · fixture could not drive it
|
MIT (permissive) |
| Personalized Visual Instruction Tuning |
9 Oct 2024 |
sterzhang/pvit/personalize-llava/llava/mm_utils.py 3999ff487573f32c |
ran · fixture could not drive it
|
no licence file found · pointer only |
| SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference |
6 Oct 2024 |
Gumpest/SparseVLMs/llava/mm_utils.py 3999ff487573f32c |
ran · fixture could not drive it
|
Apache-2.0 (permissive) |
| EMMA: Efficient Visual Alignment in Multi-Modal LLMs |
2 Oct 2024 |
saraghazanfari/emma/llava/mm_utils.py 3999ff487573f32c |
ran · fixture could not drive it
|
Apache-2.0 (permissive) |
| AVG-LLaVA: A Large Multimodal Model with Adaptive Visual Granularity |
20 Sep 2024 |
deeplearnxmu/avg-llava/llava/mm_utils.py 3999ff487573f32c |
ran · fixture could not drive it
|
Apache-2.0 (permissive) |
| One Token to Seg Them All: Language Instructed Reasoning Segmentation in Videos |
29 Sep 2024 |
showlab/videolisa/model/llava/mm_utils.py 3999ff487573f32c |
ran · fixture could not drive it
|
Apache-2.0 (permissive) |
| Explanation Bottleneck Models |
26 Sep 2024 |
yshinya6/xbm/xbm-llava/llava/mm_utils.py 3999ff487573f32c |
ran · fixture could not drive it
|
no licence file found · pointer only |
| SURf: Teaching Large Vision-Language Models to Selectively Utilize Retrieved Information |
21 Sep 2024 |
GasolSun36/SURf/llava/mm_utils.py 3999ff487573f32c |
ran · fixture could not drive it
|
no licence file found · pointer only |
| Less is More: A Simple yet Effective Token Reduction Method for Efficient Multi-modal LLMs |
17 Sep 2024 |
freedomintelligence/trim/llava/mm_utils.py 3999ff487573f32c |
ran · fixture could not drive it
|
Apache-2.0 (permissive) |
| MotIF: Motion Instruction Fine-tuning |
16 Sep 2024 |
Minyoung1005/motif/LLaVA/llava/mm_utils.py 3999ff487573f32c |
ran · fixture could not drive it
|
no licence file found · pointer only |
| LongLLaVA: Scaling Multi-modal LLMs to 1000 Images Efficiently via a Hybrid Architecture |
4 Sep 2024 |
freedomintelligence/longllava/llava/mm_utils.py 3999ff487573f32c |
ran · fixture could not drive it
|
no licence file found · pointer only |
| Advancing Multimodal Large Language Models with Quantization-Aware Scale Learning for Efficient Adaptation |
7 Aug 2024 |
xjjxmu/qslaw/llava/mm_utils.py 3999ff487573f32c |
ran · fixture could not drive it
|
Apache-2.0 (permissive) |
| MedTrinity-25M: A Large-scale Multimodal Dataset with Multigranular Annotations for Medicine |
6 Aug 2024 |
UCSC-VLAA/MedTrinity-25M/llava/mm_utils.py 3999ff487573f32c |
ran · fixture could not drive it
|
no licence file found · pointer only |
| GalleryGPT: Analyzing Paintings with Large Multimodal Models |
1 Aug 2024 |
steven640pixel/gallerygpt/llava/mm_utils.py 3999ff487573f32c |
ran · fixture could not drive it
|
Apache-2.0 (permissive) |
| SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models |
22 Jul 2024 |
apple/ml-slowfast-llava/slowfast_llava/llava/mm_utils.py 3999ff487573f32c |
ran · fixture could not drive it
|
no licence file found · pointer only |
| Efficient Large Multi-modal Models via Visual Context Compression |
28 Jun 2024 |
Beckschen/LLaVolta/llava/mm_utils.py 3999ff487573f32c |
ran · fixture could not drive it
|
Apache-2.0 (permissive) |
| EVALALIGN: Supervised Fine-Tuning Multimodal LLMs with Human-Aligned Data for Evaluating Text-to-Image Models |
24 Jun 2024 |
sais-fuxi/evalalign/evalalign/mm_utils.py 3999ff487573f32c |
ran · fixture could not drive it
|
Apache-2.0 (permissive) |
| VoCo-LLaMA: Towards Vision Compression with Large Language Models |
18 Jun 2024 |
Yxxxb/VoCo-LLaMA/llava/mm_utils.py 3999ff487573f32c |
ran · fixture could not drive it
|
Apache-2.0 (permissive) |
| On Efficient Language and Vision Assistants for Visually-Situated Natural Language Understanding: What Matters in Reading and Reasoning |
17 Jun 2024 |
naver-ai/elva/Elva/mm_utils.py 3999ff487573f32c |
ran · fixture could not drive it
|
no licence file found · pointer only |
| ClawMachine: Learning to Fetch Visual Tokens for Referential Comprehension |
17 Jun 2024 |
martian422/ClawMachine/ClawMachine/mm_utils.py 3999ff487573f32c |
ran · fixture could not drive it
|
no licence file found · pointer only |
| Concept-skill Transferability-based Data Selection for Large Vision-Language Models |
16 Jun 2024 |
g-jwlee/coincide_code/COINCIDE_cluster/tinyllava/mm_utils.py 3999ff487573f32c |
ran · fixture could not drive it
|
no licence file found · pointer only |
| Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models |
12 Jun 2024 |
yfzhang114/slime/llava/mm_utils.py 3999ff487573f32c |
ran · fixture could not drive it
|
Apache-2.0 (permissive) |
| Matryoshka Query Transformer for Large Vision-Language Models |
29 May 2024 |
gordonhu608/mqt-llava/llava/mm_utils.py 3999ff487573f32c |
ran · fixture could not drive it
|
Apache-2.0 (permissive) |
| ConvLLaVA: Hierarchical Backbones as Visual Encoder for Large Multimodal Models |
24 May 2024 |
alibaba/conv-llava/llava/mm_utils.py 3999ff487573f32c |
ran · fixture could not drive it
|
Apache-2.0 (permissive) |
| FreeVA: Offline MLLM as Training-Free Video Assistant |
13 May 2024 |
whwu95/freeva/llava/mm_utils.py 3999ff487573f32c |
ran · fixture could not drive it
|
Apache-2.0 (permissive) |
| CuMo: Scaling Multimodal LLM with Co-Upcycled Mixture-of-Experts |
9 May 2024 |
shi-labs/cumo/cumo/mm_utils.py 3999ff487573f32c |
ran · fixture could not drive it
|
Apache-2.0 (permissive) |
| Boosting Multimodal Large Language Models with Visual Tokens Withdrawal for Rapid Inference |
9 May 2024 |
lzhxmu/vtw/llava/mm_utils.py 3999ff487573f32c |
ran · fixture could not drive it
|
no licence file found · pointer only |
| SEED-Data-Edit Technical Report: A Hybrid Dataset for Instructional Image Editing |
7 May 2024 |
ailab-cvc/seed-x/src/inference/any_res.py 3999ff487573f32c |
ran · fixture could not drive it
|
no licence file found · pointer only |
| MoVA: Adapting Mixture of Vision Experts to Multimodal Context |
19 Apr 2024 |
templex98/mova/mova/mm_utils.py 3999ff487573f32c |
ran · fixture could not drive it
|
Apache-2.0 (permissive) |
| Less is More: Mitigating Multimodal Hallucination from an EOS Decision Perspective |
22 Feb 2024 |
yuezih/less-is-more/LLaVA/llava/mm_utils.py 3999ff487573f32c |
ran · fixture could not drive it
|
Apache-2.0 (permissive) |
| Aligning Modalities in Vision Large Language Models via Preference Fine-tuning |
18 Feb 2024 |
yiyangzhou/povid/llava/mm_utils.py 3999ff487573f32c |
ran · fixture could not drive it
|
Apache-2.0 (permissive) |
| Unsupervised Universal Image Segmentation |
28 Dec 2023 |
dantong88/llarva/llava/mm_utils.py 3999ff487573f32c |
ran · fixture could not drive it
|
Apache-2.0 (permissive) |
| Volcano: Mitigating Multimodal Hallucination through Self-Feedback Guided Revision |
13 Nov 2023 |
kaistai/volcano/llava/mm_utils.py 3999ff487573f32c |
ran · fixture could not drive it
|
no licence file found · pointer only |
| Octopus: Embodied Vision-Language Programmer from Environmental Feedback |
12 Oct 2023 |
dongyh20/octopus/octopus/LLaVA/llava/mm_utils.py 3999ff487573f32c |
ran · fixture could not drive it
|
no licence file found · pointer only |
| MLLM-DataEngine: An Iterative Refinement Approach for MLLM |
25 Aug 2023 |
opendatalab/mllm-dataengine/LLaVA/llava/mm_utils.py 3999ff487573f32c |
ran · fixture could not drive it
|
Apache-2.0 (permissive) |
| GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio |
13 Jun 2021 |
maikezuefle/contr-pretraining/llava/mm_utils.py 3999ff487573f32c |
ran · fixture could not drive it
|
Apache-2.0 (permissive) |
| arXiv:Zhang_Beyond_Training_Dynamic_Token_Merging_for_Zero-Shot_Video_Understanding_ICCV_2025_paper |
|
Jam1ezhang/DYTO/dyto/llava/mm_utils.py 3999ff487573f32c |
ran · fixture could not drive it
|
Apache-2.0 (permissive) |
| arXiv:Xing_Conical_Visual_Concentration_for_Efficient_Large_Vision-Language_Models_CVPR_2025_paper |
|
Cooperx521/PyramidDrop/llava/mm_utils.py 3999ff487573f32c |
ran · fixture could not drive it
|
MIT (permissive) |
| arXiv:2025.findings-acl.865 |
|
DeepLearnXMU/AVG-LLaVA/llava/mm_utils.py 3999ff487573f32c |
ran · fixture could not drive it
|
Apache-2.0 (permissive) |
| arXiv:2025.findings-acl.458 |
|
DCDmllm/Align2LLaVA/reward_model/llava/mm_utils.py 3999ff487573f32c |
ran · fixture could not drive it
|
Apache-2.0 (permissive) |