| Beyond Language Priors: Diagnosing and Fixing Visual-Origin Hallucinations in Multimodal LLM added by Syntology |
2026-09 (from id) |
zxp555/ACFT_MM26/ACFT/llava/eval/model_vqa.py 076c252c52cbb161 |
ran · our draft was wrong
fingerprinted |
no licence file found · pointer only |
| DIVE: Dynamic Iterative Visual Evidence Construction for Efficient Vision-Language Models added by Syntology |
2026-08 (from id) |
Zhong-Chenchen/DIVE/llava/eval/model_vqa.py 076c252c52cbb161 |
ran · our draft was wrong
fingerprinted |
Apache-2.0 (permissive) |
| DIVE: Dynamic Iterative Visual Evidence Construction for Efficient Vision-Language Models added by Syntology |
2026-08 (from id) |
Zhong-Chenchen/DIVE/llava/eval/model_vqa_ocrbench.py cb9eaee96a223830 |
ran
fingerprinted |
Apache-2.0 (permissive) |
| MLLM-DataEngine: Closing the Loop of Multimodal Instruction Tuning Data Generation added by Syntology |
2026-07 (from id) |
opendatalab/MLLM-DataEngine/LLaVA/llava/eval/model_vqa.py 076c252c52cbb161 |
ran · our draft was wrong
fingerprinted |
Apache-2.0 (permissive) |
| OralAgent: Integrating Reasoning, Tools, and Knowledge for Interactive Dental Image Analysis added by Syntology |
2026-05 (from id) |
isjinghao/OralAgent/oralagent/llava/eval/model_vqa.py 076c252c52cbb161 |
ran · our draft was wrong
fingerprinted |
Apache-2.0 (permissive) |
| Reasoning Portability: Guiding Continual Learning for MLLMs in the RLVR Era added by Syntology |
2026-05 (from id) |
lluosi/RDB-CL/ETrain/Eval/eval_GenealKnowledge.py c81af3c995ec5f75 |
ran
fingerprinted |
no licence file found · pointer only |
| WSVD: Weighted Low-Rank Approximation for Fast and Efficient Execution of Low-Precision Vision-Language Models added by Syntology |
2026-04 (from id) |
SAI-Lab-NYU/WSVD/e2e/infer_llava.py 076c252c52cbb161 |
ran · our draft was wrong
fingerprinted |
Apache-2.0 (permissive) |
| COVTrack++: Learning Open-Vocabulary Multi-Object Tracking from Continuous Videos via a Synergistic Paradigm added by Syntology |
2026-03 (from id) |
zekunqian/covtrack/run_later.py cd7eff5270a7a400 |
unverified |
Apache-2.0 (permissive) |
| Uncertainty-Aware Knowledge Distillation for Multimodal Large Language Models added by Syntology |
2026-03 (from id) |
Jingchensun/beta-kd/mobilevlm/eval/model_vqa_loader.py 076c252c52cbb161 |
ran · our draft was wrong
fingerprinted |
no licence file found · pointer only |
| HiDrop: Hierarchical Vision Token Reduction in MLLMs via Late Injection, Concave Pyramid Pruning, and Early Exit added by Syntology |
2026-02 (from id) |
EIT-NLP/HiDrop/llava/eval/model_vqa.py 076c252c52cbb161 |
ran · our draft was wrong
fingerprinted |
Apache-2.0 (permissive) |
| Beyond Next-Token Alignment: Distilling Multimodal Large Language Models via Token Interactions added by Syntology |
2026-02 (from id) |
lchen1019/Align-TI/alignti/eval/model_vqa.py 076c252c52cbb161 |
ran · our draft was wrong
fingerprinted |
Apache-2.0 (permissive) |
| Video-based Music Generation added by Syntology |
2026-02 (from id) |
serkansulun/emsync/utils.py 4abf480758895f1e |
unverified |
licence not identified · pointer only |
| Benchmarking Direct Preference Optimization for Medical Large Vision-Language Models added by Syntology |
2026-01 (from id) |
dmis-lab/med-vlm-dpo/inference/LLaVA-Med/llava/eval/model_vqa.py 076c252c52cbb161 |
ran · our draft was wrong
fingerprinted |
no licence file found · pointer only |
| DeepMoLM: Leveraging Visual and Geometric Structural Information for Molecule-Text Modeling added by Syntology |
2026-01 (from id) |
1anj/DeepMoLM/llava/eval/model_vqa_video.py 076c252c52cbb161 |
ran · our draft was wrong
fingerprinted |
Apache-2.0 (permissive) |
| Cross-Layer Injection for Deep Vision-Language Fusion added by Syntology |
2026-01 (from id) |
codefuse-ai/CLI/llava/eval/model_vqa.py 076c252c52cbb161 |
ran · our draft was wrong
fingerprinted |
Apache-2.0 (permissive) |
| QSVD: Efficient Low-rank Approximation for Unified Query-Key-Value Weight Compression in Low-Precision Vision-Language Models added by Syntology |
2025-10 (from id) |
SAI-Lab-NYU/QSVD/fake_quant/eval_llavanext_vizwiz.py 076c252c52cbb161 |
ran · our draft was wrong
fingerprinted |
Apache-2.0 (permissive) |
| Spatial Preference Rewarding for MLLMs Spatial Understanding added by Syntology |
2025-10 (from id) |
hanqiu-hq/SPR/construct_data/ferret_score_siglip.py 076c252c52cbb161 |
ran · our draft was wrong
fingerprinted |
no licence file found · pointer only |
| arXiv:2507.20163 |
2025-07 (from id) |
Zeyu1226-mt/LLM-IAVC/D_select_player_more_than5.py cc33a1ae9ef8204d |
unverified |
no licence file found · pointer only |
| arXiv:2507.18300 |
2025-07 (from id) |
360CVGroup/LMM-Det/llava/eval/model_coco_owlv2.py 076c252c52cbb161 |
ran · our draft was wrong
fingerprinted |
Apache-2.0 (permissive) |
| Mitigating Object Hallucinations via Sentence-Level Early Intervention |
16 Jul 2025 |
pspdada/SENTINEL/llava/utils.py d5904889c6fd4eb0 |
unverified |
Apache-2.0 (permissive) |
| LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs |
1 Jul 2025 |
CnFaker/LLaVA-SP/llava/eval/model_vqa.py 076c252c52cbb161 |
ran · our draft was wrong
fingerprinted |
Apache-2.0 (permissive) |
| UniFork: Exploring Modality Alignment for Unified Multimodal Understanding and Generation |
20 Jun 2025 |
tliby/unifork/unifork/eval/model_vqa.py 076c252c52cbb161 |
ran · our draft was wrong
fingerprinted |
no licence file found · pointer only |
| Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs |
12 Jun 2025 |
theia-4869/cdpruner/llava/eval/model_vqa.py 076c252c52cbb161 |
ran · our draft was wrong
fingerprinted |
Apache-2.0 (permissive) |
| Learning from Videos for 3D World: Enhancing MLLMs with 3D Vision Geometry Priors |
30 May 2025 |
LaVi-Lab/Video-3D-LLM/llava/eval/model_vqa.py 076c252c52cbb161 |
ran · our draft was wrong
fingerprinted |
Apache-2.0 (permissive) |
| VideoReasonBench: Can MLLMs Perform Vision-Centric Complex Video Reasoning? |
29 May 2025 |
identical code first harvested elsewhere 076c252c52cbb161 |
ran · our draft was wrong
fingerprinted |
licence of this copy not recorded |
| Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models |
27 May 2025 |
jefferyzhan/griffon/griffon/eval/model_vqa.py 076c252c52cbb161 |
ran · our draft was wrong
fingerprinted |
Apache-2.0 (permissive) |
| LaViDa: A Large Diffusion Language Model for Multimodal Understanding |
22 May 2025 |
jacklishufan/lavida/llava/eval/model_vqa.py 076c252c52cbb161 |
ran · our draft was wrong
fingerprinted |
Apache-2.0 (permissive) |
| R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning |
5 May 2025 |
yfzhang114/r1_reward/inference/MM-RLHF-Reward/r1_reward.py 076c252c52cbb161 |
ran · our draft was wrong
fingerprinted |
Apache-2.0 (permissive) |
| Unlearning Sensitive Information in Multimodal LLMs: Benchmark and Attack-Defense Evaluation |
1 May 2025 |
vaidehi99/unlok-vqa/LLaVA/llava/eval/model_vqa.py 076c252c52cbb161 |
ran · our draft was wrong
fingerprinted |
MIT (permissive) |
| Multi-Resolution Pathology-Language Pre-training Model with Text-Guided Visual Representation |
26 Apr 2025 |
BasitAlawode/MR-PLIP/generate_text.py 076c252c52cbb161 |
ran · our draft was wrong
fingerprinted |
MIT (permissive) |
| Benchmarking Large Vision-Language Models on Fine-Grained Image Tasks: A Comprehensive Evaluation |
21 Apr 2025 |
SEU-VIPGroup/FG-BMK/demo/human_evaluation/human_evaluation_demo.py 7ae822eca6f8dafe |
ran · our draft was wrong
fingerprinted |
no licence file found · pointer only |
| SoTA with Less: MCTS-Guided Sample Selection for Data-Efficient Visual Reasoning Self-Improvement |
10 Apr 2025 |
identical code first harvested elsewhere 076c252c52cbb161 |
ran · our draft was wrong
fingerprinted |
licence of this copy not recorded |
| arXiv:2504.00502 |
2025-04 (from id) |
icip-cas/ShortV/llava/eval/model_vqa.py 076c252c52cbb161 |
ran · our draft was wrong
fingerprinted |
Apache-2.0 (permissive) |
| MaTVLM: Hybrid Mamba-Transformer for Efficient Vision-Language Modeling |
17 Mar 2025 |
hustvl/MaTVLM/tinyllava/eval/model_vqa.py 076c252c52cbb161 |
ran · our draft was wrong
fingerprinted |
Apache-2.0 (permissive) |
| Safety Mirage: How Spurious Correlations Undermine VLM Safety Fine-tuning |
14 Mar 2025 |
OPTML-Group/VLM-Safety-Unlearn/llava/eval/model_vqa.py 076c252c52cbb161 |
ran · our draft was wrong
fingerprinted |
MIT (permissive) |
| Keyframe-oriented Vision Token Pruning: Enhancing Efficiency of Large Vision Language Models on Long-Form Video Processing |
13 Mar 2025 |
1999Lyd/KVTP/llava/eval/model_vqa.py 076c252c52cbb161 |
ran · our draft was wrong
fingerprinted |
MIT (permissive) |
| TokenCarve: Information-Preserving Visual Token Compression in Multimodal Large Language Models |
13 Mar 2025 |
shawntan86/tokencarve/TokenCarve/TokenCarve_model_vqa_loader.py 076c252c52cbb161 |
ran · our draft was wrong
fingerprinted |
Apache-2.0 (permissive) |
| Re-Align: Aligning Vision Language Models via Retrieval-Augmented Direct Preference Optimization |
18 Feb 2025 |
taco-group/re-align/llava/eval/model_vqa.py 076c252c52cbb161 |
ran · our draft was wrong
fingerprinted |
Apache-2.0 (permissive) |
| Stop Looking for Important Tokens in Multimodal Language Models: Duplication Matters More |
17 Feb 2025 |
zichenwen1/dart/llava/eval/model_vqa.py 076c252c52cbb161 |
ran · our draft was wrong
fingerprinted |
Apache-2.0 (permissive) |
| HealthGPT: A Medical Large Vision-Language Model for Unifying Comprehension and Generation via Heterogeneous Knowledge Adaptation |
14 Feb 2025 |
dcdmllm/healthgpt/HealthGPT/llava/eval/model_vqa.py 076c252c52cbb161 |
ran · our draft was wrong
fingerprinted |
Apache-2.0 (permissive) |
| EVEv2: Improved Baselines for Encoder-Free Vision-Language Models |
10 Feb 2025 |
baaivision/EVE/EVEv1/eve/eval/model_vqa.py 076c252c52cbb161 |
ran · our draft was wrong
fingerprinted |
MIT (permissive) |
| VideoRoPE: What Makes for Good Video Rotary Position Embedding? |
7 Feb 2025 |
wiselnn570/videorope/eval/model_longvideobench_qwen2_vl.py 076c252c52cbb161 |
ran · our draft was wrong
fingerprinted |
Apache-2.0 (permissive) |
| MedRAX: Medical Reasoning Agent for Chest X-ray |
4 Feb 2025 |
bowang-lab/medrax/medrax/llava/eval/model_vqa.py 076c252c52cbb161 |
ran · our draft was wrong
fingerprinted |
Apache-2.0 (permissive) |
| Mitigating Hallucinations in Large Vision-Language Models via DPO: On-Policy Data Hold the Key |
16 Jan 2025 |
zhyang2226/opa-dpo/eval_llava_rlhf_coco/model_vqa.py 076c252c52cbb161 |
ran · our draft was wrong
fingerprinted |
MIT recorded; this copy not marked cleared · pointer only |
| Tarsier2: Advancing Large Vision-Language Models from Detailed Video Description to Comprehensive Video Understanding |
14 Jan 2025 |
identical code first harvested elsewhere 076c252c52cbb161 |
ran · our draft was wrong
fingerprinted |
licence of this copy not recorded |
| OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning |
31 Dec 2024 |
yuliang-liu/multimodalocr/OCRBench/example.py cb9eaee96a223830 |
ran
fingerprinted |
MIT (permissive) |
| LLaVA Steering: Visual Instruction Tuning with 500x Fewer Parameters through Modality Linear Representation-Steering |
16 Dec 2024 |
bibisbar/LLaVA-Steering/tinyllava/eval/model_vqa.py 076c252c52cbb161 |
ran · our draft was wrong
fingerprinted |
Apache-2.0 (permissive) |
| LLaVA Steering: Visual Instruction Tuning with 500x Fewer Parameters through Modality Linear Representation-Steering |
16 Dec 2024 |
bibisbar/LLaVA-Steering/tinyllava/eval/model_vqa_chair.py 803512fe1dd82fde |
unverified |
Apache-2.0 (permissive) |
| Lyra: An Efficient and Speech-Centric Framework for Omni-Cognition |
12 Dec 2024 |
dvlab-research/Lyra/lyra/eval/model_lyra_image_speech.py 076c252c52cbb161 |
ran · our draft was wrong
fingerprinted |
Apache-2.0 (permissive) |
| DriveMM: All-in-One Large Multimodal Model for Autonomous Driving |
10 Dec 2024 |
zhijian11/DriveMM/llava/eval/model_vqa.py 076c252c52cbb161 |
ran · our draft was wrong
fingerprinted |
Apache-2.0 (permissive) |
| MMedPO: Aligning Medical Vision-Language Models with Clinical-Aware Multimodal Preference Optimization |
9 Dec 2024 |
aiming-lab/mmedpo/inference/llava-med-1.5_report.py 076c252c52cbb161 |
ran · our draft was wrong
fingerprinted |
Apache-2.0 (permissive) |
| [CLS] Token Tells Everything Needed for Training-free Efficient MLLMs |
8 Dec 2024 |
thu-mig/vtc-cls/llava/eval/model_vqa.py 076c252c52cbb161 |
ran · our draft was wrong
fingerprinted |
Apache-2.0 (permissive) |
| TeamCraft: A Benchmark for Multi-Modal Multi-Agent Systems in Minecraft |
6 Dec 2024 |
teamcraft-bench/teamcraft/llava_teamcraft/llava/eval/model_vqa.py 076c252c52cbb161 |
ran · our draft was wrong
fingerprinted |
MIT (permissive) |
| p-MoD: Building Mixture-of-Depths MLLMs via Progressive Ratio Decay |
5 Dec 2024 |
mcg-nju/p-mod/llava/eval/model_vqa.py 076c252c52cbb161 |
ran · our draft was wrong
fingerprinted |
Apache-2.0 (permissive) |
| FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression |
5 Dec 2024 |
codefanw/flashsloth/flashsloth/eval/model_vqa.py 076c252c52cbb161 |
ran · our draft was wrong
fingerprinted |
no licence file found · pointer only |
| AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning |
4 Dec 2024 |
lavi-lab/aim/llava/eval/model_vqa.py 076c252c52cbb161 |
ran · our draft was wrong
fingerprinted |
Apache-2.0 (permissive) |
| Agri-LLaVA: Knowledge-Infused Large Multimodal Assistant on Agricultural Pests and Diseases |
3 Dec 2024 |
kki2eve/agri-llava/agri_llava/eval/model_vqa.py 076c252c52cbb161 |
ran · our draft was wrong
fingerprinted |
Apache-2.0 (permissive) |
| Beyond Text-Visual Attention: Exploiting Visual Cues for Effective Token Pruning in VLMs |
2 Dec 2024 |
theia-4869/fastervlm/llava/eval/model_vqa.py 076c252c52cbb161 |
ran · our draft was wrong
fingerprinted |
Apache-2.0 (permissive) |
| Dynamic-LLaVA: Efficient Multimodal Large Language Models via Dynamic Vision-language Context Sparsification |
1 Dec 2024 |
Osilly/dynamic_llava/llava/dynamic_eval/model_lvis_for_meteor.py 076c252c52cbb161 |
ran · our draft was wrong
fingerprinted |
Apache-2.0 (permissive) |
| Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering |
25 Nov 2024 |
aimagelab/reflectiva/llava/eval/model_vqa.py 076c252c52cbb161 |
ran · our draft was wrong
fingerprinted |
Apache-2.0 (permissive) |
| Devils in Middle Layers of Large Vision-Language Models: Interpreting, Detecting and Mitigating Object Hallucinations via Attention Lens |
23 Nov 2024 |
zhangqijiang07/middle_layers_indicating_hallucinations/llava/eval/model_vqa.py 076c252c52cbb161 |
ran · our draft was wrong
fingerprinted |
no licence file found · pointer only |
| DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models |
22 Nov 2024 |
kd-tao/dycoke/llava/eval/model_vqa.py 076c252c52cbb161 |
ran · our draft was wrong
fingerprinted |
Apache-2.0 (permissive) |
| Insight-V: Exploring Long-Chain Visual Reasoning with Multimodal Large Language Models |
21 Nov 2024 |
dongyh20/insight-v/llava/eval/model_vqa.py 076c252c52cbb161 |
ran · our draft was wrong
fingerprinted |
no licence file found · pointer only |
| CCExpert: Advancing MLLM Capability in Remote Sensing Change Captioning with Difference-Aware Integration and a Foundational Dataset |
18 Nov 2024 |
meize0729/ccexpert/llava/eval/model_vqa.py 076c252c52cbb161 |
ran · our draft was wrong
fingerprinted |
Apache-2.0 (permissive) |
| TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models |
17 Nov 2024 |
tingyu215/ts-llava/llava/eval/run_inference_benchmark_consistency.py 076c252c52cbb161 |
ran · our draft was wrong
fingerprinted |
Apache-2.0 (permissive) |
| Multi-Stage Vision Token Dropping: Towards Efficient Multimodal Large Language Model |
16 Nov 2024 |
liuting20/mustdrop/llava/eval/model_vqa.py 076c252c52cbb161 |
ran · our draft was wrong
fingerprinted |
no licence file found · pointer only |
| V-DPO: Mitigating Hallucination in Large Vision Language Models via Vision-Guided Direct Preference Optimization |
5 Nov 2024 |
yuxixie/v-dpo/llava_dpo/eval/model_vqa.py 076c252c52cbb161 |
ran · our draft was wrong
fingerprinted |
Apache-2.0 (permissive) |
| R-CoT: Reverse Chain-of-Thought Problem Generation for Geometric Reasoning in Large Multimodal Models |
23 Oct 2024 |
dle666/r-cot/GeoQA_test/model_vqa_rcot7b.py 076c252c52cbb161 |
ran · our draft was wrong
fingerprinted |
no licence file found · pointer only |
| MIA-DPO: Multi-Image Augmented Direct Preference Optimization For Large Vision-Language Models |
23 Oct 2024 |
liuziyu77/mia-dpo/LLaVA-Hound-DPO/chatuniv/ChatUniVi/eval/model_coco_vqa.py 076c252c52cbb161 |
ran · our draft was wrong
fingerprinted |
Apache-2.0 (permissive) |
| PyramidDrop: Accelerating Your Large Vision-Language Models via Pyramid Visual Redundancy Reduction |
22 Oct 2024 |
cooperx521/pyramiddrop/llava/eval/model_vqa.py 076c252c52cbb161 |
ran · our draft was wrong
fingerprinted |
MIT (permissive) |
| LLaVA-KD: A Framework of Distilling Multimodal Large Language Models |
21 Oct 2024 |
Fantasyele/LLaVA-KD/llavakd/eval/model_vqa.py 076c252c52cbb161 |
ran · our draft was wrong
fingerprinted |
no licence file found · pointer only |
| Improve Vision Language Model Chain-of-thought Reasoning |
21 Oct 2024 |
riflezhang/llava-hound-dpo/llava_hound_dpo/inference/inference_dpo_reward.py 076c252c52cbb161 |
ran · our draft was wrong
fingerprinted |
no licence file found · pointer only |
| RAG-DDR: Optimizing Retrieval-Augmented Generation Using Differentiable Data Rewards |
17 Oct 2024 |
openmatch/rag-ddr/src/knowledgeRefinement/kr_inference.py e2f5857da2c35d73 |
ran
|
MIT (permissive) |
| MMed-RAG: Versatile Multimodal RAG System for Medical Vision Language Models |
16 Oct 2024 |
richard-peng-xia/MMed-RAG/train/dpo/povid_infer.py 076c252c52cbb161 |
ran · our draft was wrong
fingerprinted |
MIT (permissive) |
| MoH: Multi-Head Attention as Mixture-of-Head Attention |
15 Oct 2024 |
pku-yuangroup/chat-univi/ChatUniVi/eval/model_coco_vqa.py 076c252c52cbb161 |
ran · our draft was wrong
fingerprinted |
Apache-2.0 (permissive) |
| MF-LAL: Drug Compound Generation Using Multi-Fidelity Latent Space Active Learning |
15 Oct 2024 |
Rose-STL-Lab/MF-LAL/BAT.py/BAT-brd4-updated/lib/equil-sdr.py bea3d38d64526c35 |
ran
fingerprinted |
no licence file found · pointer only |
| Q-VLM: Post-training Quantization for Large Vision-Language Models |
10 Oct 2024 |
changyuanwang17/qvlm/llava/eval/model_vqa.py 076c252c52cbb161 |
ran · our draft was wrong
fingerprinted |
Apache-2.0 (permissive) |
| Deciphering Cross-Modal Alignment in Large Vision-Language Models with Modality Integration Rate |
9 Oct 2024 |
shikiw/modality-integration-rate/llava/eval/model_vqa.py 076c252c52cbb161 |
ran · our draft was wrong
fingerprinted |
MIT (permissive) |
| Personalized Visual Instruction Tuning |
9 Oct 2024 |
sterzhang/pvit/personalize-llava/llava/eval/model_vqa.py 076c252c52cbb161 |
ran · our draft was wrong
fingerprinted |
no licence file found · pointer only |
| ETA: Evaluating Then Aligning Safety of Vision Language Models at Inference Time |
9 Oct 2024 |
dripnowhy/eta/llava/eval/model_vqa_loader_eta.py 076c252c52cbb161 |
ran · our draft was wrong
fingerprinted |
no licence file found · pointer only |
| Treat Visual Tokens as Text? But Your MLLM Only Needs Fewer Efforts to See |
8 Oct 2024 |
ZhangAIPI/YOPO_MLLM_Pruning/LLaVA/llava/eval/model_vqa.py 076c252c52cbb161 |
ran · our draft was wrong
fingerprinted |
MIT (permissive) |
| SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference |
6 Oct 2024 |
Gumpest/SparseVLMs/llava/eval/model_vqa.py 076c252c52cbb161 |
ran · our draft was wrong
fingerprinted |
Apache-2.0 (permissive) |
| Look Twice Before You Answer: Memory-Space Visual Retracing for Hallucination Mitigation in Multimodal Large Language Models |
4 Oct 2024 |
1zhou-Wang/MemVR/eval/glm_model_vqa.py 076c252c52cbb161 |
ran · our draft was wrong
fingerprinted |
Apache-2.0 (permissive) |
| EMMA: Efficient Visual Alignment in Multi-Modal LLMs |
2 Oct 2024 |
saraghazanfari/emma/llava/eval/model_vqa.py 076c252c52cbb161 |
ran · our draft was wrong
fingerprinted |
Apache-2.0 (permissive) |
| AVG-LLaVA: A Large Multimodal Model with Adaptive Visual Granularity |
20 Sep 2024 |
deeplearnxmu/avg-llava/llava/eval/model_vqa.py 076c252c52cbb161 |
ran · our draft was wrong
fingerprinted |
Apache-2.0 (permissive) |
| Explanation Bottleneck Models |
26 Sep 2024 |
yshinya6/xbm/xbm-llava/llava/eval/model_vqa.py 076c252c52cbb161 |
ran · our draft was wrong
fingerprinted |
no licence file found · pointer only |
| MECD: Unlocking Multi-Event Causal Discovery in Video Reasoning |
26 Sep 2024 |
identical code first harvested elsewhere 076c252c52cbb161 |
ran · our draft was wrong
fingerprinted |
licence of this copy not recorded |
| EventHallusion: Diagnosing Event Hallucinations in Video LLMs |
25 Sep 2024 |
identical code first harvested elsewhere 076c252c52cbb161 |
ran · our draft was wrong
fingerprinted |
licence of this copy not recorded |
| CDChat: A Large Multimodal Model for Remote Sensing Change Description |
24 Sep 2024 |
techmn/cdchat/cdchat/eval/batch_cdchat_vqa.py 076c252c52cbb161 |
ran · our draft was wrong
fingerprinted |
no licence file found · pointer only |
| M$^2$PT: Multimodal Prompt Tuning for Zero-shot Instruction Learning |
24 Sep 2024 |
william-wang618/mmpt-emnlp2024/M2PT/eval/model_vqa.py 076c252c52cbb161 |
ran · our draft was wrong
fingerprinted |
no licence file found · pointer only |
| FineCops-Ref: A new Dataset and Task for Fine-Grained Compositional Referring Expression Comprehension |
23 Sep 2024 |
identical code first harvested elsewhere 076c252c52cbb161 |
ran · our draft was wrong
fingerprinted |
licence of this copy not recorded |
| SURf: Teaching Large Vision-Language Models to Selectively Utilize Retrieved Information |
21 Sep 2024 |
GasolSun36/SURf/initial/generate_initial_data.py 076c252c52cbb161 |
ran · our draft was wrong
fingerprinted |
no licence file found · pointer only |
| SURf: Teaching Large Vision-Language Models to Selectively Utilize Retrieved Information |
21 Sep 2024 |
gasolsun36/surf/initial/tool_evaluate.py abe4f5d245fa207a |
ran
fingerprinted |
no licence file found · pointer only |
| JourneyBench: A Challenging One-Stop Vision-Language Understanding Benchmark of Generated Images |
19 Sep 2024 |
journeybench/journeybench/automatic-qa-generator/baseline/llava_model_vcr.py 076c252c52cbb161 |
ran · our draft was wrong
fingerprinted |
no licence file found · pointer only |
| Less is More: A Simple yet Effective Token Reduction Method for Efficient Multi-modal LLMs |
17 Sep 2024 |
freedomintelligence/trim/llava/eval/model_vqa.py 076c252c52cbb161 |
ran · our draft was wrong
fingerprinted |
Apache-2.0 (permissive) |
| MotIF: Motion Instruction Fine-tuning |
16 Sep 2024 |
Minyoung1005/motif/LLaVA/llava/eval/model_vqa.py 076c252c52cbb161 |
ran · our draft was wrong
fingerprinted |
no licence file found · pointer only |
| Fit and Prune: Fast and Training-free Visual Token Pruning for Multi-modal Large Language Models |
16 Sep 2024 |
ywh187/fitprune/LLaVA_1.5/llava/eval/model_vqa.py 076c252c52cbb161 |
ran · our draft was wrong
fingerprinted |
no licence file found · pointer only |
| LIME: Less Is More for MLLM Evaluation |
10 Sep 2024 |
kangreen0210/lime/llava_next_110B.py 076c252c52cbb161 |
ran · our draft was wrong
fingerprinted |
no licence file found · pointer only |
| LLaMA-Omni: Seamless Speech Interaction with Large Language Models |
10 Sep 2024 |
identical code first harvested elsewhere 076c252c52cbb161 |
ran · our draft was wrong
fingerprinted |
licence of this copy not recorded |
| HiPrompt: Tuning-free Higher-Resolution Generation with Hierarchical MLLM Prompts |
4 Sep 2024 |
Liuxinyv/HiPrompt/LLaVA/llava/eval/model_vqa.py 076c252c52cbb161 |
ran · our draft was wrong
fingerprinted |
no licence file found · pointer only |
| LongLLaVA: Scaling Multi-modal LLMs to 1000 Images Efficiently via a Hybrid Architecture |
4 Sep 2024 |
freedomintelligence/longllava/benchmarks/MVBench/model_mvbench_qa.py 076c252c52cbb161 |
ran · our draft was wrong
fingerprinted |
no licence file found · pointer only |
| Recoverable Compression: A Multimodal Vision Token Recovery Mechanism Guided by Text Information |
2 Sep 2024 |
banjiuyufen/Recoverable-Compression/llava/eval/model_vqa.py 076c252c52cbb161 |
ran · our draft was wrong
fingerprinted |
Apache-2.0 (permissive) |
| MultiMath: Bridging Visual and Mathematical Reasoning for Large Language Models |
30 Aug 2024 |
pengshuai-rin/multimath/eval_mathverse/infer.py 076c252c52cbb161 |
ran · our draft was wrong
fingerprinted |
Apache-2.0 (permissive) |