| Beyond Language Priors: Diagnosing and Fixing Visual-Origin Hallucinations in Multimodal LLM added by Syntology |
2026-09 (from id) |
zxp555/ACFT_MM26/ACFT/llava/eval/model_vqa.py 42a46570620cd9fa |
ran · our draft was wrong
fingerprinted |
no licence file found · pointer only |
| Certified but Private: Scalable Zero-Knowledge Proofs for Neural Network Guarantees added by Syntology |
2026-08 (from id) |
youweizhong/PANDA/evaluation/config.py c39616a6f06d7a49 |
unverified |
MIT (permissive) |
| DIVE: Dynamic Iterative Visual Evidence Construction for Efficient Vision-Language Models added by Syntology |
2026-08 (from id) |
Zhong-Chenchen/DIVE/llava/eval/model_vqa.py 42a46570620cd9fa |
ran · our draft was wrong
fingerprinted |
Apache-2.0 (permissive) |
| MLLM-DataEngine: Closing the Loop of Multimodal Instruction Tuning Data Generation added by Syntology |
2026-07 (from id) |
opendatalab/MLLM-DataEngine/LLaVA/llava/eval/model_vqa.py 42a46570620cd9fa |
ran · our draft was wrong
fingerprinted |
Apache-2.0 (permissive) |
| OralAgent: Integrating Reasoning, Tools, and Knowledge for Interactive Dental Image Analysis added by Syntology |
2026-05 (from id) |
isjinghao/OralAgent/oralagent/llava/eval/model_vqa.py 42a46570620cd9fa |
ran · our draft was wrong
fingerprinted |
Apache-2.0 (permissive) |
| WSVD: Weighted Low-Rank Approximation for Fast and Efficient Execution of Low-Precision Vision-Language Models added by Syntology |
2026-04 (from id) |
SAI-Lab-NYU/WSVD/e2e/infer_llava.py 42a46570620cd9fa |
ran · our draft was wrong
fingerprinted |
Apache-2.0 (permissive) |
| Uncertainty-Aware Knowledge Distillation for Multimodal Large Language Models added by Syntology |
2026-03 (from id) |
Jingchensun/beta-kd/mobilevlm/eval/model_vqa_loader.py 42a46570620cd9fa |
ran · our draft was wrong
fingerprinted |
no licence file found · pointer only |
| HiDrop: Hierarchical Vision Token Reduction in MLLMs via Late Injection, Concave Pyramid Pruning, and Early Exit added by Syntology |
2026-02 (from id) |
EIT-NLP/HiDrop/llava/eval/model_vqa.py 42a46570620cd9fa |
ran · our draft was wrong
fingerprinted |
Apache-2.0 (permissive) |
| Beyond Next-Token Alignment: Distilling Multimodal Large Language Models via Token Interactions added by Syntology |
2026-02 (from id) |
lchen1019/Align-TI/alignti/eval/model_vqa.py 42a46570620cd9fa |
ran · our draft was wrong
fingerprinted |
Apache-2.0 (permissive) |
| Benchmarking Direct Preference Optimization for Medical Large Vision-Language Models added by Syntology |
2026-01 (from id) |
dmis-lab/med-vlm-dpo/inference/LLaVA-Med/llava/eval/model_vqa.py 42a46570620cd9fa |
ran · our draft was wrong
fingerprinted |
no licence file found · pointer only |
| DeepMoLM: Leveraging Visual and Geometric Structural Information for Molecule-Text Modeling added by Syntology |
2026-01 (from id) |
1anj/DeepMoLM/llava/eval/model_vqa_video.py 42a46570620cd9fa |
ran · our draft was wrong
fingerprinted |
Apache-2.0 (permissive) |
| Cross-Layer Injection for Deep Vision-Language Fusion added by Syntology |
2026-01 (from id) |
codefuse-ai/CLI/llava/eval/model_vqa.py 42a46570620cd9fa |
ran · our draft was wrong
fingerprinted |
Apache-2.0 (permissive) |
| QSVD: Efficient Low-rank Approximation for Unified Query-Key-Value Weight Compression in Low-Precision Vision-Language Models added by Syntology |
2025-10 (from id) |
SAI-Lab-NYU/QSVD/fake_quant/eval_llavanext_vizwiz.py 42a46570620cd9fa |
ran · our draft was wrong
fingerprinted |
Apache-2.0 (permissive) |
| Spatial Preference Rewarding for MLLMs Spatial Understanding added by Syntology |
2025-10 (from id) |
hanqiu-hq/SPR/construct_data/ferret_score_siglip.py 42a46570620cd9fa |
ran · our draft was wrong
fingerprinted |
no licence file found · pointer only |
| arXiv:2507.18300 |
2025-07 (from id) |
360CVGroup/LMM-Det/llava/eval/model_coco_owlv2.py 42a46570620cd9fa |
ran · our draft was wrong
fingerprinted |
Apache-2.0 (permissive) |
| LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs |
1 Jul 2025 |
CnFaker/LLaVA-SP/llava/eval/model_vqa.py 42a46570620cd9fa |
ran · our draft was wrong
fingerprinted |
Apache-2.0 (permissive) |
| UniFork: Exploring Modality Alignment for Unified Multimodal Understanding and Generation |
20 Jun 2025 |
tliby/unifork/unifork/eval/model_vqa.py 42a46570620cd9fa |
ran · our draft was wrong
fingerprinted |
no licence file found · pointer only |
| Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs |
12 Jun 2025 |
theia-4869/cdpruner/llava/eval/model_vqa.py 42a46570620cd9fa |
ran · our draft was wrong
fingerprinted |
Apache-2.0 (permissive) |
| Learning from Videos for 3D World: Enhancing MLLMs with 3D Vision Geometry Priors |
30 May 2025 |
LaVi-Lab/Video-3D-LLM/llava/eval/model_vqa.py 42a46570620cd9fa |
ran · our draft was wrong
fingerprinted |
Apache-2.0 (permissive) |
| VideoReasonBench: Can MLLMs Perform Vision-Centric Complex Video Reasoning? |
29 May 2025 |
llyx97/video_reason_bench/eval_api.py cda3e7eebb6da37a |
ran · our draft was wrong
fingerprinted |
no licence file found · pointer only |
| Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models |
27 May 2025 |
jefferyzhan/griffon/griffon/eval/model_vqa.py 42a46570620cd9fa |
ran · our draft was wrong
fingerprinted |
Apache-2.0 (permissive) |
| LaViDa: A Large Diffusion Language Model for Multimodal Understanding |
22 May 2025 |
jacklishufan/lavida/llava/eval/model_vqa.py 42a46570620cd9fa |
ran · our draft was wrong
fingerprinted |
Apache-2.0 (permissive) |
| R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning |
5 May 2025 |
yfzhang114/r1_reward/inference/MM-RLHF-Reward/r1_reward.py 42a46570620cd9fa |
ran · our draft was wrong
fingerprinted |
Apache-2.0 (permissive) |
| Unlearning Sensitive Information in Multimodal LLMs: Benchmark and Attack-Defense Evaluation |
1 May 2025 |
vaidehi99/unlok-vqa/LLaVA/llava/eval/model_vqa.py 42a46570620cd9fa |
ran · our draft was wrong
fingerprinted |
MIT (permissive) |
| Multi-Resolution Pathology-Language Pre-training Model with Text-Guided Visual Representation |
26 Apr 2025 |
BasitAlawode/MR-PLIP/generate_text.py 42a46570620cd9fa |
ran · our draft was wrong
fingerprinted |
MIT (permissive) |
| Benchmarking Large Vision-Language Models on Fine-Grained Image Tasks: A Comprehensive Evaluation |
21 Apr 2025 |
SEU-VIPGroup/FG-BMK/demo/human_evaluation/human_evaluation_demo.py 49a93847dfeff479 |
ran · our draft was wrong
fingerprinted |
no licence file found · pointer only |
| SoTA with Less: MCTS-Guided Sample Selection for Data-Efficient Visual Reasoning Self-Improvement |
10 Apr 2025 |
identical code first harvested elsewhere 42a46570620cd9fa |
ran · our draft was wrong
fingerprinted |
licence of this copy not recorded |
| SoTA with Less: MCTS-Guided Sample Selection for Data-Efficient Visual Reasoning Self-Improvement |
10 Apr 2025 |
si0wang/thinklite-vl/eval/model_ai2d_qwen.py e6e072225bd8f400 |
ran · our draft was wrong
fingerprinted |
no licence file found · pointer only |
| arXiv:2504.00502 |
2025-04 (from id) |
icip-cas/ShortV/llava/eval/model_vqa.py 42a46570620cd9fa |
ran · our draft was wrong
fingerprinted |
Apache-2.0 (permissive) |
| MaTVLM: Hybrid Mamba-Transformer for Efficient Vision-Language Modeling |
17 Mar 2025 |
hustvl/MaTVLM/tinyllava/eval/model_vqa.py 42a46570620cd9fa |
ran · our draft was wrong
fingerprinted |
Apache-2.0 (permissive) |
| Safety Mirage: How Spurious Correlations Undermine VLM Safety Fine-tuning |
14 Mar 2025 |
OPTML-Group/VLM-Safety-Unlearn/llava/eval/model_vqa.py 42a46570620cd9fa |
ran · our draft was wrong
fingerprinted |
MIT (permissive) |
| Keyframe-oriented Vision Token Pruning: Enhancing Efficiency of Large Vision Language Models on Long-Form Video Processing |
13 Mar 2025 |
1999Lyd/KVTP/llava/eval/model_vqa.py 42a46570620cd9fa |
ran · our draft was wrong
fingerprinted |
MIT (permissive) |
| TokenCarve: Information-Preserving Visual Token Compression in Multimodal Large Language Models |
13 Mar 2025 |
shawntan86/tokencarve/TokenCarve/TokenCarve_model_vqa_loader.py 42a46570620cd9fa |
ran · our draft was wrong
fingerprinted |
Apache-2.0 (permissive) |
| Re-Align: Aligning Vision Language Models via Retrieval-Augmented Direct Preference Optimization |
18 Feb 2025 |
taco-group/re-align/llava/eval/model_vqa.py 42a46570620cd9fa |
ran · our draft was wrong
fingerprinted |
Apache-2.0 (permissive) |
| Stop Looking for Important Tokens in Multimodal Language Models: Duplication Matters More |
17 Feb 2025 |
zichenwen1/dart/llava/eval/model_vqa.py 42a46570620cd9fa |
ran · our draft was wrong
fingerprinted |
Apache-2.0 (permissive) |
| HealthGPT: A Medical Large Vision-Language Model for Unifying Comprehension and Generation via Heterogeneous Knowledge Adaptation |
14 Feb 2025 |
dcdmllm/healthgpt/HealthGPT/llava/eval/model_vqa.py 42a46570620cd9fa |
ran · our draft was wrong
fingerprinted |
Apache-2.0 (permissive) |
| EVEv2: Improved Baselines for Encoder-Free Vision-Language Models |
10 Feb 2025 |
baaivision/EVE/EVEv1/eve/eval/model_vqa.py 42a46570620cd9fa |
ran · our draft was wrong
fingerprinted |
MIT (permissive) |
| VideoRoPE: What Makes for Good Video Rotary Position Embedding? |
7 Feb 2025 |
wiselnn570/videorope/eval/model_longvideobench_qwen2_vl.py e32e49fba862aa72 |
unverified |
Apache-2.0 (permissive) |
| VideoRoPE: What Makes for Good Video Rotary Position Embedding? |
7 Feb 2025 |
wiselnn570/videorope/eval/model_videohallucer.py b6f76fcc05062c30 |
unverified |
Apache-2.0 (permissive) |
| MedRAX: Medical Reasoning Agent for Chest X-ray |
4 Feb 2025 |
bowang-lab/medrax/medrax/llava/eval/model_vqa.py 42a46570620cd9fa |
ran · our draft was wrong
fingerprinted |
Apache-2.0 (permissive) |
| Mitigating Hallucinations in Large Vision-Language Models via DPO: On-Policy Data Hold the Key |
16 Jan 2025 |
zhyang2226/opa-dpo/eval_llava_rlhf_coco/model_vqa.py 42a46570620cd9fa |
ran · our draft was wrong
fingerprinted |
MIT recorded; this copy not marked cleared · pointer only |
| Tarsier2: Advancing Large Vision-Language Models from Detailed Video Description to Comprehensive Video Understanding |
14 Jan 2025 |
identical code first harvested elsewhere 42a46570620cd9fa |
ran · our draft was wrong
fingerprinted |
licence of this copy not recorded |
| LLaVA Steering: Visual Instruction Tuning with 500x Fewer Parameters through Modality Linear Representation-Steering |
16 Dec 2024 |
bibisbar/LLaVA-Steering/tinyllava/eval/model_vqa.py 42a46570620cd9fa |
ran · our draft was wrong
fingerprinted |
Apache-2.0 (permissive) |
| LLaVA Steering: Visual Instruction Tuning with 500x Fewer Parameters through Modality Linear Representation-Steering |
16 Dec 2024 |
bibisbar/LLaVA-Steering/tinyllava/eval/model_vqa_chair.py 01c6b696dad5f567 |
unverified |
Apache-2.0 (permissive) |
| Lyra: An Efficient and Speech-Centric Framework for Omni-Cognition |
12 Dec 2024 |
dvlab-research/Lyra/lyra/eval/model_lyra_image_speech.py 42a46570620cd9fa |
ran · our draft was wrong
fingerprinted |
Apache-2.0 (permissive) |
| DriveMM: All-in-One Large Multimodal Model for Autonomous Driving |
10 Dec 2024 |
zhijian11/DriveMM/llava/eval/model_vqa.py 42a46570620cd9fa |
ran · our draft was wrong
fingerprinted |
Apache-2.0 (permissive) |
| MMedPO: Aligning Medical Vision-Language Models with Clinical-Aware Multimodal Preference Optimization |
9 Dec 2024 |
aiming-lab/mmedpo/inference/llava-med-1.5_report.py 42a46570620cd9fa |
ran · our draft was wrong
fingerprinted |
Apache-2.0 (permissive) |
| [CLS] Token Tells Everything Needed for Training-free Efficient MLLMs |
8 Dec 2024 |
thu-mig/vtc-cls/llava/eval/model_vqa.py 42a46570620cd9fa |
ran · our draft was wrong
fingerprinted |
Apache-2.0 (permissive) |
| TeamCraft: A Benchmark for Multi-Modal Multi-Agent Systems in Minecraft |
6 Dec 2024 |
teamcraft-bench/teamcraft/llava_teamcraft/llava/eval/model_vqa.py 42a46570620cd9fa |
ran · our draft was wrong
fingerprinted |
MIT (permissive) |
| p-MoD: Building Mixture-of-Depths MLLMs via Progressive Ratio Decay |
5 Dec 2024 |
mcg-nju/p-mod/llava/eval/model_vqa.py 42a46570620cd9fa |
ran · our draft was wrong
fingerprinted |
Apache-2.0 (permissive) |
| FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression |
5 Dec 2024 |
codefanw/flashsloth/flashsloth/eval/model_vqa.py 42a46570620cd9fa |
ran · our draft was wrong
fingerprinted |
no licence file found · pointer only |
| AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning |
4 Dec 2024 |
lavi-lab/aim/llava/eval/model_vqa.py 42a46570620cd9fa |
ran · our draft was wrong
fingerprinted |
Apache-2.0 (permissive) |
| Agri-LLaVA: Knowledge-Infused Large Multimodal Assistant on Agricultural Pests and Diseases |
3 Dec 2024 |
kki2eve/agri-llava/agri_llava/eval/model_vqa.py 42a46570620cd9fa |
ran · our draft was wrong
fingerprinted |
Apache-2.0 (permissive) |
| Beyond Text-Visual Attention: Exploiting Visual Cues for Effective Token Pruning in VLMs |
2 Dec 2024 |
theia-4869/fastervlm/llava/eval/model_vqa.py 42a46570620cd9fa |
ran · our draft was wrong
fingerprinted |
Apache-2.0 (permissive) |
| Dynamic-LLaVA: Efficient Multimodal Large Language Models via Dynamic Vision-language Context Sparsification |
1 Dec 2024 |
Osilly/dynamic_llava/llava/dynamic_eval/model_lvis_for_meteor.py 42a46570620cd9fa |
ran · our draft was wrong
fingerprinted |
Apache-2.0 (permissive) |
| Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering |
25 Nov 2024 |
aimagelab/reflectiva/llava/eval/model_vqa.py 42a46570620cd9fa |
ran · our draft was wrong
fingerprinted |
Apache-2.0 (permissive) |
| Devils in Middle Layers of Large Vision-Language Models: Interpreting, Detecting and Mitigating Object Hallucinations via Attention Lens |
23 Nov 2024 |
zhangqijiang07/middle_layers_indicating_hallucinations/llava/eval/model_vqa.py 42a46570620cd9fa |
ran · our draft was wrong
fingerprinted |
no licence file found · pointer only |
| ZoomEye: Enhancing Multimodal LLMs with Human-Like Zooming Capabilities through Tree-Based Image Exploration |
25 Nov 2024 |
om-ai-lab/ZoomEye/ZoomEye/eval/perform_zoom_eye.py d61cdd0662f59fed |
ran · our draft was wrong
fingerprinted |
no licence file found · pointer only |
| DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models |
22 Nov 2024 |
kd-tao/dycoke/llava/eval/model_vqa.py 42a46570620cd9fa |
ran · our draft was wrong
fingerprinted |
Apache-2.0 (permissive) |
| Insight-V: Exploring Long-Chain Visual Reasoning with Multimodal Large Language Models |
21 Nov 2024 |
dongyh20/insight-v/llava/eval/model_vqa.py 42a46570620cd9fa |
ran · our draft was wrong
fingerprinted |
no licence file found · pointer only |
| CCExpert: Advancing MLLM Capability in Remote Sensing Change Captioning with Difference-Aware Integration and a Foundational Dataset |
18 Nov 2024 |
meize0729/ccexpert/llava/eval/model_vqa.py 42a46570620cd9fa |
ran · our draft was wrong
fingerprinted |
Apache-2.0 (permissive) |
| TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models |
17 Nov 2024 |
tingyu215/ts-llava/llava/eval/run_inference_benchmark_consistency.py 42a46570620cd9fa |
ran · our draft was wrong
fingerprinted |
Apache-2.0 (permissive) |
| Multi-Stage Vision Token Dropping: Towards Efficient Multimodal Large Language Model |
16 Nov 2024 |
liuting20/mustdrop/llava/eval/model_vqa.py 42a46570620cd9fa |
ran · our draft was wrong
fingerprinted |
no licence file found · pointer only |
| V-DPO: Mitigating Hallucination in Large Vision Language Models via Vision-Guided Direct Preference Optimization |
5 Nov 2024 |
yuxixie/v-dpo/llava_dpo/eval/model_vqa.py 42a46570620cd9fa |
ran · our draft was wrong
fingerprinted |
Apache-2.0 (permissive) |
| R-CoT: Reverse Chain-of-Thought Problem Generation for Geometric Reasoning in Large Multimodal Models |
23 Oct 2024 |
dle666/r-cot/GeoQA_test/model_vqa_rcot7b.py 42a46570620cd9fa |
ran · our draft was wrong
fingerprinted |
no licence file found · pointer only |
| MIA-DPO: Multi-Image Augmented Direct Preference Optimization For Large Vision-Language Models |
23 Oct 2024 |
liuziyu77/mia-dpo/LLaVA-Hound-DPO/chatuniv/ChatUniVi/eval/model_coco_vqa.py 42a46570620cd9fa |
ran · our draft was wrong
fingerprinted |
Apache-2.0 (permissive) |
| PyramidDrop: Accelerating Your Large Vision-Language Models via Pyramid Visual Redundancy Reduction |
22 Oct 2024 |
cooperx521/pyramiddrop/llava/eval/model_vqa.py 42a46570620cd9fa |
ran · our draft was wrong
fingerprinted |
MIT (permissive) |
| LLaVA-KD: A Framework of Distilling Multimodal Large Language Models |
21 Oct 2024 |
Fantasyele/LLaVA-KD/llavakd/eval/model_vqa.py 42a46570620cd9fa |
ran · our draft was wrong
fingerprinted |
no licence file found · pointer only |
| Improve Vision Language Model Chain-of-thought Reasoning |
21 Oct 2024 |
riflezhang/llava-hound-dpo/llava_hound_dpo/inference/run_inference_video_caption.py 42a46570620cd9fa |
ran · our draft was wrong
fingerprinted |
no licence file found · pointer only |
| MMed-RAG: Versatile Multimodal RAG System for Medical Vision Language Models |
16 Oct 2024 |
richard-peng-xia/MMed-RAG/train/dpo/povid_infer.py 42a46570620cd9fa |
ran · our draft was wrong
fingerprinted |
MIT (permissive) |
| MoH: Multi-Head Attention as Mixture-of-Head Attention |
15 Oct 2024 |
pku-yuangroup/chat-univi/ChatUniVi/eval/model_coco_vqa.py 42a46570620cd9fa |
ran · our draft was wrong
fingerprinted |
Apache-2.0 (permissive) |
| Text4Seg: Reimagining Image Segmentation as Text Generation |
13 Oct 2024 |
mc-lan/Text4Seg/ms-swift/text4seg/infer_refer_seg.py fec9d7679b542a15 |
ran
|
licence not identified · pointer only |
| Q-VLM: Post-training Quantization for Large Vision-Language Models |
10 Oct 2024 |
changyuanwang17/qvlm/llava/eval/model_vqa.py 42a46570620cd9fa |
ran · our draft was wrong
fingerprinted |
Apache-2.0 (permissive) |
| Deciphering Cross-Modal Alignment in Large Vision-Language Models with Modality Integration Rate |
9 Oct 2024 |
shikiw/modality-integration-rate/llava/eval/model_vqa.py 42a46570620cd9fa |
ran · our draft was wrong
fingerprinted |
MIT (permissive) |
| Personalized Visual Instruction Tuning |
9 Oct 2024 |
sterzhang/pvit/personalize-llava/llava/eval/model_vqa.py 42a46570620cd9fa |
ran · our draft was wrong
fingerprinted |
no licence file found · pointer only |
| ETA: Evaluating Then Aligning Safety of Vision Language Models at Inference Time |
9 Oct 2024 |
dripnowhy/eta/llava/eval/model_vqa_loader_eta.py 42a46570620cd9fa |
ran · our draft was wrong
fingerprinted |
no licence file found · pointer only |
| Treat Visual Tokens as Text? But Your MLLM Only Needs Fewer Efforts to See |
8 Oct 2024 |
ZhangAIPI/YOPO_MLLM_Pruning/LLaVA/llava/eval/model_vqa.py 42a46570620cd9fa |
ran · our draft was wrong
fingerprinted |
MIT (permissive) |
| SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference |
6 Oct 2024 |
Gumpest/SparseVLMs/llava/eval/model_vqa.py 42a46570620cd9fa |
ran · our draft was wrong
fingerprinted |
Apache-2.0 (permissive) |
| Look Twice Before You Answer: Memory-Space Visual Retracing for Hallucination Mitigation in Multimodal Large Language Models |
4 Oct 2024 |
1zhou-Wang/MemVR/eval/glm_model_vqa.py 42a46570620cd9fa |
ran · our draft was wrong
fingerprinted |
Apache-2.0 (permissive) |
| EMMA: Efficient Visual Alignment in Multi-Modal LLMs |
2 Oct 2024 |
saraghazanfari/emma/llava/eval/model_vqa.py 42a46570620cd9fa |
ran · our draft was wrong
fingerprinted |
Apache-2.0 (permissive) |
| AVG-LLaVA: A Large Multimodal Model with Adaptive Visual Granularity |
20 Sep 2024 |
deeplearnxmu/avg-llava/llava/eval/model_vqa.py 42a46570620cd9fa |
ran · our draft was wrong
fingerprinted |
Apache-2.0 (permissive) |
| Explanation Bottleneck Models |
26 Sep 2024 |
yshinya6/xbm/xbm-llava/llava/eval/model_vqa.py 42a46570620cd9fa |
ran · our draft was wrong
fingerprinted |
no licence file found · pointer only |
| MECD: Unlocking Multi-Event Causal Discovery in Video Reasoning |
26 Sep 2024 |
identical code first harvested elsewhere 42a46570620cd9fa |
ran · our draft was wrong
fingerprinted |
licence of this copy not recorded |
| EventHallusion: Diagnosing Event Hallucinations in Video LLMs |
25 Sep 2024 |
identical code first harvested elsewhere 42a46570620cd9fa |
ran · our draft was wrong
fingerprinted |
licence of this copy not recorded |
| CDChat: A Large Multimodal Model for Remote Sensing Change Description |
24 Sep 2024 |
techmn/cdchat/cdchat/eval/batch_cdchat_vqa.py 42a46570620cd9fa |
ran · our draft was wrong
fingerprinted |
no licence file found · pointer only |
| M$^2$PT: Multimodal Prompt Tuning for Zero-shot Instruction Learning |
24 Sep 2024 |
william-wang618/mmpt-emnlp2024/M2PT/eval/model_vqa.py 42a46570620cd9fa |
ran · our draft was wrong
fingerprinted |
no licence file found · pointer only |
| FineCops-Ref: A new Dataset and Task for Fine-Grained Compositional Referring Expression Comprehension |
23 Sep 2024 |
identical code first harvested elsewhere 42a46570620cd9fa |
ran · our draft was wrong
fingerprinted |
licence of this copy not recorded |
| SURf: Teaching Large Vision-Language Models to Selectively Utilize Retrieved Information |
21 Sep 2024 |
GasolSun36/SURf/initial/generate_initial_data.py 42a46570620cd9fa |
ran · our draft was wrong
fingerprinted |
no licence file found · pointer only |
| SURf: Teaching Large Vision-Language Models to Selectively Utilize Retrieved Information |
21 Sep 2024 |
GasolSun36/SURf/eval/pope.py 06d40f4711e62357 |
ran
fingerprinted |
no licence file found · pointer only |
| SURf: Teaching Large Vision-Language Models to Selectively Utilize Retrieved Information |
21 Sep 2024 |
gasolsun36/surf/initial/tool_evaluate.py 8a203abf91abfcad |
ran
fingerprinted |
no licence file found · pointer only |
| JourneyBench: A Challenging One-Stop Vision-Language Understanding Benchmark of Generated Images |
19 Sep 2024 |
journeybench/journeybench/automatic-qa-generator/baseline/llava_model_vcr.py 42a46570620cd9fa |
ran · our draft was wrong
fingerprinted |
no licence file found · pointer only |
| Less is More: A Simple yet Effective Token Reduction Method for Efficient Multi-modal LLMs |
17 Sep 2024 |
freedomintelligence/trim/llava/eval/model_vqa.py 42a46570620cd9fa |
ran · our draft was wrong
fingerprinted |
Apache-2.0 (permissive) |
| MotIF: Motion Instruction Fine-tuning |
16 Sep 2024 |
Minyoung1005/motif/LLaVA/llava/eval/model_vqa.py 42a46570620cd9fa |
ran · our draft was wrong
fingerprinted |
no licence file found · pointer only |
| Fit and Prune: Fast and Training-free Visual Token Pruning for Multi-modal Large Language Models |
16 Sep 2024 |
ywh187/fitprune/LLaVA_1.5/llava/eval/model_vqa.py 42a46570620cd9fa |
ran · our draft was wrong
fingerprinted |
no licence file found · pointer only |
| LIME: Less Is More for MLLM Evaluation |
10 Sep 2024 |
kangreen0210/lime/llava_next_110B.py 42a46570620cd9fa |
ran · our draft was wrong
fingerprinted |
no licence file found · pointer only |
| HiPrompt: Tuning-free Higher-Resolution Generation with Hierarchical MLLM Prompts |
4 Sep 2024 |
Liuxinyv/HiPrompt/LLaVA/llava/eval/model_vqa.py 42a46570620cd9fa |
ran · our draft was wrong
fingerprinted |
no licence file found · pointer only |
| LongLLaVA: Scaling Multi-modal LLMs to 1000 Images Efficiently via a Hybrid Architecture |
4 Sep 2024 |
freedomintelligence/longllava/benchmarks/MVBench/model_mvbench_qa.py 42a46570620cd9fa |
ran · our draft was wrong
fingerprinted |
no licence file found · pointer only |
| Recoverable Compression: A Multimodal Vision Token Recovery Mechanism Guided by Text Information |
2 Sep 2024 |
banjiuyufen/Recoverable-Compression/llava/eval/model_vqa.py 42a46570620cd9fa |
ran · our draft was wrong
fingerprinted |
Apache-2.0 (permissive) |
| MultiMath: Bridging Visual and Mathematical Reasoning for Large Language Models |
30 Aug 2024 |
pengshuai-rin/multimath/eval_mathverse/infer.py 42a46570620cd9fa |
ran · our draft was wrong
fingerprinted |
Apache-2.0 (permissive) |
| IAA: Inner-Adaptor Architecture Empowers Frozen Large Language Model with Multimodal Capabilities |
23 Aug 2024 |
360cvgroup/inner-adaptor-architecture/iaa/eval/model_vqa_loader_llama3.py 42a46570620cd9fa |
ran · our draft was wrong
fingerprinted |
Apache-2.0 (permissive) |
| Visual Agents as Fast and Slow Thinkers |
16 Aug 2024 |
guangyans/sys2-llava/ROILLaVA/llava/eval/model_vqa.py 42a46570620cd9fa |
ran · our draft was wrong
fingerprinted |
MIT (permissive) |
| Advancing Multimodal Large Language Models with Quantization-Aware Scale Learning for Efficient Adaptation |
7 Aug 2024 |
xjjxmu/qslaw/llava/eval/model_vqa.py 42a46570620cd9fa |
ran · our draft was wrong
fingerprinted |
Apache-2.0 (permissive) |
| MedTrinity-25M: A Large-scale Multimodal Dataset with Multigranular Annotations for Medicine |
6 Aug 2024 |
UCSC-VLAA/MedTrinity-25M/llava/eval/model_vqa.py 42a46570620cd9fa |
ran · our draft was wrong
fingerprinted |
no licence file found · pointer only |
| Paying More Attention to Image: A Training-Free Method for Alleviating Hallucination in LVLMs |
31 Jul 2024 |
hasanar1f/llava-hallunication-fix/modPAI/llava/eval/model_vqa.py 42a46570620cd9fa |
ran · our draft was wrong
fingerprinted |
Apache-2.0 (permissive) |
| MLLM Is a Strong Reranker: Advancing Multimodal Retrieval-augmented Generation via Knowledge-enhanced Reranking and Noise-injected Training |
31 Jul 2024 |
idea-finai/ragllava/llava/eval/model_vqa.py 42a46570620cd9fa |
ran · our draft was wrong
fingerprinted |
MIT (permissive) |