| Focus When Necessary: Adaptive Routing and Collaborative Grounding for Training-Free Visual Grounding added by Syntology |
2026-06 (from id) |
TencentBAC/LazyMCoT/Internvl/utiles_internvl.py 3ff7a3c1eec6ef6c |
ran
|
Apache-2.0 (permissive) |
| Closed-Form Spectral Regularization for Multi-Task Model Merging added by Syntology |
2026-06 (from id) |
WalkerWorldPeace/MLLMerging/InternVL/internvl_chat/model_merging.py 6dc77bb22fb9d660 |
ran · our draft was wrong
|
no licence file found · pointer only |
| VIABLE: A Visually Impaired Assistance Benchmark for VLM-as-a-Judge Evaluation added by Syntology |
2026-05 (from id) |
YiyiyiZhao/VIABLE/viable/run_effectiveness/judge_infer_direct/utils_effectiveness/inference_internvl.py 4d0f271ceb8d4d5d |
ran
|
no licence file found · pointer only |
| When Do Diffusion Models Learn to Generate Multiple Objects? added by Syntology |
2026-05 (from id) |
eugene6923/MOSAIC/mosaic/comfort_utils/model_utils/intern_vl.py d610b5eabe0c9db1 |
ran · our draft was wrong
|
no licence file found · pointer only |
| MLLM-4D: Towards Visual-based Spatial-Temporal Intelligence added by Syntology |
2026-03 (from id) |
GVCLab/MLLM-4D/evaluation/model_inference/internvideo2_5.py d610b5eabe0c9db1 |
ran · our draft was wrong
|
no licence file found · pointer only |
| Tone Matters: The Impact of Linguistic Tone on Hallucination in VLMs added by Syntology |
2026-01 (from id) |
bli1/tone-matters/VLM_Benchmark_GitHub_Ready/benchmark/model_runners/benchmark_internvl25_freeform_with_judge.py f1f772b363bd357c |
unverified |
MIT (permissive) |
| Enhancing the Outcome Reward-based RL Training of MLLMs with Self-Consistency Sampling added by Syntology |
2025-11 (from id) |
GenuineWWD/SCS/evaluation/eval_internvl_m3cot.py 6dc77bb22fb9d660 |
ran · our draft was wrong
|
Apache-2.0 (permissive) |
| Enhancing the Outcome Reward-based RL Training of MLLMs with Self-Consistency Sampling added by Syntology |
2025-11 (from id) |
GenuineWWD/SCS/evaluation/eval_internvl_mathverse.py 23c7844d4646bf83 |
unverified |
Apache-2.0 (permissive) |
| arXiv:2507.17539 |
2025-07 (from id) |
MeteorElf/FundusExpert/src/quick_start.py 6dc77bb22fb9d660 |
ran · our draft was wrong
|
Apache-2.0 (permissive) |
| LD-RPS: Zero-Shot Unified Image Restoration via Latent Diffusion Recurrent Posterior Sampling |
1 Jul 2025 |
AMAP-ML/LD-RPS/get_prompts.py 6dc77bb22fb9d660 |
ran · our draft was wrong
|
no licence file found · pointer only |
| Draw ALL Your Imagine: A Holistic Benchmark and Agent Framework for Complex Instruction-based Image Generation |
30 May 2025 |
yczhou001/longbench-t2i/utils/evaluator.py f81449fe50195856 |
unverified |
MIT (permissive) |
| MineAnyBuild: Benchmarking Spatial Planning for Open-world AI Agents |
26 May 2025 |
identical code first harvested elsewhere 6dc77bb22fb9d660 |
ran · our draft was wrong
|
licence of this copy not recorded |
| Unifying Multimodal Large Language Model Capabilities and Modalities via Model Merging |
26 May 2025 |
identical code first harvested elsewhere 6dc77bb22fb9d660 |
ran · our draft was wrong
|
licence of this copy not recorded |
| Decoupled Visual Interpretation and Linguistic Reasoning for Math Problem Solving |
23 May 2025 |
guozix/dvlr/eval/MathVerse/evaluation/generate_response_geo_text_qwen_2.py 93f012e28ba3ca67 |
ran · our draft was wrong
|
MIT (permissive) |
| arXiv:2504.15485 |
2025-04 (from id) |
atinpothiraj/CAPTURe/occluded_scripts/intern.py 6dc77bb22fb9d660 |
ran · our draft was wrong
|
MIT (permissive) |
| Pixel-SAIL: Single Transformer For Pixel-Grounded Understanding |
14 Apr 2025 |
magic-research/Sa2VA/projects/sa2va/models/utils.py 0540d954e21b5549 |
unverified |
Apache-2.0 (permissive) |
| CoE: Chain-of-Explanation via Automatic Visual Concept Circuit Description and Polysemanticity Quantification |
19 Mar 2025 |
YuWLong666/CoE/models/internvl.py d610b5eabe0c9db1 |
ran · our draft was wrong
|
BSD-3-Clause (permissive) |
| ViSpeak: Visual Instruction Feedback in Streaming Videos |
17 Mar 2025 |
thunlp-mt/streamingbench/src/model/InternVL.py d610b5eabe0c9db1 |
ran · our draft was wrong
|
MIT (permissive) |
| CalibQuant: 1-Bit KV Cache Quantization for Multimodal LLMs |
15 Feb 2025 |
identical code first harvested elsewhere 6dc77bb22fb9d660 |
ran · our draft was wrong
|
licence of this copy not recorded |
| Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model |
30 Dec 2024 |
identical code first harvested elsewhere d610b5eabe0c9db1 |
ran · our draft was wrong
|
licence of this copy not recorded |
| SPHERE: A Hierarchical Evaluation on Spatial Perception and Reasoning for Vision-Language Models |
17 Dec 2024 |
zwenyu/SPHERE-VLM/models/vision_language_models/intern_vl2_5.py 6dc77bb22fb9d660 |
ran · our draft was wrong
|
no licence file found · pointer only |
| Holmes-VAU: Towards Long-term Video Anomaly Understanding at Any Granularity |
9 Dec 2024 |
pipixin321/holmesvau/holmesvau/internvl_utils.py 6dc77bb22fb9d660 |
ran · our draft was wrong
|
MIT (permissive) |
| R-CoT: Reverse Chain-of-Thought Problem Generation for Geometric Reasoning in Large Multimodal Models |
23 Oct 2024 |
dle666/r-cot/GeoQA_test/model_vqa_rcot2b.py 93f012e28ba3ca67 |
ran · our draft was wrong
|
no licence file found · pointer only |
| Do Vision-Language Models Represent Space and How? Evaluating Spatial Frame of Reference Under Ambiguities |
22 Oct 2024 |
sled-group/COMFORT/comfort_utils/model_utils/intern_vl.py d610b5eabe0c9db1 |
ran · our draft was wrong
|
no licence file found · pointer only |
| MultiChartQA: Benchmarking Vision-Language Models on Multi-Chart Problems |
18 Oct 2024 |
zivenzhu/multi-chart-qa/code/evaluate_internvl15.py 6dc77bb22fb9d660 |
ran · our draft was wrong
|
no licence file found · pointer only |
| MMAD: The First-Ever Comprehensive Benchmark for Multimodal Large Language Models in Industrial Anomaly Detection |
12 Oct 2024 |
jam-cc/MMAD/evaluation/examples/Transformers/internvl_query.py 6dc77bb22fb9d660 |
ran · our draft was wrong
|
no licence file found · pointer only |
| Visual Perception in Text Strings |
2 Oct 2024 |
JiaQiSJTU/VisionInText/src/evaluation_mm.py 6dc77bb22fb9d660 |
ran · our draft was wrong
|
no licence file found · pointer only |
| Med-PMC: Medical Personalized Multi-modal Consultation with a Proactive Ask-First-Observe-Next Paradigm |
16 Aug 2024 |
liuhc0428/med-pmc/src/models/InternVL.py d610b5eabe0c9db1 |
ran · our draft was wrong
|
no licence file found · pointer only |
| BIGbench: A Unified Benchmark for Evaluating Multi-dimensional Social Biases in Text-to-Image Models |
21 Jul 2024 |
bigbench2024/bigbench2024/benchmark/internViT_pkg/internvl_detection.py 4d05bcc50be7d25e |
ran
|
GPL-3.0 (copyleft) · pointer only |
| InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output |
3 Jul 2024 |
identical code first harvested elsewhere d610b5eabe0c9db1 |
ran · our draft was wrong
|
licence of this copy not recorded |
| MMLongBench-Doc: Benchmarking Long-context Document Understanding with Visualizations |
1 Jul 2024 |
mayubo2333/mmlongbench-doc/models/internvl_chat.py d610b5eabe0c9db1 |
ran · our draft was wrong
|
Apache-2.0 (permissive) |
| MMDU: A Multi-Turn Multi-Image Dialog Understanding Benchmark and Instruction-Tuning Dataset for LVLMs |
17 Jun 2024 |
liuziyu77/mmdu/model_generation/InternVL_chat_gen_ans.py d610b5eabe0c9db1 |
ran · our draft was wrong
|
Apache-2.0 (permissive) |
| Needle In A Multimodal Haystack |
11 Jun 2024 |
OpenGVLab/MM-NIAH/eval_internvl.py d610b5eabe0c9db1 |
ran · our draft was wrong
|
no licence file found · pointer only |
| Unveiling the Tapestry of Consistency in Large Vision-Language Models |
23 May 2024 |
foundation-multimodal-models/conbench/eval/InternVL-Chat-V1-5-26B.py d610b5eabe0c9db1 |
ran · our draft was wrong
|
Apache-2.0 (permissive) |
| IllusionVQA: A Challenging Optical Illusion Dataset for Vision Language Models |
23 Mar 2024 |
csebuetnlp/illusionvqa/inference_code/open_source/internvlm_inference.py 6dc77bb22fb9d660 |
ran · our draft was wrong
|
no licence file found · pointer only |
| Repaint123: Fast and High-quality One Image to 3D Generation with Progressive Controllable 2D Repainting |
20 Dec 2023 |
junwuzhang19/repaint123/main2.py d610b5eabe0c9db1 |
ran · our draft was wrong
|
MIT (permissive) |
| DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter |
2 Oct 2019 |
reycn/multi-modal-scale/script/3.evaluation.batch.py 6dc77bb22fb9d660 |
ran · our draft was wrong
|
MIT (permissive) |
| arXiv:Zhang_Holmes-VAU_Towards_Long-term_Video_Anomaly_Understanding_at_Any_Granularity_CVPR_2025_paper |
|
pipixin321/HolmesVAU/holmesvau/internvl_utils.py 6dc77bb22fb9d660 |
ran · our draft was wrong
|
MIT (permissive) |
| arXiv:Yang_PVC_Progressive_Visual_Token_Compression_for_Unified_Image_and_Video_CVPR_2025_paper |
|
OpenGVLab/PVC/utils/preprocess.py 6dc77bb22fb9d660 |
ran · our draft was wrong
|
MIT (permissive) |
| arXiv:Guo_Integrating_Visual_Interpretation_and_Linguistic_Reasoning_for_Geometric_Problem_Solving_ICCV_2025_paper |
|
guozix/DVLR/eval/MathVerse/evaluation/generate_response_geo_text_qwen_2.py 93f012e28ba3ca67 |
ran · our draft was wrong
|
MIT (permissive) |
| arXiv:2025.acl-long.663 |
|
BlueZeros/ReflecTool/reflectool/models/InternVLChat.py 6dc77bb22fb9d660 |
ran · our draft was wrong
|
Apache-2.0 (permissive) |