| Focus When Necessary: Adaptive Routing and Collaborative Grounding for Training-Free Visual Grounding added by Syntology |
2026-06 (from id) |
TencentBAC/LazyMCoT/Internvl/utiles_internvl.py bd77f5f8067f18e9 |
ran · fixture could not drive it
|
Apache-2.0 (permissive) |
| Closed-Form Spectral Regularization for Multi-Task Model Merging added by Syntology |
2026-06 (from id) |
WalkerWorldPeace/MLLMerging/InternVL/internvl_chat/model_merging.py bd77f5f8067f18e9 |
ran · fixture could not drive it
|
no licence file found · pointer only |
| VIABLE: A Visually Impaired Assistance Benchmark for VLM-as-a-Judge Evaluation added by Syntology |
2026-05 (from id) |
YiyiyiZhao/VIABLE/viable/run_effectiveness/judge_infer_direct/utils_effectiveness/inference_internvl.py bd77f5f8067f18e9 |
ran · fixture could not drive it
|
no licence file found · pointer only |
| When Do Diffusion Models Learn to Generate Multiple Objects? added by Syntology |
2026-05 (from id) |
eugene6923/MOSAIC/mosaic/comfort_utils/model_utils/intern_vl.py bd77f5f8067f18e9 |
ran · fixture could not drive it
|
no licence file found · pointer only |
| MLLM-4D: Towards Visual-based Spatial-Temporal Intelligence added by Syntology |
2026-03 (from id) |
GVCLab/MLLM-4D/evaluation/model_inference/internvideo2_5.py bd77f5f8067f18e9 |
ran · fixture could not drive it
|
no licence file found · pointer only |
| Enhancing the Outcome Reward-based RL Training of MLLMs with Self-Consistency Sampling added by Syntology |
2025-11 (from id) |
GenuineWWD/SCS/evaluation/eval_internvl_m3cot.py bd77f5f8067f18e9 |
ran · fixture could not drive it
|
Apache-2.0 (permissive) |
| LD-RPS: Zero-Shot Unified Image Restoration via Latent Diffusion Recurrent Posterior Sampling |
1 Jul 2025 |
AMAP-ML/LD-RPS/get_prompts.py bd77f5f8067f18e9 |
ran · fixture could not drive it
|
no licence file found · pointer only |
| MineAnyBuild: Benchmarking Spatial Planning for Open-world AI Agents |
26 May 2025 |
identical code first harvested elsewhere bd77f5f8067f18e9 |
ran · fixture could not drive it
|
licence of this copy not recorded |
| Unifying Multimodal Large Language Model Capabilities and Modalities via Model Merging |
26 May 2025 |
identical code first harvested elsewhere bd77f5f8067f18e9 |
ran · fixture could not drive it
|
licence of this copy not recorded |
| Decoupled Visual Interpretation and Linguistic Reasoning for Math Problem Solving |
23 May 2025 |
identical code first harvested elsewhere bd77f5f8067f18e9 |
ran · fixture could not drive it
|
licence of this copy not recorded |
| arXiv:2504.15485 |
2025-04 (from id) |
atinpothiraj/CAPTURe/occluded_scripts/intern.py bd77f5f8067f18e9 |
ran · fixture could not drive it
|
MIT (permissive) |
| Pixel-SAIL: Single Transformer For Pixel-Grounded Understanding |
14 Apr 2025 |
magic-research/Sa2VA/projects/sa2va/models/utils.py bd77f5f8067f18e9 |
ran · fixture could not drive it
|
Apache-2.0 (permissive) |
| CoE: Chain-of-Explanation via Automatic Visual Concept Circuit Description and Polysemanticity Quantification |
19 Mar 2025 |
YuWLong666/CoE/models/internvl.py bd77f5f8067f18e9 |
ran · fixture could not drive it
|
BSD-3-Clause (permissive) |
| ViSpeak: Visual Instruction Feedback in Streaming Videos |
17 Mar 2025 |
thunlp-mt/streamingbench/src/model/InternVL.py bd77f5f8067f18e9 |
ran · fixture could not drive it
|
MIT (permissive) |
| Forgotten Polygons: Multimodal Large Language Models are Shape-Blind |
21 Feb 2025 |
identical code first harvested elsewhere bd77f5f8067f18e9 |
ran · fixture could not drive it
|
licence of this copy not recorded |
| CalibQuant: 1-Bit KV Cache Quantization for Multimodal LLMs |
15 Feb 2025 |
identical code first harvested elsewhere bd77f5f8067f18e9 |
ran · fixture could not drive it
|
licence of this copy not recorded |
| OVO-Bench: How Far is Your Video-LLMs from Real-World Online Video Understanding? |
9 Jan 2025 |
JoeLeelyf/OVO-Bench/models/InternVL2.py bd77f5f8067f18e9 |
ran · fixture could not drive it
|
MIT (permissive) |
| Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model |
30 Dec 2024 |
identical code first harvested elsewhere bd77f5f8067f18e9 |
ran · fixture could not drive it
|
licence of this copy not recorded |
| SPHERE: A Hierarchical Evaluation on Spatial Perception and Reasoning for Vision-Language Models |
17 Dec 2024 |
zwenyu/SPHERE-VLM/models/vision_language_models/intern_vl2_5.py bd77f5f8067f18e9 |
ran · fixture could not drive it
|
no licence file found · pointer only |
| Holmes-VAU: Towards Long-term Video Anomaly Understanding at Any Granularity |
9 Dec 2024 |
pipixin321/holmesvau/holmesvau/internvl_utils.py bd77f5f8067f18e9 |
ran · fixture could not drive it
|
MIT (permissive) |
| LinVT: Empower Your Image-level Large Language Model to Understand Videos |
6 Dec 2024 |
gls0425/linvt/streamlit_demo/model_worker.py bd77f5f8067f18e9 |
ran · fixture could not drive it
|
no licence file found · pointer only |
| Safe + Safe = Unsafe? Exploring How Safe Images Can Be Exploited to Jailbreak Large Vision-Language Models |
18 Nov 2024 |
gzcch/safety_snowball_agent/InternVL_assitant.py bd77f5f8067f18e9 |
ran · fixture could not drive it
|
no licence file found · pointer only |
| Both Text and Images Leaked! A Systematic Analysis of Multimodal LLM Data Contamination |
6 Nov 2024 |
MLLM-Data-Contamination/MM-Detect/mm_detect/mllms/internvl2.py bd77f5f8067f18e9 |
ran · fixture could not drive it
|
Apache-2.0 (permissive) |
| R-CoT: Reverse Chain-of-Thought Problem Generation for Geometric Reasoning in Large Multimodal Models |
23 Oct 2024 |
dle666/r-cot/GeoQA_test/model_vqa_rcot2b.py bd77f5f8067f18e9 |
ran · fixture could not drive it
|
no licence file found · pointer only |
| Do Vision-Language Models Represent Space and How? Evaluating Spatial Frame of Reference Under Ambiguities |
22 Oct 2024 |
sled-group/COMFORT/comfort_utils/model_utils/intern_vl.py bd77f5f8067f18e9 |
ran · fixture could not drive it
|
no licence file found · pointer only |
| MultiChartQA: Benchmarking Vision-Language Models on Multi-Chart Problems |
18 Oct 2024 |
zivenzhu/multi-chart-qa/code/evaluate_internvl15.py bd77f5f8067f18e9 |
ran · fixture could not drive it
|
no licence file found · pointer only |
| MMAD: The First-Ever Comprehensive Benchmark for Multimodal Large Language Models in Industrial Anomaly Detection |
12 Oct 2024 |
identical code first harvested elsewhere bd77f5f8067f18e9 |
ran · fixture could not drive it
|
licence of this copy not recorded |
| Visual Perception in Text Strings |
2 Oct 2024 |
JiaQiSJTU/VisionInText/src/evaluation_mm.py bd77f5f8067f18e9 |
ran · fixture could not drive it
|
no licence file found · pointer only |
| Med-PMC: Medical Personalized Multi-modal Consultation with a Proactive Ask-First-Observe-Next Paradigm |
16 Aug 2024 |
liuhc0428/med-pmc/src/models/InternVL.py bd77f5f8067f18e9 |
ran · fixture could not drive it
|
no licence file found · pointer only |
| BIGbench: A Unified Benchmark for Evaluating Multi-dimensional Social Biases in Text-to-Image Models |
21 Jul 2024 |
bigbench2024/bigbench2024/benchmark/internViT_pkg/internvl_detection.py d0487594e4d5c510 |
ran
|
GPL-3.0 (copyleft) · pointer only |
| MMSci: A Dataset for Graduate-Level Multi-Discipline Multimodal Scientific Understanding |
6 Jul 2024 |
leezekun/mmsci/mmsci-exps/model_loader.py bd77f5f8067f18e9 |
ran · fixture could not drive it
|
no licence file found · pointer only |
| InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output |
3 Jul 2024 |
identical code first harvested elsewhere bd77f5f8067f18e9 |
ran · fixture could not drive it
|
licence of this copy not recorded |
| MMLongBench-Doc: Benchmarking Long-context Document Understanding with Visualizations |
1 Jul 2024 |
mayubo2333/mmlongbench-doc/models/internvl_chat.py bd77f5f8067f18e9 |
ran · fixture could not drive it
|
Apache-2.0 (permissive) |
| MMDU: A Multi-Turn Multi-Image Dialog Understanding Benchmark and Instruction-Tuning Dataset for LVLMs |
17 Jun 2024 |
liuziyu77/mmdu/model_generation/InternVL_chat_gen_ans.py bd77f5f8067f18e9 |
ran · fixture could not drive it
|
Apache-2.0 (permissive) |
| DataComp-LM: In search of the next generation of training sets for language models |
17 Jun 2024 |
jieyuz2/taskmeanything/tma/models/qa_model/imageqa_model.py bd77f5f8067f18e9 |
ran · fixture could not drive it
|
Apache-2.0 (permissive) |
| Needle In A Multimodal Haystack |
11 Jun 2024 |
OpenGVLab/MM-NIAH/eval_internvl.py bd77f5f8067f18e9 |
ran · fixture could not drive it
|
no licence file found · pointer only |
| VCR: A Task for Pixel-Level Complex Reasoning in Vision Language Models via Restoring Occluded Text |
10 Jun 2024 |
tianyu-z/vcr/src/evaluation/utils.py bd77f5f8067f18e9 |
ran · fixture could not drive it
|
CC-BY-SA-4.0 (copyleft) · pointer only |
| Unveiling the Tapestry of Consistency in Large Vision-Language Models |
23 May 2024 |
foundation-multimodal-models/conbench/eval/InternVL-Chat-V1-5-26B.py bd77f5f8067f18e9 |
ran · fixture could not drive it
|
Apache-2.0 (permissive) |
| IllusionVQA: A Challenging Optical Illusion Dataset for Vision Language Models |
23 Mar 2024 |
csebuetnlp/illusionvqa/inference_code/open_source/internvlm_inference.py bd77f5f8067f18e9 |
ran · fixture could not drive it
|
no licence file found · pointer only |
| Repaint123: Fast and High-quality One Image to 3D Generation with Progressive Controllable 2D Repainting |
20 Dec 2023 |
junwuzhang19/repaint123/main2.py bd77f5f8067f18e9 |
ran · fixture could not drive it
|
MIT (permissive) |
| DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter |
2 Oct 2019 |
reycn/multi-modal-scale/script/3.evaluation.batch.py bd77f5f8067f18e9 |
ran · fixture could not drive it
|
MIT (permissive) |
| arXiv:Zhang_Holmes-VAU_Towards_Long-term_Video_Anomaly_Understanding_at_Any_Granularity_CVPR_2025_paper |
|
pipixin321/HolmesVAU/holmesvau/internvl_utils.py bd77f5f8067f18e9 |
ran · fixture could not drive it
|
MIT (permissive) |
| arXiv:Yang_PVC_Progressive_Visual_Token_Compression_for_Unified_Image_and_Video_CVPR_2025_paper |
|
OpenGVLab/PVC/utils/preprocess.py bd77f5f8067f18e9 |
ran · fixture could not drive it
|
MIT (permissive) |
| arXiv:Guo_Integrating_Visual_Interpretation_and_Linguistic_Reasoning_for_Geometric_Problem_Solving_ICCV_2025_paper |
|
guozix/DVLR/eval/MathVerse/evaluation/generate_response_geo_text_qwen_2.py bd77f5f8067f18e9 |
ran · fixture could not drive it
|
MIT (permissive) |
| arXiv:2025.acl-long.663 |
|
BlueZeros/ReflecTool/reflectool/models/InternVLChat.py bd77f5f8067f18e9 |
ran · fixture could not drive it
|
Apache-2.0 (permissive) |