| Do Video Generators Track the World Across Segments? A Benchmark and Method for World-State Reasoning in Video Continuation added by Syntology |
2026-09 (from id) |
AMAP-ML/StateAgent/stateagent/utils.py d9a8f3eeba92ed3f |
unverified |
MIT (permissive) |
| WIDER-FAIR: An Annotated Version of the WIDER-FACE Dataset for Fairness Evaluation added by Syntology |
2026-06 (from id) |
bronval/Wider-Fair-Dataset/paper_experiments/Annotator/ui.py 703a6b572693d8f8 |
ran
fingerprinted |
no licence file found · pointer only |
| Unlimited OCR Works Welcome the Era of One-shot Long-horizon Parsing added by Syntology |
2026-06 (from id) |
baidu/Unlimited-OCR/infer.py 85dc65aa2564f6c1 |
unverified |
MIT (permissive) |
| BLUEX v2: Benchmarking LLMs on Open-Ended Questions from Brazilian University Entrance Exams added by Syntology |
2026-06 (from id) |
TropicAI-Research/BLUEXv2/dataset_pipeline/generate_captions.py 7628c496f9145193 |
unverified |
no licence file found · pointer only |
| KCSAT-ML: Probing Reasoning Models with Nationwide-Cohort Human Difficulty added by Syntology |
2026-06 (from id) |
naver-ai/KCSAT-ML/src/utils/image_utils.py 0cf01235db87302e |
ran
|
AGPL-3.0 (copyleft) · pointer only |
| PACT: Self-Evolving Physical Safety Alignment for Diffusion Policies in Embodied Manipulation added by Syntology |
2026-06 (from id) |
robotwin-Platform/RoboTwin/code_gen/observation_agent.py 73c070397fc27811 |
ran
|
MIT (permissive) |
| CRAFTER: A Multi-Agent Harness for Editable Scientific Figure Generation from Diverse Inputs added by Syntology |
2026-05 (from id) |
HaozheZhao/Crafter/crafter/editor/raster_to_svg/model_router.py 5639ab5d5241466a |
ran
|
MIT (permissive) |
| WindowsWorld: A Process-Centric Benchmark of Autonomous GUI Agents in Professional Cross-Application Environments added by Syntology |
2026-04 (from id) |
HITsz-TMG/WindowsWorld/mm_agents/agent.py 78ee6b46cebd99ee |
ran · our draft was wrong
|
Apache-2.0 (permissive) |
| More Than Sum of Its Parts: Deciphering Intent Shifts in Multimodal Hate Speech Detection added by Syntology |
2026-03 (from id) |
Sayur1n/H-VLI/utils.py 74b640f100721700 |
unverified |
licence not identified · pointer only |
| Mobile-Agent-v3.5: Multi-platform Fundamental GUI Agents added by Syntology |
2026-02 (from id) |
X-PLUG/MobileAgent/Mobile-Agent-v1/MobileAgent/api.py f41cb1a19b154297 |
ran · our draft was wrong
|
MIT (permissive) |
| Natural Language Instructions for Scene-Responsive Human-in-the-Loop Motion Planning in Autonomous Driving using Vision-Language-Action Models added by Syntology |
2026-02 (from id) |
Mi3-Lab/doScenes-VLM-Planning/src/utils.py f41cb1a19b154297 |
ran · our draft was wrong
|
AGPL-3.0 (copyleft) · pointer only |
| Brazilian Portuguese Image Captioning with Transformers: A Study on Cross-Native-Translated Dataset added by Syntology |
2026-02 (from id) |
laicsiifes/transformer-caption-ptbr/vlm_zero_shot/src/inference_sambanova_openai.py 9303356715a06b3c |
unverified |
no licence file found · pointer only |
| Q-Hawkeye: Reliable Visual Policy Optimization for Image Quality Assessment added by Syntology |
30 Jan 2026 |
AMAP-ML/Q-Hawkeye/src/Dataset/Degradation_Dataset/VLM_filter.py 587a00453d7c872e |
unverified |
no licence file found · pointer only |
| UniFinEval: Towards Unified Evaluation of Financial Multimodal Models across Text, Images and Videos added by Syntology |
2026-01 (from id) |
aifinlab/UniFinEval/evaluate_py/model_api.py 625809a36981482c |
unverified |
Apache-2.0 (permissive) |
| TEMPVIZ: On the Evaluation of Temporal Knowledge in Text-to-Image Models added by Syntology |
2026-01 (from id) |
TAI-HAMBURG/TempViz/code/get_answers_openai.py f41cb1a19b154297 |
ran · our draft was wrong
|
no licence file found · pointer only |
| INFINITEWEB: Scalable Web Environment Synthesis for GUI Agent Training added by Syntology |
2026-01 (from id) |
microsoft/FIVE-UI-Evol/InfiniteWeb/src/llm_caller.py d6d6f575b423041b |
unverified |
MIT (permissive) |
| RAG-IGBench: Innovative Evaluation for RAG-based Interleaved Generation in Open-domain Question Answering added by Syntology |
2025-12 (from id) |
USTC-StarTeam/RAG-IGBench/model_generation/claude.py f41cb1a19b154297 |
ran · our draft was wrong
|
Apache-2.0 (permissive) |
| ImageSentinel: Protecting Visual Datasets from Unauthorized Retrieval-Augmented Image Generation added by Syntology |
2025-10 (from id) |
luo-ziyuan/ImageSentinel/ImageSentinel/utils.py f41cb1a19b154297 |
ran · our draft was wrong
|
no licence file found · pointer only |
| GUI-Spotlight: Adaptive Iterative Focus Refinement for Enhanced GUI Visual Grounding added by Syntology |
5 Oct 2025 |
bin123apple/GUI_Spotlight/screenspot_pro_evaluation.py 78ee6b46cebd99ee |
ran · our draft was wrong
|
MIT (permissive) |
| GenExam: A Multidisciplinary Text-to-Image Exam added by Syntology |
2025-09 (from id) |
OpenGVLab/GenExam/run_eval.py c780d21931c485a1 |
unverified |
MIT (permissive) |
| Effective Training Data Synthesis for Improving MLLM Chart Understanding added by Syntology |
2025-08 (from id) |
yuweiyang-anu/ECD/data_generation_pipeline/chart_image_filtering.py 852ef01643a7f02e |
unverified |
MIT (permissive) |
| arXiv:2507.19969 |
2025-07 (from id) |
vis-nlp/Text2Vis/eval_predictions.py 6ad02ea0ba44bcaf |
unverified |
GPL-3.0 (copyleft) · pointer only |
| arXiv:2507.17539 |
2025-07 (from id) |
MeteorElf/FundusExpert/src/eval/eval_api/call_api.py 6ad02ea0ba44bcaf |
unverified |
Apache-2.0 (permissive) |
| GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning |
1 Jul 2025 |
thudm/glm-4.1v-thinking/glmv_reward/src/glmv_reward/utils/image.py 107e2008fcd35451 |
unverified |
Apache-2.0 (permissive) |
| MIRAGE: A Benchmark for Multimodal Information-Seeking and Reasoning in Agricultural Expert-Guided Conversations |
25 Jun 2025 |
mirage-benchmark/mirage-benchmark/MMST/chat_models/Client.py f41cb1a19b154297 |
ran · our draft was wrong
|
no licence file found · pointer only |
| Detecting Harmful Memes with Decoupled Understanding and Guided CoT Reasoning |
10 Jun 2025 |
panFJCharlotte98/HMC/call_gpt.py f41cb1a19b154297 |
ran · our draft was wrong
|
MIT (permissive) |
| RedTeamCUA: Realistic Adversarial Testing of Computer-Use Agents in Hybrid Web-OS Environments |
28 May 2025 |
osu-nlp-group/redteamcua/mm_agents/agent.py 78ee6b46cebd99ee |
ran · our draft was wrong
|
Apache-2.0 (permissive) |
| VisCRA: A Visual Chain Reasoning Attack for Jailbreaking Multimodal Large Language Models |
26 May 2025 |
DyMessi/VisCRA/evaluation/gemini.py f41cb1a19b154297 |
ran · our draft was wrong
|
Apache-2.0 (permissive) |
| VisCRA: A Visual Chain Reasoning Attack for Jailbreaking Multimodal Large Language Models |
26 May 2025 |
DyMessi/VisCRA/evaluation/QvQ_Max.py cfe427db2f2f5fcc |
unverified |
Apache-2.0 (permissive) |
| VTool-R1: VLMs Learn to Think with Images via Reinforcement Learning on Multimodal Tool Use |
25 May 2025 |
VTOOL-R1/vtool-r1/eval/eval_gpt_no_tool.py 1104c28afa7dd8d2 |
unverified |
Apache-2.0 (permissive) |
| VTool-R1: VLMs Learn to Think with Images via Reinforcement Learning on Multimodal Tool Use |
25 May 2025 |
VTOOL-R1/vtool-r1/eval/eval_gpt_tableqa_with_tool.py 67161499de84ab96 |
unverified |
Apache-2.0 (permissive) |
| Can Multimodal Large Language Models Understand Spatial Relations? |
25 May 2025 |
ziyan-xiaoyu/spatialmqa/Code/close_models/gpt4_1_shot.py b60704e5254f6958 |
unverified |
Apache-2.0 (permissive) |
| DanmakuTPPBench: A Multi-modal Benchmark for Temporal Point Process Modeling and Understanding |
23 May 2025 |
identical code first harvested elsewhere f41cb1a19b154297 |
ran · our draft was wrong
|
licence of this copy not recorded |
| ComfyMind: Toward General-Purpose Generation via Tree-Based Planning and Reactive Feedback |
23 May 2025 |
EnVision-Research/ComfyMind/script/evaluation.py 0ea99cab0cdb51ab |
ran · our draft was wrong
|
MIT (permissive) |
| ComfyMind: Toward General-Purpose Generation via Tree-Based Planning and Reactive Feedback |
23 May 2025 |
EnVision-Research/ComfyMind/script/evaluation_wise.py 393004fc099595af |
unverified |
MIT (permissive) |
| DriveAgent: Multi-Agent Structured Reasoning with LLM and Multimodal Sensor Fusion for Autonomous Driving |
2025-05 (from id) |
paparare/driveagent/enviroment.py 735a954e94f37a3f |
unverified |
MIT (permissive) |
| arXiv:2504.15485 |
2025-04 (from id) |
atinpothiraj/CAPTURe/occluded_scripts/gpt.py f41cb1a19b154297 |
ran · our draft was wrong
|
MIT (permissive) |
| Why We Feel: Breaking Boundaries in Emotional Reasoning with Multimodal Large Language Models |
10 Apr 2025 |
lum1104/eibench/EIBench/baselines/ChatGPT-4/gpt4-score-complex.py f41cb1a19b154297 |
ran · our draft was wrong
|
Apache-2.0 (permissive) |
| An Illusion of Progress? Assessing the Current State of Web Agents |
2 Apr 2025 |
osu-nlp-group/online-mind2web/src/utils.py 6bffa3dffc4ecf77 |
unverified |
MIT (permissive) |
| CoE: Chain-of-Explanation via Automatic Visual Concept Circuit Description and Polysemanticity Quantification |
19 Mar 2025 |
YuWLong666/CoE/closeai.py f41cb1a19b154297 |
ran · our draft was wrong
|
BSD-3-Clause (permissive) |
| MDocAgent: A Multi-Modal Multi-Agent Framework for Document Understanding |
18 Mar 2025 |
aiming-lab/mdocagent/models/openai.py f41cb1a19b154297 |
ran · our draft was wrong
|
MIT (permissive) |
| CoLMDriver: LLM-based Negotiation Benefits Cooperative Autonomous Driving |
11 Mar 2025 |
cxliu0314/CoLMDriver/simulation/leaderboard/team_code/colmdriver_action.py f41cb1a19b154297 |
ran · our draft was wrong
|
Apache-2.0 (permissive) |
| VEM: Environment-Free Exploration for Training GUI Agent with Value Environment Model |
26 Feb 2025 |
microsoft/gui-agent-rl/data_preprocess/gpt.py 7802770e6818ba35 |
unverified |
MIT (permissive) |
| ArtMentor: AI-Assisted Evaluation of Artworks to Explore Multimodal Large Language Models Capabilities |
2025-02 (from id) |
identical code first harvested elsewhere f41cb1a19b154297 |
ran · our draft was wrong
|
licence of this copy not recorded |
| Manual2Skill: Learning to Read Manuals and Acquire Robotic Skills for Furniture Assembly Using Vision-Language Models |
14 Feb 2025 |
owensun2004/Manual2Skill/VLM_assembly_plan_gen/inference/utils.py c15f0525475c47c9 |
unverified |
Apache-2.0 (permissive) |
| Kimi k1.5: Scaling Reinforcement Learning with LLMs |
22 Jan 2025 |
mathllm/math-v/models/GPT4V.py a7b8d145010c3bfa |
unverified |
MIT (permissive) |
| Kimi k1.5: Scaling Reinforcement Learning with LLMs |
22 Jan 2025 |
mathllm/math-v/models/GPT_with_caption.py 7f16ba077ded04a3 |
unverified |
MIT (permissive) |
| EmbodiedEval: Evaluate Multimodal LLMs as Embodied Agents |
21 Jan 2025 |
thunlp/embodiedeval/agent.py b72ccf9df612250c |
unverified |
MIT (permissive) |
| AutoPresent: Designing Structured Visuals from Scratch |
1 Jan 2025 |
para-lost/AutoPresent/evaluate/reference_free_eval.py f41cb1a19b154297 |
ran · our draft was wrong
|
MIT (permissive) |
| AutoTrust: Benchmarking Trustworthiness in Large Vision Language Models for Autonomous Driving |
19 Dec 2024 |
taco-group/autotrust/utils.py f41cb1a19b154297 |
ran · our draft was wrong
|
Apache-2.0 (permissive) |
| SPHERE: A Hierarchical Evaluation on Spatial Perception and Reasoning for Vision-Language Models |
17 Dec 2024 |
zwenyu/SPHERE-VLM/models/vision_language_models/gpt.py f41cb1a19b154297 |
ran · our draft was wrong
|
no licence file found · pointer only |
| TeamCraft: A Benchmark for Multi-Modal Multi-Agent Systems in Minecraft |
6 Dec 2024 |
teamcraft-bench/teamcraft/teamcraft/openai_api.py 65bac3f0167d9a59 |
unverified |
MIT (permissive) |
| VLSBench: Unveiling Visual Leakage in Multimodal Safety |
29 Nov 2024 |
ai45lab/vlsbench/eval_utils.py f41cb1a19b154297 |
ran · our draft was wrong
|
no licence file found · pointer only |
| MovieBench: A Hierarchical Movie Level Dataset for Long Video Generation |
22 Nov 2024 |
showlab/moviebecnh/MovieBench/utils.py f41cb1a19b154297 |
ran · our draft was wrong
|
no licence file found · pointer only |
| The Dawn of GUI Agent: A Preliminary Case Study with Claude 3.5 Computer Use |
15 Nov 2024 |
showlab/computer_use_ootb/computer_use_demo/gui_agent/llm_utils/llm_utils.py fef879c33995e69e |
unverified |
Apache-2.0 (permissive) |
| The Dawn of GUI Agent: A Preliminary Case Study with Claude 3.5 Computer Use |
15 Nov 2024 |
showlab/computer_use_ootb/computer_use_demo/gui_agent/llm_utils/qwen.py f6a0e025afd2766e |
unverified |
Apache-2.0 (permissive) |
| Is Your LLM Secretly a World Model of the Internet? Model-Based Planning for Web Agents |
10 Nov 2024 |
identical code first harvested elsewhere f41cb1a19b154297 |
ran · our draft was wrong
|
licence of this copy not recorded |
| HourVideo: 1-Hour Video-Language Understanding |
7 Nov 2024 |
keshik6/HourVideo/hourvideo/gpt4_utils.py f41cb1a19b154297 |
ran · our draft was wrong
|
Apache-2.0 (permissive) |
| Both Text and Images Leaked! A Systematic Analysis of Multimodal LLM Data Contamination |
6 Nov 2024 |
MLLM-Data-Contamination/MM-Detect/mm_detect/mllms/gpt.py e7b40194e0800766 |
unverified |
Apache-2.0 (permissive) |
| Constrained Human-AI Cooperation: An Inclusive Embodied Social Intelligence Challenge |
4 Nov 2024 |
UMass-Embodied-AGI/CHAIC/LM_agent/VLM.py f41cb1a19b154297 |
ran · our draft was wrong
|
no licence file found · pointer only |
| TurtleBench: A Visual Programming Benchmark in Turtle Geometry |
31 Oct 2024 |
sinaris76/turtlebench/models/gpt.py f41cb1a19b154297 |
ran · our draft was wrong
|
no licence file found · pointer only |
| AndroidLab: Training and Systematic Benchmarking of Android Autonomous Agents |
31 Oct 2024 |
THUDM/Android-Lab/agent/utils.py f41cb1a19b154297 |
ran · our draft was wrong
|
MIT (permissive) |
| AndroidLab: Training and Systematic Benchmarking of Android Autonomous Agents |
31 Oct 2024 |
THUDM/Android-Lab/evaluation/definition.py 392fd89baad47e29 |
unverified |
MIT (permissive) |
| Implementation and Application of an Intelligibility Protocol for Interaction with an LLM |
27 Oct 2024 |
karannb/interact/src/utils.py d193df7c20c2973f |
unverified |
MIT (permissive) |
| OpenWebVoyager: Building Multimodal Web Agents via Iterative Real-World Exploration, Feedback and Optimization |
25 Oct 2024 |
minorjerry/openwebvoyager/WebVoyager/utils.py f41cb1a19b154297 |
ran · our draft was wrong
|
Apache-2.0 (permissive) |
| UGotMe: An Embodied System for Affective Human-Robot Interaction |
2024-10 (from id) |
lipzh5/amecavle/models/emotion_rec.py f41cb1a19b154297 |
ran · our draft was wrong
|
no licence file found · pointer only |
| Chain of Ideas: Revolutionizing Research Via Novel Idea Development with LLM Agents |
17 Oct 2024 |
damo-nlp-sg/coi-agent/LLM.py f41cb1a19b154297 |
ran · our draft was wrong
|
Apache-2.0 (permissive) |
| VidEgoThink: Assessing Egocentric Video Understanding Capabilities for Embodied AI |
15 Oct 2024 |
adacheng/egothink/gpt_eval.py f41cb1a19b154297 |
ran · our draft was wrong
|
Apache-2.0 (permissive) |
| Towards Foundation Models for 3D Vision: How Close Are We? |
14 Oct 2024 |
princeton-vl/uniqa-3d/LLM_evaluations/clevr_vqa/generate_gpt_response.py f41cb1a19b154297 |
ran · our draft was wrong
|
BSD-3-Clause (permissive) |
| LiveXiv -- A Multi-Modal Live Benchmark Based on Arxiv Papers Content |
14 Oct 2024 |
nimrodshabtay/livexiv/vqa_generation/model_utils/claude_utils.py beb4056b7230f14f |
ran
|
Apache-2.0 (permissive) |
| LiveXiv -- A Multi-Modal Live Benchmark Based on Arxiv Papers Content |
14 Oct 2024 |
nimrodshabtay/livexiv/vqa_generation/model_utils/gpt_utils.py 826cad294a32186a |
ran
|
Apache-2.0 (permissive) |
| VideoAgent: Self-Improving Video Generation |
14 Oct 2024 |
video-as-agent/videoagent/flowdiffusion/feedback_binary_rf.py f41cb1a19b154297 |
ran · our draft was wrong
|
MIT (permissive) |
| VLFeedback: A Large-Scale AI Feedback Dataset for Large Vision-Language Models Alignment |
12 Oct 2024 |
yuweihao/MM-Vet/inference/utils.py f41cb1a19b154297 |
ran · our draft was wrong
|
Apache-2.0 (permissive) |
| ETA: Evaluating Then Aligning Safety of Vision Language Models at Inference Time |
9 Oct 2024 |
dripnowhy/eta/helpfulscore_ai.py f41cb1a19b154297 |
ran · our draft was wrong
|
no licence file found · pointer only |
| Unlocking the Capabilities of Thought: A Reasoning Boundary Framework to Quantify and Optimize Chain-of-Thought |
8 Oct 2024 |
LightChen233/reasoning-boundary/request_multimodal.py f41cb1a19b154297 |
ran · our draft was wrong
|
no licence file found · pointer only |
| Towards World Simulator: Crafting Physical Commonsense-Based Benchmark for Video Generation |
7 Oct 2024 |
opengvlab/phygenbench/PhyGenEval/multi/GPT4o.py f41cb1a19b154297 |
ran · our draft was wrong
|
no licence file found · pointer only |
| Visual Perception in Text Strings |
2 Oct 2024 |
JiaQiSJTU/VisionInText/src/utils/data_utils.py f41cb1a19b154297 |
ran · our draft was wrong
|
no licence file found · pointer only |
| MECD: Unlocking Multi-Event Causal Discovery in Video Reasoning |
26 Sep 2024 |
tychen-SJTU/MECD-Benchmark/mecd_llm_fewshot/gpt4o.py 6163d22216816e96 |
ran
|
MIT (permissive) |
| MediConfusion: Can you trust your AI radiologist? Probing the reliability of multimodal medical foundation models |
23 Sep 2024 |
AIF4S/MediConfusion/Models/gpt.py f41cb1a19b154297 |
ran · our draft was wrong
|
no licence file found · pointer only |
| LIME: Less Is More for MLLM Evaluation |
10 Sep 2024 |
kangreen0210/lime/data_curation_pipeline/gpt.py f41cb1a19b154297 |
ran · our draft was wrong
|
no licence file found · pointer only |
| CMM-Math: A Chinese Multimodal Math Dataset To Evaluate and Enhance the Mathematics Reasoning of Large Multimodal Models |
4 Sep 2024 |
ecnu-icalk/educhat-math/model/answer_in_testdata/GPT4o-shot.py a7b8d145010c3bfa |
unverified |
no licence file found · pointer only |
| CMM-Math: A Chinese Multimodal Math Dataset To Evaluate and Enhance the Mathematics Reasoning of Large Multimodal Models |
4 Sep 2024 |
ecnu-icalk/educhat-math/model/answer_in_testdata/GPT4o.py dcdb5c80a7b5cffa |
unverified |
no licence file found · pointer only |
| ComfyBench: Benchmarking LLM-based Agents in ComfyUI for Autonomously Designing Collaborative AI Systems |
2 Sep 2024 |
xxyQwQ/ComfyBench/script/evaluation.py 0ea99cab0cdb51ab |
ran · our draft was wrong
|
no licence file found · pointer only |
| CogVLM2: Visual Language Models for Image and Video Understanding |
29 Aug 2024 |
thudm/cogvlm2/basic_demo/openai_api_request.py c0b1c7a482944767 |
ran
|
Apache-2.0 (permissive) |
| Measuring Agreeableness Bias in Multimodal Models |
17 Aug 2024 |
jasonlim131/looksRdeceiving/python-src/evaluate_gpt4_v1.py 2d5f690652247562 |
ran · our draft was wrong
|
MIT (permissive) |
| Imagen 3 |
13 Aug 2024 |
linzhiqiu/t2i_metrics/t2v_metrics/models/vqascore_models/gpt4v_model.py f41cb1a19b154297 |
ran · our draft was wrong
|
Apache-2.0 (permissive) |
| Imagen 3 |
13 Aug 2024 |
linzhiqiu/t2i_metrics/t2v_metrics/models/vqascore_models/gemini_model.py 2c2b1eb6c4d90f6c |
ran
|
Apache-2.0 (permissive) |
| EARBench: Towards Evaluating Physical Risk Awareness for Task Planning of Foundation Model-based Embodied AI Agents |
8 Aug 2024 |
zihao-ai/eairiskbench/image_judger.py f41cb1a19b154297 |
ran · our draft was wrong
|
MIT (permissive) |
| Learning Video Context as Interleaved Multimodal Sequences |
31 Jul 2024 |
showlab/movieseq/utils.py f41cb1a19b154297 |
ran · our draft was wrong
|
no licence file found · pointer only |
| AMONGAGENTS: Evaluating Large Language Models in the Interactive Text-Based Social Deduction Game |
23 Jul 2024 |
cyzus/among-agents/amongagents/evaluation/evaluate.py f41cb1a19b154297 |
ran · our draft was wrong
|
MIT (permissive) |
| CoCoG-2: Controllable generation of visual stimuli for understanding human concept representation |
20 Jul 2024 |
ncclab-sustech/cocog-2/customized_pipe.py 62f91e0069da2565 |
unverified |
no licence file found · pointer only |
| Spider2-V: How Far Are Multimodal Agents From Automating Data Science and Engineering Workflows? |
15 Jul 2024 |
xlang-ai/spider2-v/mm_agents/agent.py 78ee6b46cebd99ee |
ran · our draft was wrong
|
Apache-2.0 (permissive) |
| LLaMAR: Long-Horizon Planning for Multi-Agent Robots in Partially Observable Environments |
2024-07 (from id) |
nsidn98/llamar/SAR/baselines/llamar_utils_multiagent.py 292346e4cf05533b |
ran
|
MIT (permissive) |
| SEED-Story: Multimodal Long Story Generation with Large Language Model |
11 Jul 2024 |
tencentarc/seed-story/StoryStream/build_story_v2.py f41cb1a19b154297 |
ran · our draft was wrong
|
no licence file found · pointer only |
| MMSci: A Dataset for Graduate-Level Multi-Discipline Multimodal Scientific Understanding |
6 Jul 2024 |
leezekun/mmsci/mmsci-exps/model_loader.py f41cb1a19b154297 |
ran · our draft was wrong
|
no licence file found · pointer only |
| Cross-Modality Safety Alignment |
21 Jun 2024 |
identical code first harvested elsewhere f41cb1a19b154297 |
ran · our draft was wrong
|
licence of this copy not recorded |
| Dissecting Adversarial Robustness of Multimodal LM Agents |
18 Jun 2024 |
ChenWu98/agent-attack/agent_attack/models/claude.py b594d1567e135be0 |
ran
|
MIT (permissive) |
| Dissecting Adversarial Robustness of Multimodal LM Agents |
18 Jun 2024 |
ChenWu98/agent-attack/agent_attack/models/gemini.py 72cb0591b722709a |
ran
|
MIT (permissive) |
| ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools |
18 Jun 2024 |
thudm/chatglm4/inference/glm4v_api_request.py c0b1c7a482944767 |
ran
|
Apache-2.0 (permissive) |
| Multimodal Needle in a Haystack: Benchmarking Long-Context Capability of Multimodal Large Language Models |
17 Jun 2024 |
wang-ml-lab/multimodal-needle-in-a-haystack/utils.py 09468b7dfba57e92 |
ran
|
MIT (permissive) |
| AutoHallusion: Automatic Generation of Hallucination Benchmarks for Vision-Language Models |
16 Jun 2024 |
identical code first harvested elsewhere f41cb1a19b154297 |
ran · our draft was wrong
|
licence of this copy not recorded |
| Unveiling the Ignorance of MLLMs: Seeing Clearly, Answering Incorrectly |
15 Jun 2024 |
baai-dcai/multimodal-robustness-benchmark/dataset/data_generation.py f41cb1a19b154297 |
ran · our draft was wrong
|
Apache-2.0 (permissive) |
| Details Make a Difference: Object State-Sensitive Neurorobotic Task Planning |
14 Jun 2024 |
xiao-wen-sun/ossa/utils.py 1e90c2ddcf6921e2 |
unverified |
Apache-2.0 (permissive) |
| Too Many Frames, Not All Useful: Efficient Strategies for Long-Form Video QA |
13 Jun 2024 |
jongwoopark7978/LVNet/fineKeyframeDetector.py f41cb1a19b154297 |
ran · our draft was wrong
|
no licence file found · pointer only |
| VLind-Bench: Measuring Language Priors in Large Vision-Language Models |
13 Jun 2024 |
klee972/vlind-bench/eval/gpt4o_eval.py f41cb1a19b154297 |
ran · our draft was wrong
|
Apache-2.0 (permissive) |
| TC-Bench: Benchmarking Temporal Compositionality in Text-to-Video and Image-to-Video Generation |
12 Jun 2024 |
weixi-feng/tc-bench/vlm_eval.py 3d39217a41c77766 |
ran
|
MIT (permissive) |
| Mobile-Agent-v2: Mobile Device Operation Assistant with Effective Navigation via Multi-Agent Collaboration |
3 Jun 2024 |
identical code first harvested elsewhere f41cb1a19b154297 |
ran · our draft was wrong
|
licence of this copy not recorded |
| G3: An Effective and Adaptive Framework for Worldwide Geolocalization Using Large Multi-Modality Models |
23 May 2024 |
applied-machine-learning-lab/g3/llm_predict.py f41cb1a19b154297 |
ran · our draft was wrong
|
Apache-2.0 (permissive) |
| Unveiling the Tapestry of Consistency in Large Vision-Language Models |
23 May 2024 |
foundation-multimodal-models/conbench/eval/GPT-4o.py f41cb1a19b154297 |
ran · our draft was wrong
|
Apache-2.0 (permissive) |
| OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments |
11 Apr 2024 |
xlang-ai/OSWorld/mm_agents/agent.py 78ee6b46cebd99ee |
ran · our draft was wrong
|
Apache-2.0 (permissive) |