| OralAgent: Integrating Reasoning, Tools, and Knowledge for Interactive Dental Image Analysis added by Syntology |
2026-05 (from id) |
isjinghao/OralAgent/oralagent/llava/utils.py 37899f22fb191b37 |
unverified |
Apache-2.0 (permissive) |
| Natural Language Instructions for Scene-Responsive Human-in-the-Loop Motion Planning in Autonomous Driving using Vision-Language-Action Models added by Syntology |
2026-02 (from id) |
Mi3-Lab/doScenes-VLM-Planning/src/llava/utils.py 37899f22fb191b37 |
unverified |
AGPL-3.0 (copyleft) · pointer only |
| Benchmarking Direct Preference Optimization for Medical Large Vision-Language Models added by Syntology |
2026-01 (from id) |
dmis-lab/med-vlm-dpo/inference/LLaVA-Med/llava/utils.py 37899f22fb191b37 |
unverified |
no licence file found · pointer only |
| arXiv:2507.18300 |
2025-07 (from id) |
360CVGroup/LMM-Det/llava/utils.py 37899f22fb191b37 |
unverified |
Apache-2.0 (permissive) |
| Hunyuan3D 2.5: Towards High-Fidelity 3D Assets Generation with Ultimate Details |
19 Jun 2025 |
tencent/hunyuan3d-2/api_server.py 37899f22fb191b37 |
unverified |
no licence file found · pointer only |
| Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval |
26 May 2025 |
friedrichor/UNITE/unite/utils.py 37899f22fb191b37 |
unverified |
Apache-2.0 (permissive) |
| SegEarth-R1: Geospatial Pixel Reasoning via Large Language Model |
13 Apr 2025 |
earth-insights/segearth-r1/segearth_r1/utils.py 37899f22fb191b37 |
unverified |
Apache-2.0 (permissive) |
| MedRAX: Medical Reasoning Agent for Chest X-ray |
4 Feb 2025 |
bowang-lab/medrax/medrax/llava/utils.py 37899f22fb191b37 |
unverified |
Apache-2.0 (permissive) |
| Lyra: An Efficient and Speech-Centric Framework for Omni-Cognition |
12 Dec 2024 |
dvlab-research/Lyra/lyra/utils.py 37899f22fb191b37 |
unverified |
Apache-2.0 (permissive) |
| Agri-LLaVA: Knowledge-Infused Large Multimodal Assistant on Agricultural Pests and Diseases |
3 Dec 2024 |
kki2eve/agri-llava/agri_llava/utils.py 37899f22fb191b37 |
unverified |
Apache-2.0 (permissive) |
| LongVALE: Vision-Audio-Language-Event Benchmark Towards Time-Aware Omni-Modal Perception of Long Videos |
29 Nov 2024 |
ttgeng233/LongVALE/longvalellm/utils.py 37899f22fb191b37 |
unverified |
MIT (permissive) |
| HyperSeg: Towards Universal Visual Segmentation with Large Language Model |
26 Nov 2024 |
congvvc/HyperSeg/hyperseg/model/mipha/utils.py 37899f22fb191b37 |
unverified |
Apache-2.0 (permissive) |
| TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models |
17 Nov 2024 |
tingyu215/ts-llava/llava/utils.py 37899f22fb191b37 |
unverified |
Apache-2.0 (permissive) |
| VidEgoThink: Assessing Egocentric Video Understanding Capabilities for Embodied AI |
15 Oct 2024 |
adacheng/egothink/models/lego/utils.py 37899f22fb191b37 |
unverified |
Apache-2.0 (permissive) |
| Reconstructive Visual Instruction Tuning |
12 Oct 2024 |
haochen-wang409/ross/ross/utils.py 37899f22fb191b37 |
unverified |
Apache-2.0 (permissive) |
| TRACE: Temporal Grounding Video LLM via Causal Event Modeling |
8 Oct 2024 |
gyxxyg/TRACE/trace/utils.py 37899f22fb191b37 |
unverified |
Apache-2.0 (permissive) |
| One Token to Seg Them All: Language Instructed Reasoning Segmentation in Videos |
29 Sep 2024 |
showlab/videolisa/model/llava/utils.py 37899f22fb191b37 |
unverified |
Apache-2.0 (permissive) |
| CDChat: A Large Multimodal Model for Remote Sensing Change Description |
24 Sep 2024 |
techmn/cdchat/cdchat/utils.py 37899f22fb191b37 |
unverified |
no licence file found · pointer only |
| Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution |
19 Sep 2024 |
Oryx-mllm/Oryx/oryx/utils.py 37899f22fb191b37 |
unverified |
MIT (permissive) |
| IAA: Inner-Adaptor Architecture Empowers Frozen Large Language Model with Multimodal Capabilities |
23 Aug 2024 |
360cvgroup/inner-adaptor-architecture/iaa/utils.py 37899f22fb191b37 |
unverified |
Apache-2.0 (permissive) |
| GalleryGPT: Analyzing Paintings with Large Multimodal Models |
1 Aug 2024 |
steven640pixel/gallerygpt/llava/utils.py 37899f22fb191b37 |
unverified |
Apache-2.0 (permissive) |
| SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models |
22 Jul 2024 |
apple/ml-slowfast-llava/slowfast_llava/llava/utils.py 37899f22fb191b37 |
unverified |
no licence file found · pointer only |
| EVALALIGN: Supervised Fine-Tuning Multimodal LLMs with Human-Aligned Data for Evaluating Text-to-Image Models |
24 Jun 2024 |
sais-fuxi/evalalign/evalalign/utils.py 37899f22fb191b37 |
unverified |
Apache-2.0 (permissive) |
| Biomedical Visual Instruction Tuning with Clinician Preference Alignment |
19 Jun 2024 |
mao1207/BioMed-VITAL/backbone/utils.py 37899f22fb191b37 |
unverified |
no licence file found · pointer only |
| mDPO: Conditional Preference Optimization for Multimodal Large Language Models |
17 Jun 2024 |
luka-group/mDPO/bunny/bunny_utils/util/utils.py 37899f22fb191b37 |
unverified |
no licence file found · pointer only |
| ClawMachine: Learning to Fetch Visual Tokens for Referential Comprehension |
17 Jun 2024 |
martian422/ClawMachine/ClawMachine/utils.py 37899f22fb191b37 |
unverified |
no licence file found · pointer only |
| Concept-skill Transferability-based Data Selection for Large Vision-Language Models |
16 Jun 2024 |
g-jwlee/coincide_code/COINCIDE_cluster/tinyllava/utils.py 37899f22fb191b37 |
unverified |
no licence file found · pointer only |
| RLAIF-V: Open-Source AI Feedback Leads to Super GPT-4V Trustworthiness |
27 May 2024 |
openbmb/omnilmm/omnilmm/utils.py 37899f22fb191b37 |
unverified |
Apache-2.0 (permissive) |
| ConvLLaVA: Hierarchical Backbones as Visual Encoder for Large Multimodal Models |
24 May 2024 |
alibaba/conv-llava/llava/utils.py 37899f22fb191b37 |
unverified |
Apache-2.0 (permissive) |
| FreeVA: Offline MLLM as Training-Free Video Assistant |
13 May 2024 |
whwu95/freeva/llava/utils.py 37899f22fb191b37 |
unverified |
Apache-2.0 (permissive) |
| Detecting and Mitigating Hallucination in Large Vision Language Models via Fine-Grained AI Feedback |
22 Apr 2024 |
Mr-Loevan/HSA-DPO/hsa_dpo/models/llava-v1_5/llava/utils.py 37899f22fb191b37 |
unverified |
no licence file found · pointer only |
| Groma: Localized Visual Tokenization for Grounding Multimodal Large Language Models |
19 Apr 2024 |
FoundationVision/Groma/groma/utils.py 37899f22fb191b37 |
unverified |
Apache-2.0 (permissive) |
| LaSagnA: Language-based Segmentation Assistant for Complex Queries |
12 Apr 2024 |
congvvc/lasagna/model/llava/utils.py 37899f22fb191b37 |
unverified |
Apache-2.0 (permissive) |
| PSALM: Pixelwise SegmentAtion with Large Multi-Modal Model |
21 Mar 2024 |
zamling/PSALM/psalm/utils.py 37899f22fb191b37 |
unverified |
Apache-2.0 (permissive) |
| Towards a clinically accessible radiology foundation model: open-access and lightweight, with automated evaluation |
12 Mar 2024 |
microsoft/llava-rad/llava/utils.py 37899f22fb191b37 |
unverified |
no licence file found · pointer only |
| CAT: Enhancing Multimodal Large Language Model to Answer Questions in Dynamic Audio-Visual Scenarios |
7 Mar 2024 |
rikeilong/bay-cat/ADPO_CAT/model/utils.py 37899f22fb191b37 |
unverified |
Apache-2.0 (permissive) |
| ImgTrojan: Jailbreaking Vision-Language Models with ONE Image |
5 Mar 2024 |
xijia-tao/imgtrojan/finetune/llava/utils.py 37899f22fb191b37 |
unverified |
no licence file found · pointer only |
| Here Comes The AI Worm: Unleashing Zero-click Worms that Target GenAI-Powered Applications |
2024-03 (from id) |
stavc/compromptmized/Legacy_Arxiv_V1/FlowSteering/llava/utils.py 37899f22fb191b37 |
unverified |
no licence file found · pointer only |
| GroundingGPT:Language Enhanced Multi-modal Grounding Model |
11 Jan 2024 |
lzw-lzw/groundinggpt/lego/utils.py 37899f22fb191b37 |
unverified |
Apache-2.0 (permissive) |
| LISA++: An Improved Baseline for Reasoning Segmentation with Large Language Model |
28 Dec 2023 |
dvlab-research/lisa/model/llava/utils.py 37899f22fb191b37 |
unverified |
Apache-2.0 (permissive) |
| Osprey: Pixel Understanding with Visual Instruction Tuning |
15 Dec 2023 |
circleradon/osprey/osprey/utils.py 37899f22fb191b37 |
unverified |
Apache-2.0 (permissive) |
| SMILE: Multimodal Dataset for Understanding Laughter in Video with Language Models |
15 Dec 2023 |
postech-ami/smile-dataset/FastChat/fastchat/utils.py 37899f22fb191b37 |
unverified |
no licence file found · pointer only |
| VTimeLLM: Empower LLM to Grasp Video Moments |
30 Nov 2023 |
huangb23/vtimellm/vtimellm/utils.py 37899f22fb191b37 |
unverified |
no licence file found · pointer only |
| GeoChat: Grounded Large Vision-Language Model for Remote Sensing |
24 Nov 2023 |
mbzuai-oryx/geochat/geochat/utils.py 37899f22fb191b37 |
unverified |
no licence file found · pointer only |
| PG-Video-LLaVA: Pixel Grounding Large Video-Language Models |
22 Nov 2023 |
mbzuai-oryx/video-llava/video_chatgpt/utils.py 37899f22fb191b37 |
unverified |
no licence file found · pointer only |
| Exploring the Reliability of Large Language Models as Customized Evaluators for Diverse NLP Tasks |
30 Oct 2023 |
qtli/coeval/webapp/utils.py 37899f22fb191b37 |
unverified |
Apache-2.0 (permissive) |
| Qilin-Med-VL: Towards Chinese Large Vision-Language Model for General Healthcare |
27 Oct 2023 |
williamliujl/qilin-med-vl/llava/utils.py 37899f22fb191b37 |
unverified |
no licence file found · pointer only |
| Reformulating Vision-Language Foundation Models and Datasets Towards Universal Multimodal Assistants |
1 Oct 2023 |
thunlp/muffin/muffin/utils.py 37899f22fb191b37 |
unverified |
no licence file found · pointer only |
| BT-Adapter: Video Conversation is Feasible Without Video Instruction Tuning |
27 Sep 2023 |
farewellthree/BT-Adapter/llava_base/bt_adapter/utils.py 37899f22fb191b37 |
unverified |
Apache-2.0 (permissive) |
| Sight Beyond Text: Multi-Modal Training Enhances LLMs in Truthfulness and Ethics |
13 Sep 2023 |
ucsc-vlaa/sight-beyond-text/llava/utils.py 37899f22fb191b37 |
unverified |
Apache-2.0 (permissive) |
| Position-Enhanced Visual Instruction Tuning for Multimodal Large Language Models |
25 Aug 2023 |
pvit-official/pvit/pvit/utils.py 37899f22fb191b37 |
unverified |
no licence file found · pointer only |
| GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio |
13 Jun 2021 |
maikezuefle/contr-pretraining/llava/utils.py 37899f22fb191b37 |
unverified |
Apache-2.0 (permissive) |
| arXiv:Zhang_Beyond_Training_Dynamic_Token_Merging_for_Zero-Shot_Video_Understanding_ICCV_2025_paper |
|
Jam1ezhang/DYTO/dyto/llava/utils.py 37899f22fb191b37 |
unverified |
Apache-2.0 (permissive) |
| arXiv:2025.findings-acl.458 |
|
DCDmllm/Align2LLaVA/reward_model/llava/utils.py 37899f22fb191b37 |
unverified |
Apache-2.0 (permissive) |
| arXiv:2024.findings-emnlp.268 |
|
HZQ950419/Math-LLaVA/llava/utils.py 37899f22fb191b37 |
unverified |
Apache-2.0 (permissive) |