| Generative Retrieval for Unsupervised Text-Based Person Search added by Syntology |
2026-09 (from id) |
Flame-Chasers/GTR/models/blip_ps.py 0ec9fc2025c16f65 |
unverified |
no licence file found · pointer only |
| RetroMPA: A Molecular Property-Aware Auxiliary Framework for Enhancing Retrosynthesis Prediction added by Syntology |
2026-08 (from id) |
MengzhouLu/RetroMPA/lavis/models/base_model.py 0ec9fc2025c16f65 |
unverified |
no licence file found · pointer only |
| Do We Really Need Multimodal Emotion Language Models Larger Than 1B Parameters? added by Syntology |
2026-07 (from id) |
GAIR-Lab/Light-MER/my_affectgpt/models/base_model.py 0ec9fc2025c16f65 |
unverified |
Apache-2.0 (permissive) |
| Multimodal Knowledge Edit-Scoped Generalization for Online Recursive MLLM Editing added by Syntology |
2026-07 (from id) |
lab-klc/ScopeEdit/easyeditor/trainer/blip2_models/base_model.py 0ec9fc2025c16f65 |
unverified |
MIT (permissive) |
| Concept-Guided Noisy Negative Suppression for Zero-Shot Classification and Grounding of Chest X-Ray Findings added by Syntology |
2026-05 (from id) |
DopamineLcy/conns/conns/loss.py 8a0642d4189e86e7 |
ran · honoured contract
fingerprinted |
no licence file found · pointer only |
| arXiv:2507.12416 |
2025-07 (from id) |
jackwaky/QuRe/lavis/models/base_model.py 0ec9fc2025c16f65 |
unverified |
no licence file found · pointer only |
| TimeChat-Online: 80% Visual Tokens are Naturally Redundant in Streaming Videos |
24 Apr 2025 |
renshuhuai-andy/timechat/timechat/models/base_model.py 0ec9fc2025c16f65 |
unverified |
BSD-3-Clause (permissive) |
| Seedream 2.0: A Native Chinese-English Bilingual Image Generation Foundation Model |
10 Mar 2025 |
DYEvaLab/EvalMuse/lavis/models/base_model.py 0ec9fc2025c16f65 |
unverified |
Apache-2.0 (permissive) |
| Mitigating Hallucinations in Large Vision-Language Models by Adaptively Constraining Information Flow |
28 Feb 2025 |
jiaqi5598/adavib/minigpt4/models/base_model.py 0ec9fc2025c16f65 |
unverified |
MIT (permissive) |
| When and How Does CLIP Enable Domain and Compositional Generalization? |
13 Feb 2025 |
salesforce/LAVIS/lavis/models/base_model.py 0ec9fc2025c16f65 |
unverified |
BSD-3-Clause (permissive) |
| Implicit Location-Caption Alignment via Complementary Masking for Weakly-Supervised Dense Video Captioning |
17 Dec 2024 |
ShipingGe/ILCACM/src/util/dist.py 0ec9fc2025c16f65 |
unverified |
MIT (permissive) |
| Gramian Multimodal Representation Learning and Alignment |
16 Dec 2024 |
ispamm/GRAM/utils/distributed.py 0ec9fc2025c16f65 |
unverified |
MIT (permissive) |
| PPLLaVA: Varied Video Sequence Understanding With Prompt Guidance |
4 Nov 2024 |
farewellthree/ppllava/ppllava/models/base_model.py 0ec9fc2025c16f65 |
unverified |
Apache-2.0 (permissive) |
| Enhancing Temporal Modeling of Video LLMs via Time Gating |
8 Oct 2024 |
lavi-lab/tg-vid/stllm/models/base_model.py 0ec9fc2025c16f65 |
unverified |
no licence file found · pointer only |
| Image Over Text: Transforming Formula Recognition Evaluation with Character Detection Matching |
5 Sep 2024 |
opendatalab/unimernet/unimernet/models/base_model.py 0ec9fc2025c16f65 |
unverified |
Apache-2.0 (permissive) |
| HERMES: temporal-coHERent long-forM understanding with Episodes and Semantics |
30 Aug 2024 |
joslefaure/HERMES/lavis/models/base_model.py 0ec9fc2025c16f65 |
unverified |
MIT (permissive) |
| ConVis: Contrastive Decoding with Hallucination Visualization for Mitigating Hallucinations in Multimodal Large Language Models |
25 Aug 2024 |
yejipark-m/convis/minigpt4/models/base_model.py 0ec9fc2025c16f65 |
unverified |
MIT (permissive) |
| ParGo: Bridging Vision-Language with Partial and Global Views |
23 Aug 2024 |
bytedance/pargo/pargo/utils/disttools.py 0ec9fc2025c16f65 |
unverified |
BSD-3-Clause (permissive) |
| AGLA: Mitigating Object Hallucinations in Large Vision-Language Models with Assembly of Global and Local Attention |
18 Jun 2024 |
lackel/agla/lavis/models/base_model.py 0ec9fc2025c16f65 |
unverified |
no licence file found · pointer only |
| VRSBench: A Versatile Vision-Language Benchmark Dataset for Remote Sensing Image Understanding |
18 Jun 2024 |
lavender105/rsgpt/rsgpt/models/base_model.py 0ec9fc2025c16f65 |
unverified |
no licence file found · pointer only |
| Jailbreak Vision Language Models via Bi-Modal Adversarial Prompt |
6 Jun 2024 |
NY1024/BAP-Jailbreak-Vision-Language-Models-via-Bi-Modal-Adversarial-Prompt/LAVIS/lavis/models/base_model.py 0ec9fc2025c16f65 |
unverified |
no licence file found · pointer only |
| Text-like Encoding of Collaborative Information in Large Language Models for Recommendation |
5 Jun 2024 |
zyang1580/binllm/minigpt4/models/base_model.py 0ec9fc2025c16f65 |
unverified |
no licence file found · pointer only |
| Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis |
31 May 2024 |
PhysGame/PhysGame/physvlm/models/base_model.py 0ec9fc2025c16f65 |
unverified |
Apache-2.0 (permissive) |
| White-box Multimodal Jailbreaks Against Large Vision-Language Models |
28 May 2024 |
roywang021/UMK/minigpt4/models/base_model.py 0ec9fc2025c16f65 |
unverified |
no licence file found · pointer only |
| Reason3D: Searching and Reasoning 3D Segmentation via Large Language Model |
27 May 2024 |
kuanchihhuang/reason3d/lavis/models/base_model.py 0ec9fc2025c16f65 |
unverified |
no licence file found · pointer only |
| Hawk: Learning to Understand Open-World Video Anomalies |
27 May 2024 |
jqtangust/hawk/hawk/models/base_model.py 0ec9fc2025c16f65 |
unverified |
no licence file found · pointer only |
| Adapting Multi-modal Large Language Model to Concept Drift From Pre-training Onwards |
22 May 2024 |
XiaoyuYoung/ConceptDriftMLLMs/lavis/models/base_model.py 0ec9fc2025c16f65 |
unverified |
BSD-3-Clause (permissive) |
| MovieChat+: Question-aware Sparse Memory for Long Video Question Answering |
26 Apr 2024 |
rese1f/MovieChat/MovieChat/models/base_model.py 0ec9fc2025c16f65 |
unverified |
BSD-3-Clause (permissive) |
| SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension |
25 Apr 2024 |
ailab-cvc/seed-bench/lavis/models/base_model.py 0ec9fc2025c16f65 |
unverified |
no licence file found · pointer only |
| Self-Bootstrapped Visual-Language Model for Knowledge Selection and Question Answering |
22 Apr 2024 |
haodongze/self-ksel-qans/lavis/models/base_model.py 0ec9fc2025c16f65 |
unverified |
MIT (permissive) |
| MA-LMM: Memory-Augmented Large Multimodal Model for Long-Term Video Understanding |
8 Apr 2024 |
boheumd/MA-LMM/lavis/models/base_model.py 0ec9fc2025c16f65 |
unverified |
MIT (permissive) |
| From Pixels to Graphs: Open-Vocabulary Scene Graph Generation with Vision-Language Models |
1 Apr 2024 |
shtuplus/pix2grp_cvpr2024/lavis/models/base_model.py 0ec9fc2025c16f65 |
unverified |
BSD-3-Clause (permissive) |
| ST-LLM: Large Language Models Are Effective Temporal Learners |
30 Mar 2024 |
TencentARC/ST-LLM/stllm/models/base_model.py 0ec9fc2025c16f65 |
unverified |
Apache-2.0 (permissive) |
| Generative Multi-modal Models are Good Class-Incremental Learners |
27 Mar 2024 |
DoubleClass/GMM/minigpt4/models/base_model.py 0ec9fc2025c16f65 |
unverified |
no licence file found · pointer only |
| OmniVid: A Generative Framework for Universal Video Understanding |
26 Mar 2024 |
wangjk666/OmniVid/lavis/models/base_model.py 0ec9fc2025c16f65 |
unverified |
no licence file found · pointer only |
| CAT: Enhancing Multimodal Large Language Model to Answer Questions in Dynamic Audio-Visual Scenarios |
7 Mar 2024 |
rikeilong/bay-cat/ADPO_CAT/model/base_model.py 0ec9fc2025c16f65 |
unverified |
Apache-2.0 (permissive) |
| Embodied Understanding of Driving Scenarios |
7 Mar 2024 |
opendrivelab/elm/lavis/models/base_model.py 0ec9fc2025c16f65 |
unverified |
no licence file found · pointer only |
| MADTP: Multimodal Alignment-Guided Dynamic Token Pruning for Accelerating Vision-Language Transformer |
5 Mar 2024 |
double125/madtp/models/blip_retrieval.py 0ec9fc2025c16f65 |
unverified |
Apache-2.0 (permissive) |
| Grounding Language Models for Visual Entity Recognition |
28 Feb 2024 |
mrzilinxiao/autover/common_utils/dist_utils.py 0ec9fc2025c16f65 |
unverified |
Apache-2.0 (permissive) |
| Stop Reasoning! When Multimodal LLM with Chain-of-Thought Reasoning Meets Adversarial Image |
22 Feb 2024 |
aipenguin/stopreasoning/minigpt4/models/base_model.py 0ec9fc2025c16f65 |
unverified |
no licence file found · pointer only |
| VisLingInstruct: Elevating Zero-Shot Learning in Multi-Modal Language Models with Autonomous Instruction Optimization |
12 Feb 2024 |
zhudongsheng75/vislinginstruct/vislinginstruct/models/base_model.py 0ec9fc2025c16f65 |
unverified |
no licence file found · pointer only |
| GraphTranslator: Aligning Graph Model to Large Language Model for Open-ended Tasks |
11 Feb 2024 |
alibaba/graphtranslator/Translator/models/base_model.py 5d118a6ac70d4e06 |
ran
fingerprinted |
BSD-3-Clause (permissive) |
| CREMA: Generalizable and Efficient Video-Language Reasoning via Multimodal Modular Fusion |
8 Feb 2024 |
Yui010206/CREMA/lavis/models/base_model.py 0ec9fc2025c16f65 |
unverified |
BSD-3-Clause (permissive) |
| Inducing High Energy-Latency of Large Vision-Language Models with Verbose Images |
20 Jan 2024 |
KuofengGao/Verbose_Images/lavis/models/base_model.py 0ec9fc2025c16f65 |
unverified |
no licence file found · pointer only |
| GroundingGPT:Language Enhanced Multi-modal Grounding Model |
11 Jan 2024 |
lzw-lzw/groundinggpt/video_llama/models/base_model.py 0ec9fc2025c16f65 |
unverified |
Apache-2.0 (permissive) |
| Model Editing Harms General Abilities of Large Language Models: Regularization to the Rescue |
9 Jan 2024 |
jasonforjoy/model-editing-hurt/easyeditor/trainer/blip2_models/base_model.py 0ec9fc2025c16f65 |
unverified |
no licence file found · pointer only |
| Moonshot: Towards Controllable Video Generation and Editing with Multimodal Conditions |
3 Jan 2024 |
salesforce/lavis/lavis/models/base_model.py 0ec9fc2025c16f65 |
unverified |
BSD-3-Clause (permissive) |
| When Parameter-efficient Tuning Meets General-purpose Vision-language Models |
16 Dec 2023 |
melonking32/petal/lavis/models/base_model.py 0ec9fc2025c16f65 |
unverified |
no licence file found · pointer only |
| Image Content Generation with Causal Reasoning |
12 Dec 2023 |
ieit-agi/mix-shannon/lavis/models/base_model.py 0ec9fc2025c16f65 |
unverified |
Apache-2.0 (permissive) |
| Localized Symbolic Knowledge Distillation for Visual Commonsense Models |
8 Dec 2023 |
jamespark3922/lskd/lavis/models/base_model.py 0ec9fc2025c16f65 |
unverified |
BSD-3-Clause (permissive) |
| GPT4Point: A Unified Framework for Point-Language Understanding and Generation |
5 Dec 2023 |
Pointcept/GPT4Point/lavis/models/base_model.py 0ec9fc2025c16f65 |
unverified |
MIT (permissive) |
| TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video Understanding |
4 Dec 2023 |
lntzm/cvpr24track-longvideo/timechat/models/base_model.py 0ec9fc2025c16f65 |
unverified |
BSD-3-Clause (permissive) |
| Synthesize, Diagnose, and Optimize: Towards Fine-Grained Vision-Language Understanding |
30 Nov 2023 |
wjpoom/spec/spec/models/blip_utils/blip_retrieval.py 0ec9fc2025c16f65 |
unverified |
no licence file found · pointer only |
| X-InstructBLIP: A Framework for aligning X-Modal instruction-aware representations to LLMs and Emergent Cross-modal Reasoning |
30 Nov 2023 |
artemisp/lavis-xinstructblip/lavis/models/base_model.py 0ec9fc2025c16f65 |
unverified |
BSD-3-Clause (permissive) |
| Mitigating Object Hallucinations in Large Vision-Language Models through Visual Contrastive Decoding |
28 Nov 2023 |
damo-nlp-sg/vcd/experiments/lavis/models/base_model.py 0ec9fc2025c16f65 |
unverified |
Apache-2.0 (permissive) |
| Embodied Multi-Modal Agent trained by an LLM from a Parallel TextWorld |
28 Nov 2023 |
stevenyangyj/emma-alfworld/LAVIS/lavis/models/base_model.py 0ec9fc2025c16f65 |
unverified |
no licence file found · pointer only |
| MultiDelete for Multimodal Machine Unlearning |
18 Nov 2023 |
chengjiali/Multimodal-Machine-Unlearning/lavis/models/base_model.py 0ec9fc2025c16f65 |
unverified |
no licence file found · pointer only |
| MMC: Advancing Multimodal Chart Understanding with Large-scale Instruction Tuning |
15 Nov 2023 |
FuxiaoLiu/LRV-Instruction/MiniGPT-4/minigpt4/models/base_model.py 0ec9fc2025c16f65 |
unverified |
BSD-3-Clause (permissive) |
| CoLLM: Integrating Collaborative Embeddings into Large Language Models for Recommendation |
30 Oct 2023 |
zyang1580/collm/minigpt4/models/base_model.py 0ec9fc2025c16f65 |
unverified |
BSD-3-Clause (permissive) |
| Myriad: Large Multimodal Model by Applying Vision Experts for Industrial Anomaly Detection |
29 Oct 2023 |
tzjtatata/myriad/minigpt4/models/base_model.py 0ec9fc2025c16f65 |
unverified |
no licence file found · pointer only |
| TESTA: Temporal-Spatial Token Aggregation for Long-form Video-Language Understanding |
29 Oct 2023 |
renshuhuai-andy/testa/models/testa_retrieval.py 0ec9fc2025c16f65 |
unverified |
MIT (permissive) |
| MarineGPT: Unlocking Secrets of Ocean to the Public |
20 Oct 2023 |
hkust-vgd/marinegpt/marinegpt/models/base_model.py 0ec9fc2025c16f65 |
unverified |
no licence file found · pointer only |
| Fine-grained Audio-Visual Joint Representations for Multimodal Large Language Models |
9 Oct 2023 |
the-anonymous-bs/favor/video_llama/models/base_model.py 0ec9fc2025c16f65 |
unverified |
no licence file found · pointer only |
| Rephrase, Augment, Reason: Visual Grounding of Questions for Vision-Language Models |
9 Oct 2023 |
archiki/repare/Lavis/lavis/models/base_model.py 0ec9fc2025c16f65 |
unverified |
MIT (permissive) |
| VLATTACK: Multimodal Adversarial Attacks on Vision-Language Tasks via Pre-trained Models |
7 Oct 2023 |
ericyinyzy/vlattack/BLIP_attack/models/blip_retrieval.py 0ec9fc2025c16f65 |
unverified |
BSD-3-Clause (permissive) |
| Expedited Training of Visual Conditioned Language Generation via Redundancy Reduction |
5 Oct 2023 |
yiren-jian/evlgen/lavis/models/base_model.py 0ec9fc2025c16f65 |
unverified |
BSD-3-Clause (permissive) |
| MiniGPT-5: Interleaved Vision-and-Language Generation via Generative Vokens |
3 Oct 2023 |
eric-ai-lab/minigpt-5/minigpt4/models/base_model.py 0ec9fc2025c16f65 |
unverified |
Apache-2.0 (permissive) |
| Analyzing and Mitigating Object Hallucination in Large Vision-Language Models |
1 Oct 2023 |
YiyangZhou/LURE/minigpt4/models/base_model.py 0ec9fc2025c16f65 |
unverified |
no licence file found · pointer only |
| Box-based Refinement for Weakly Supervised and Unsupervised Localization Tasks |
7 Sep 2023 |
eyalgomel/box-based-refinement/BLIP/models/blip_retrieval.py 0ec9fc2025c16f65 |
unverified |
no licence file found · pointer only |
| Sparkles: Unlocking Chats Across Multiple Images for Multimodal Instruction-Following Models |
31 Aug 2023 |
HYPJUDY/Sparkles/sparkles/models/base_model.py 0ec9fc2025c16f65 |
unverified |
BSD-3-Clause (permissive) |
| VIGC: Visual Instruction Generation and Correction |
24 Aug 2023 |
opendatalab/vigc/vigc/models/base_model.py 0ec9fc2025c16f65 |
unverified |
Apache-2.0 (permissive) |
| Simple Baselines for Interactive Video Retrieval with Questions and Answers |
21 Aug 2023 |
kevinliang888/IVR-QA-baselines/models/blip_retrieval.py 0ec9fc2025c16f65 |
unverified |
no licence file found · pointer only |
| BLIVA: A Simple Multimodal LLM for Better Handling of Text-Rich Visual Questions |
19 Aug 2023 |
mlpc-ucsd/bliva/bliva/models/base_model.py 0ec9fc2025c16f65 |
unverified |
BSD-3-Clause (permissive) |
| CTP: Towards Vision-Language Continual Pretraining via Compatible Momentum Contrast and Topology Preservation |
14 Aug 2023 |
kevinlight831/ctp/models/clip_pretrain.py 0ec9fc2025c16f65 |
unverified |
MIT (permissive) |
| Fine-tuning Multimodal LLMs to Follow Zero-shot Demonstrative Instructions |
8 Aug 2023 |
DCDmllm/Cheetah/Cheetah/cheetah/models/base_model.py 0ec9fc2025c16f65 |
unverified |
BSD-3-Clause (permissive) |
| 3D-LLM: Injecting the 3D World into Large Language Models |
24 Jul 2023 |
umass-foundation-model/3d-llm/3DLLM_BLIP2-base/lavis/models/base_model.py 0ec9fc2025c16f65 |
unverified |
MIT recorded; this copy not marked cleared · pointer only |
| FigCaps-HF: A Figure-to-Caption Generative Framework and Benchmark with Human Feedback |
20 Jul 2023 |
figcapshf/figcapshf/models/blip_retrieval.py 0ec9fc2025c16f65 |
unverified |
no licence file found · pointer only |
| BuboGPT: Enabling Visual Grounding in Multi-Modal LLMs |
17 Jul 2023 |
magic-research/bubogpt/bubogpt/models/base_model.py 0ec9fc2025c16f65 |
unverified |
BSD-3-Clause (permissive) |
| Bootstrapping Vision-Language Learning with Decoupled Language Pre-training |
13 Jul 2023 |
yiren-jian/BLIText/lavis/models/base_model.py 0ec9fc2025c16f65 |
unverified |
BSD-3-Clause (permissive) |
| Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding |
5 Jun 2023 |
damo-nlp-sg/video-llama/video_llama/models/base_model.py 0ec9fc2025c16f65 |
unverified |
BSD-3-Clause recorded; this copy not marked cleared · pointer only |
| VAST: A Vision-Audio-Subtitle-Text Omni-Modality Foundation Model and Dataset |
29 May 2023 |
txh-mercury/vast/model/vast.py 0ec9fc2025c16f65 |
unverified |
MIT (permissive) |
| ChatBridge: Bridging Modalities with Large Language Model as a Language Catalyst |
25 May 2023 |
joez17/chatbridge/chatbridge/models/base_model.py 0ec9fc2025c16f65 |
unverified |
BSD-3-Clause (permissive) |
| DetGPT: Detect What You Need via Reasoning |
23 May 2023 |
optimalscale/detgpt/detgpt/models/base_model.py 0ec9fc2025c16f65 |
unverified |
BSD-3-Clause (permissive) |
| X-LLM: Bootstrapping Advanced Large Language Models by Treating Multi-Modalities as Foreign Languages |
7 May 2023 |
phellonchen/x-llm/xllm/models/base_model.py 0ec9fc2025c16f65 |
unverified |
Apache-2.0 (permissive) |
| UniAdapter: Unified Parameter-Efficient Transfer Learning for Cross-modal Modeling |
13 Feb 2023 |
rerv/uniadapter/models/blip_retrieval.py 0ec9fc2025c16f65 |
unverified |
BSD-3-Clause recorded; this copy not marked cleared · pointer only |
| Plug-and-Play VQA: Zero-shot VQA by Conjoining Large Pretrained Models with Zero Training |
17 Oct 2022 |
Tzoulio/Large_Models_Dialogue_for_Active_Perception/llm-vqa_dialogue/lavis/models/base_model.py 0ec9fc2025c16f65 |
unverified |
MIT recorded; this copy not marked cleared · pointer only |
| Masked Unsupervised Self-training for Label-free Image Classification |
7 Jun 2022 |
salesforce/must/utils.py 0ec9fc2025c16f65 |
unverified |
BSD-3-Clause (permissive) |
| arXiv:openreview_yMYRr00HFS |
|
XLearning-SCU/2025-ICLR-TCR/ddp.py aaf3b629a2583e5a |
unverified |
Apache-2.0 (permissive) |
| arXiv:aaai_35495 |
|
salesforce/BLIP/models/blip_retrieval.py 0ec9fc2025c16f65 |
unverified |
BSD-3-Clause (permissive) |
| arXiv:aaai_27999 |
|
mlpc-ucsd/BLIVA/bliva/models/base_model.py 0ec9fc2025c16f65 |
unverified |
BSD-3-Clause (permissive) |
| arXiv:Huang_OPERA_Alleviating_Hallucination_in_Multi-Modal_Large_Language_Models_via_Over-Trust_CVPR_2024_paper |
|
shikiw/OPERA/minigpt4/models/base_model.py 0ec9fc2025c16f65 |
unverified |
MIT (permissive) |
| arXiv:2024.emnlp-main.722 |
|
xyh97/UNICORN/lavis/models/base_model.py 0ec9fc2025c16f65 |
unverified |
BSD-3-Clause recorded; this copy not marked cleared · pointer only |
| arXiv:2024.acl-long.19 |
|
yiren-jian/EVLGen/lavis/models/base_model.py 0ec9fc2025c16f65 |
unverified |
BSD-3-Clause recorded; this copy not marked cleared · pointer only |