| Decoupled Vision-Language System for Multimodal Understanding and Generation added by Syntology |
2026-08 (from id) |
YifanXu74/Libra/libra/common/config.py 791c72070c0b1cfb |
ran
|
Apache-2.0 (permissive) |
| CSMCIR: CoT-Enhanced Symmetric Alignment with Memory Bank for Composed Image Retrieval added by Syntology |
2026-01 (from id) |
qzp2018/CSMCIR/src/lavis/common/config.py 791c72070c0b1cfb |
ran
|
no licence file found · pointer only |
| TimeChat-Online: 80% Visual Tokens are Naturally Redundant in Streaming Videos |
24 Apr 2025 |
renshuhuai-andy/timechat/timechat/common/config.py 791c72070c0b1cfb |
ran
|
BSD-3-Clause (permissive) |
| Seedream 2.0: A Native Chinese-English Bilingual Image Generation Foundation Model |
10 Mar 2025 |
DYEvaLab/EvalMuse/lavis/common/config.py 791c72070c0b1cfb |
ran
|
Apache-2.0 (permissive) |
| When and How Does CLIP Enable Domain and Compositional Generalization? |
13 Feb 2025 |
salesforce/LAVIS/lavis/common/config.py 791c72070c0b1cfb |
ran
|
BSD-3-Clause (permissive) |
| Image Over Text: Transforming Formula Recognition Evaluation with Character Detection Matching |
5 Sep 2024 |
opendatalab/unimernet/unimernet/common/config.py 791c72070c0b1cfb |
ran
|
Apache-2.0 (permissive) |
| HERMES: temporal-coHERent long-forM understanding with Episodes and Semantics |
30 Aug 2024 |
joslefaure/HERMES/lavis/common/config.py 791c72070c0b1cfb |
ran
|
MIT (permissive) |
| SZTU-CMU at MER2024: Improving Emotion-LLaMA with Conv-Attention for Multimodal Emotion Recognition |
20 Aug 2024 |
zebangcheng/emotion-llama/minigpt4/common/config.py 791c72070c0b1cfb |
ran
|
BSD-3-Clause (permissive) |
| Cross-modality Information Check for Detecting Jailbreaking in Multimodal Large Language Models |
31 Jul 2024 |
pandragonxiii/cider/code/models/minigpt4/common/config.py 791c72070c0b1cfb |
ran
|
Apache-2.0 (permissive) |
| VRSBench: A Versatile Vision-Language Benchmark Dataset for Remote Sensing Image Understanding |
18 Jun 2024 |
lavender105/rsgpt/rsgpt/common/config.py 791c72070c0b1cfb |
ran
|
no licence file found · pointer only |
| Reason3D: Searching and Reasoning 3D Segmentation via Large Language Model |
27 May 2024 |
kuanchihhuang/reason3d/lavis/common/config.py 791c72070c0b1cfb |
ran
|
no licence file found · pointer only |
| Hawk: Learning to Understand Open-World Video Anomalies |
27 May 2024 |
jqtangust/hawk/hawk/common/config.py 791c72070c0b1cfb |
ran
|
no licence file found · pointer only |
| Libra: Building Decoupled Vision System on Large Language Models |
16 May 2024 |
yifanxu74/libra/libra/common/config.py 791c72070c0b1cfb |
ran
|
Apache-2.0 (permissive) |
| SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension |
25 Apr 2024 |
ailab-cvc/seed-bench/lavis/common/config.py 791c72070c0b1cfb |
ran
|
no licence file found · pointer only |
| Self-Bootstrapped Visual-Language Model for Knowledge Selection and Question Answering |
22 Apr 2024 |
haodongze/self-ksel-qans/lavis/common/config.py 791c72070c0b1cfb |
ran
|
MIT (permissive) |
| Embodied Understanding of Driving Scenarios |
7 Mar 2024 |
opendrivelab/elm/lavis/common/config.py 791c72070c0b1cfb |
ran
|
no licence file found · pointer only |
| OmniMedVQA: A New Large-Scale Comprehensive Evaluation Benchmark for Medical LVLM |
14 Feb 2024 |
opengvlab/multi-modality-arena/LVLM_evaluation/Multi_turn_Reasoning/lib/minigpt4_config.py 791c72070c0b1cfb |
ran
|
no licence file found · pointer only |
| VisLingInstruct: Elevating Zero-Shot Learning in Multi-Modal Language Models with Autonomous Instruction Optimization |
12 Feb 2024 |
zhudongsheng75/vislinginstruct/vislinginstruct/common/config.py 791c72070c0b1cfb |
ran
|
no licence file found · pointer only |
| GraphTranslator: Aligning Graph Model to Large Language Model for Open-ended Tasks |
11 Feb 2024 |
alibaba/graphtranslator/Translator/common/config.py 791c72070c0b1cfb |
ran
|
BSD-3-Clause (permissive) |
| CREMA: Generalizable and Efficient Video-Language Reasoning via Multimodal Modular Fusion |
8 Feb 2024 |
Yui010206/CREMA/lavis/common/config.py 791c72070c0b1cfb |
ran
|
BSD-3-Clause (permissive) |
| Inducing High Energy-Latency of Large Vision-Language Models with Verbose Images |
20 Jan 2024 |
KuofengGao/Verbose_Images/lavis/common/config.py 791c72070c0b1cfb |
ran
|
no licence file found · pointer only |
| Moonshot: Towards Controllable Video Generation and Editing with Multimodal Conditions |
3 Jan 2024 |
salesforce/lavis/lavis/common/config.py 791c72070c0b1cfb |
ran
|
BSD-3-Clause (permissive) |
| When Parameter-efficient Tuning Meets General-purpose Vision-language Models |
16 Dec 2023 |
melonking32/petal/lavis/common/config.py 791c72070c0b1cfb |
ran
|
no licence file found · pointer only |
| EgoPlan-Bench: Benchmarking Multimodal Large Language Models for Human-Level Planning |
11 Dec 2023 |
chenyi99/egoplan/src/video_llama/video_llama/common/config.py 791c72070c0b1cfb |
ran
|
BSD-3-Clause (permissive) |
| X-InstructBLIP: A Framework for aligning X-Modal instruction-aware representations to LLMs and Emergent Cross-modal Reasoning |
30 Nov 2023 |
artemisp/lavis-xinstructblip/lavis/common/config.py 791c72070c0b1cfb |
ran
|
BSD-3-Clause (permissive) |
| MMC: Advancing Multimodal Chart Understanding with Large-scale Instruction Tuning |
15 Nov 2023 |
FuxiaoLiu/LRV-Instruction/MiniGPT-4/minigpt4/common/config.py 791c72070c0b1cfb |
ran
|
BSD-3-Clause (permissive) |
| MarineGPT: Unlocking Secrets of Ocean to the Public |
20 Oct 2023 |
hkust-vgd/marinegpt/marinegpt/common/config.py 791c72070c0b1cfb |
ran
|
no licence file found · pointer only |
| Fine-grained Audio-Visual Joint Representations for Multimodal Large Language Models |
9 Oct 2023 |
the-anonymous-bs/favor/video_llama/common/config.py 791c72070c0b1cfb |
ran
|
no licence file found · pointer only |
| Sentence-level Prompts Benefit Composed Image Retrieval |
9 Oct 2023 |
chunmeifeng/sprc/src/lavis/common/config.py 791c72070c0b1cfb |
ran
|
no licence file found · pointer only |
| Expedited Training of Visual Conditioned Language Generation via Redundancy Reduction |
5 Oct 2023 |
yiren-jian/evlgen/lavis/common/config.py 791c72070c0b1cfb |
ran
|
BSD-3-Clause (permissive) |
| Sparkles: Unlocking Chats Across Multiple Images for Multimodal Instruction-Following Models |
31 Aug 2023 |
HYPJUDY/Sparkles/sparkles/common/config.py 791c72070c0b1cfb |
ran
|
BSD-3-Clause (permissive) |
| VIGC: Visual Instruction Generation and Correction |
24 Aug 2023 |
opendatalab/vigc/vigc/common/config.py 791c72070c0b1cfb |
ran
|
Apache-2.0 (permissive) |
| Fine-tuning Multimodal LLMs to Follow Zero-shot Demonstrative Instructions |
8 Aug 2023 |
DCDmllm/Cheetah/Cheetah/cheetah/common/config.py 791c72070c0b1cfb |
ran
|
BSD-3-Clause (permissive) |
| BuboGPT: Enabling Visual Grounding in Multi-Modal LLMs |
17 Jul 2023 |
magic-research/bubogpt/bubogpt/common/config.py 791c72070c0b1cfb |
ran
|
BSD-3-Clause (permissive) |
| Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding |
5 Jun 2023 |
damo-nlp-sg/video-llama/video_llama/common/config.py 791c72070c0b1cfb |
ran
|
BSD-3-Clause recorded; this copy not marked cleared · pointer only |
| ChatBridge: Bridging Modalities with Large Language Model as a Language Catalyst |
25 May 2023 |
joez17/chatbridge/chatbridge/common/config.py 791c72070c0b1cfb |
ran
|
BSD-3-Clause (permissive) |
| DetGPT: Detect What You Need via Reasoning |
23 May 2023 |
optimalscale/detgpt/detgpt/common/config.py 791c72070c0b1cfb |
ran
|
BSD-3-Clause (permissive) |
| X-LLM: Bootstrapping Advanced Large Language Models by Treating Multi-Modalities as Foreign Languages |
7 May 2023 |
phellonchen/x-llm/xllm/common/config.py 791c72070c0b1cfb |
ran
|
Apache-2.0 (permissive) |
| arXiv:2024.findings-emnlp.803 |
|
PandragonXIII/CIDER/code/models/minigpt4/common/config.py 791c72070c0b1cfb |
ran
|
Apache-2.0 (permissive) |
| arXiv:2024.acl-long.19 |
|
yiren-jian/EVLGen/lavis/common/config.py 791c72070c0b1cfb |
ran
|
BSD-3-Clause recorded; this copy not marked cleared · pointer only |