| Decoupled Vision-Language System for Multimodal Understanding and Generation added by Syntology |
2026-08 (from id) |
YifanXu74/Libra/libra/common/dist_utils.py 643539116e049080 |
unverified |
Apache-2.0 (permissive) |
| CSMCIR: CoT-Enhanced Symmetric Alignment with Memory Bank for Composed Image Retrieval added by Syntology |
2026-01 (from id) |
qzp2018/CSMCIR/src/lavis/common/dist_utils.py 643539116e049080 |
unverified |
no licence file found · pointer only |
| TimeChat-Online: 80% Visual Tokens are Naturally Redundant in Streaming Videos |
24 Apr 2025 |
renshuhuai-andy/timechat/timechat/common/dist_utils.py 643539116e049080 |
unverified |
BSD-3-Clause (permissive) |
| When and How Does CLIP Enable Domain and Compositional Generalization? |
13 Feb 2025 |
salesforce/LAVIS/lavis/common/dist_utils.py 643539116e049080 |
unverified |
BSD-3-Clause (permissive) |
| Learning to Merge Tokens via Decoupled Embedding for Efficient Vision Transformers |
13 Dec 2024 |
huggingface/pytorch-image-models/timm/models/_hub.py 5a0d348f6fe9632c |
unverified |
Apache-2.0 (permissive) |
| Image Over Text: Transforming Formula Recognition Evaluation with Character Detection Matching |
5 Sep 2024 |
opendatalab/unimernet/unimernet/common/dist_utils.py 643539116e049080 |
unverified |
Apache-2.0 (permissive) |
| HERMES: temporal-coHERent long-forM understanding with Episodes and Semantics |
30 Aug 2024 |
joslefaure/HERMES/lavis/common/dist_utils.py aee911f7b46c600a |
ran
|
MIT (permissive) |
| SZTU-CMU at MER2024: Improving Emotion-LLaMA with Conv-Attention for Multimodal Emotion Recognition |
20 Aug 2024 |
zebangcheng/emotion-llama/minigpt4/common/dist_utils.py 643539116e049080 |
unverified |
BSD-3-Clause (permissive) |
| Cross-modality Information Check for Detecting Jailbreaking in Multimodal Large Language Models |
31 Jul 2024 |
pandragonxiii/cider/code/models/minigpt4/common/dist_utils.py 643539116e049080 |
unverified |
Apache-2.0 (permissive) |
| VRSBench: A Versatile Vision-Language Benchmark Dataset for Remote Sensing Image Understanding |
18 Jun 2024 |
lavender105/rsgpt/rsgpt/common/dist_utils.py 643539116e049080 |
unverified |
no licence file found · pointer only |
| Hawk: Learning to Understand Open-World Video Anomalies |
27 May 2024 |
jqtangust/hawk/hawk/common/dist_utils.py 643539116e049080 |
unverified |
no licence file found · pointer only |
| UniRAG: Universal Retrieval Augmentation for Large Vision Language Models |
16 May 2024 |
castorini/unirag/src/unirag/utils.py 643539116e049080 |
unverified |
no licence file found · pointer only |
| Libra: Building Decoupled Vision System on Large Language Models |
16 May 2024 |
yifanxu74/libra/libra/common/dist_utils.py 643539116e049080 |
unverified |
Apache-2.0 (permissive) |
| Self-Bootstrapped Visual-Language Model for Knowledge Selection and Question Answering |
22 Apr 2024 |
haodongze/self-ksel-qans/lavis/common/dist_utils.py 643539116e049080 |
unverified |
MIT (permissive) |
| Embodied Understanding of Driving Scenarios |
7 Mar 2024 |
opendrivelab/elm/lavis/common/dist_utils.py 643539116e049080 |
unverified |
no licence file found · pointer only |
| VisLingInstruct: Elevating Zero-Shot Learning in Multi-Modal Language Models with Autonomous Instruction Optimization |
12 Feb 2024 |
zhudongsheng75/vislinginstruct/vislinginstruct/common/dist_utils.py 643539116e049080 |
unverified |
no licence file found · pointer only |
| GraphTranslator: Aligning Graph Model to Large Language Model for Open-ended Tasks |
11 Feb 2024 |
alibaba/graphtranslator/Translator/common/dist_utils.py 643539116e049080 |
unverified |
BSD-3-Clause (permissive) |
| CREMA: Generalizable and Efficient Video-Language Reasoning via Multimodal Modular Fusion |
8 Feb 2024 |
Yui010206/CREMA/lavis/common/dist_utils.py 643539116e049080 |
unverified |
BSD-3-Clause (permissive) |
| Video-LaVIT: Unified Video-Language Pre-training with Decoupled Visual-Motional Tokenization |
5 Feb 2024 |
jy0205/lavit/LaVIT/utils.py 643539116e049080 |
unverified |
no licence file found · pointer only |
| Inducing High Energy-Latency of Large Vision-Language Models with Verbose Images |
20 Jan 2024 |
KuofengGao/Verbose_Images/lavis/common/dist_utils.py 643539116e049080 |
unverified |
no licence file found · pointer only |
| Moonshot: Towards Controllable Video Generation and Editing with Multimodal Conditions |
3 Jan 2024 |
salesforce/lavis/lavis/common/dist_utils.py 643539116e049080 |
unverified |
BSD-3-Clause (permissive) |
| When Parameter-efficient Tuning Meets General-purpose Vision-language Models |
16 Dec 2023 |
melonking32/petal/lavis/common/dist_utils.py 643539116e049080 |
unverified |
no licence file found · pointer only |
| EgoPlan-Bench: Benchmarking Multimodal Large Language Models for Human-Level Planning |
11 Dec 2023 |
chenyi99/egoplan/src/video_llama/video_llama/common/dist_utils.py 643539116e049080 |
unverified |
BSD-3-Clause (permissive) |
| X-InstructBLIP: A Framework for aligning X-Modal instruction-aware representations to LLMs and Emergent Cross-modal Reasoning |
30 Nov 2023 |
artemisp/lavis-xinstructblip/lavis/common/dist_utils.py 643539116e049080 |
unverified |
BSD-3-Clause (permissive) |
| MarineGPT: Unlocking Secrets of Ocean to the Public |
20 Oct 2023 |
hkust-vgd/marinegpt/marinegpt/common/dist_utils.py 643539116e049080 |
unverified |
no licence file found · pointer only |
| Fine-grained Audio-Visual Joint Representations for Multimodal Large Language Models |
9 Oct 2023 |
the-anonymous-bs/favor/video_llama/common/dist_utils.py 643539116e049080 |
unverified |
no licence file found · pointer only |
| Sentence-level Prompts Benefit Composed Image Retrieval |
9 Oct 2023 |
chunmeifeng/sprc/src/lavis/common/dist_utils.py 643539116e049080 |
unverified |
no licence file found · pointer only |
| Expedited Training of Visual Conditioned Language Generation via Redundancy Reduction |
5 Oct 2023 |
yiren-jian/evlgen/lavis/common/dist_utils.py 643539116e049080 |
unverified |
BSD-3-Clause (permissive) |
| Making LLaMA SEE and Draw with SEED Tokenizer |
2 Oct 2023 |
ailab-cvc/seed/models/seed_qformer/utils.py 643539116e049080 |
unverified |
no licence file found · pointer only |
| Sparkles: Unlocking Chats Across Multiple Images for Multimodal Instruction-Following Models |
31 Aug 2023 |
HYPJUDY/Sparkles/sparkles/common/dist_utils.py 643539116e049080 |
unverified |
BSD-3-Clause (permissive) |
| VIGC: Visual Instruction Generation and Correction |
24 Aug 2023 |
opendatalab/vigc/vigc/common/dist_utils.py 643539116e049080 |
unverified |
Apache-2.0 (permissive) |
| Fine-tuning Multimodal LLMs to Follow Zero-shot Demonstrative Instructions |
8 Aug 2023 |
DCDmllm/Cheetah/Cheetah/cheetah/common/dist_utils.py 643539116e049080 |
unverified |
BSD-3-Clause (permissive) |
| BuboGPT: Enabling Visual Grounding in Multi-Modal LLMs |
17 Jul 2023 |
magic-research/bubogpt/bubogpt/common/dist_utils.py 643539116e049080 |
unverified |
BSD-3-Clause (permissive) |
| Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding |
5 Jun 2023 |
damo-nlp-sg/video-llama/video_llama/common/dist_utils.py 643539116e049080 |
unverified |
BSD-3-Clause recorded; this copy not marked cleared · pointer only |
| ChatBridge: Bridging Modalities with Large Language Model as a Language Catalyst |
25 May 2023 |
joez17/chatbridge/chatbridge/common/dist_utils.py 643539116e049080 |
unverified |
BSD-3-Clause (permissive) |
| DetGPT: Detect What You Need via Reasoning |
23 May 2023 |
optimalscale/detgpt/detgpt/common/dist_utils.py 643539116e049080 |
unverified |
BSD-3-Clause (permissive) |
| X-LLM: Bootstrapping Advanced Large Language Models by Treating Multi-Modalities as Foreign Languages |
7 May 2023 |
phellonchen/x-llm/xllm/common/dist_utils.py 643539116e049080 |
unverified |
Apache-2.0 (permissive) |
| Can CNNs Be More Robust Than Transformers? |
7 Jun 2022 |
ucsc-vlaa/robustcnn/timm/models/hub.py 32933dcf801320f2 |
unverified |
MIT (permissive) |
| arXiv:2024.findings-emnlp.803 |
|
PandragonXIII/CIDER/code/models/minigpt4/common/dist_utils.py 643539116e049080 |
unverified |
Apache-2.0 (permissive) |
| arXiv:2024.acl-long.19 |
|
yiren-jian/EVLGen/lavis/common/dist_utils.py 643539116e049080 |
unverified |
BSD-3-Clause recorded; this copy not marked cleared · pointer only |