| Decoupled Vision-Language System for Multimodal Understanding and Generation added by Syntology |
2026-08 (from id) |
YifanXu74/Libra/libra/common/dist_utils.py 98589643273920ba |
ran
|
Apache-2.0 (permissive) |
| CSMCIR: CoT-Enhanced Symmetric Alignment with Memory Bank for Composed Image Retrieval added by Syntology |
2026-01 (from id) |
qzp2018/CSMCIR/src/lavis/common/dist_utils.py 98589643273920ba |
ran
|
no licence file found · pointer only |
| arXiv:2507.12001 |
2025-07 (from id) |
wslh852/AUBlendNet/base/utilities.py 06fca4c705811abe |
ran · violated contract
|
MIT (permissive) |
| TimeChat-Online: 80% Visual Tokens are Naturally Redundant in Streaming Videos |
24 Apr 2025 |
renshuhuai-andy/timechat/timechat/common/dist_utils.py 98589643273920ba |
ran
|
BSD-3-Clause (permissive) |
| When and How Does CLIP Enable Domain and Compositional Generalization? |
13 Feb 2025 |
salesforce/LAVIS/lavis/common/dist_utils.py 98589643273920ba |
ran
|
BSD-3-Clause (permissive) |
| Image Over Text: Transforming Formula Recognition Evaluation with Character Detection Matching |
5 Sep 2024 |
opendatalab/unimernet/unimernet/common/dist_utils.py 98589643273920ba |
ran
|
Apache-2.0 (permissive) |
| HERMES: temporal-coHERent long-forM understanding with Episodes and Semantics |
30 Aug 2024 |
joslefaure/HERMES/lavis/common/dist_utils.py 98589643273920ba |
ran
|
MIT (permissive) |
| SZTU-CMU at MER2024: Improving Emotion-LLaMA with Conv-Attention for Multimodal Emotion Recognition |
20 Aug 2024 |
zebangcheng/emotion-llama/minigpt4/common/dist_utils.py 98589643273920ba |
ran
|
BSD-3-Clause (permissive) |
| Cross-modality Information Check for Detecting Jailbreaking in Multimodal Large Language Models |
31 Jul 2024 |
pandragonxiii/cider/code/models/minigpt4/common/dist_utils.py 98589643273920ba |
ran
|
Apache-2.0 (permissive) |
| VRSBench: A Versatile Vision-Language Benchmark Dataset for Remote Sensing Image Understanding |
18 Jun 2024 |
lavender105/rsgpt/rsgpt/common/dist_utils.py 98589643273920ba |
ran
|
no licence file found · pointer only |
| Hawk: Learning to Understand Open-World Video Anomalies |
27 May 2024 |
jqtangust/hawk/hawk/common/dist_utils.py 98589643273920ba |
ran
|
no licence file found · pointer only |
| Libra: Building Decoupled Vision System on Large Language Models |
16 May 2024 |
yifanxu74/libra/libra/common/dist_utils.py 98589643273920ba |
ran
|
Apache-2.0 (permissive) |
| Self-Bootstrapped Visual-Language Model for Knowledge Selection and Question Answering |
22 Apr 2024 |
haodongze/self-ksel-qans/lavis/common/dist_utils.py 98589643273920ba |
ran
|
MIT (permissive) |
| MTP: Advancing Remote Sensing Foundation Model via Multi-Task Pretraining |
20 Mar 2024 |
vitae-transformer/mtp/Multi-Task_Pretrain/main_pretrain.py e26853276bb3dd32 |
ran · violated contract
|
MIT (permissive) |
| Embodied Understanding of Driving Scenarios |
7 Mar 2024 |
opendrivelab/elm/lavis/common/dist_utils.py 98589643273920ba |
ran
|
no licence file found · pointer only |
| VisLingInstruct: Elevating Zero-Shot Learning in Multi-Modal Language Models with Autonomous Instruction Optimization |
12 Feb 2024 |
zhudongsheng75/vislinginstruct/vislinginstruct/common/dist_utils.py 98589643273920ba |
ran
|
no licence file found · pointer only |
| GraphTranslator: Aligning Graph Model to Large Language Model for Open-ended Tasks |
11 Feb 2024 |
alibaba/graphtranslator/Translator/common/dist_utils.py 98589643273920ba |
ran
|
BSD-3-Clause (permissive) |
| CREMA: Generalizable and Efficient Video-Language Reasoning via Multimodal Modular Fusion |
8 Feb 2024 |
Yui010206/CREMA/lavis/common/dist_utils.py 98589643273920ba |
ran
|
BSD-3-Clause (permissive) |
| Inducing High Energy-Latency of Large Vision-Language Models with Verbose Images |
20 Jan 2024 |
KuofengGao/Verbose_Images/lavis/common/dist_utils.py 98589643273920ba |
ran
|
no licence file found · pointer only |
| Moonshot: Towards Controllable Video Generation and Editing with Multimodal Conditions |
3 Jan 2024 |
salesforce/lavis/lavis/common/dist_utils.py 98589643273920ba |
ran
|
BSD-3-Clause (permissive) |
| When Parameter-efficient Tuning Meets General-purpose Vision-language Models |
16 Dec 2023 |
melonking32/petal/lavis/common/dist_utils.py 98589643273920ba |
ran
|
no licence file found · pointer only |
| EgoPlan-Bench: Benchmarking Multimodal Large Language Models for Human-Level Planning |
11 Dec 2023 |
chenyi99/egoplan/src/video_llama/video_llama/common/dist_utils.py 98589643273920ba |
ran
|
BSD-3-Clause (permissive) |
| Aligning and Prompting Everything All at Once for Universal Visual Perception |
4 Dec 2023 |
earth-insights/ClassTrans/src/utils.py 3baca7953e57a8e0 |
unverified |
MIT (permissive) |
| X-InstructBLIP: A Framework for aligning X-Modal instruction-aware representations to LLMs and Emergent Cross-modal Reasoning |
30 Nov 2023 |
artemisp/lavis-xinstructblip/lavis/common/dist_utils.py 98589643273920ba |
ran
|
BSD-3-Clause (permissive) |
| MarineGPT: Unlocking Secrets of Ocean to the Public |
20 Oct 2023 |
hkust-vgd/marinegpt/marinegpt/common/dist_utils.py 98589643273920ba |
ran
|
no licence file found · pointer only |
| Fine-grained Audio-Visual Joint Representations for Multimodal Large Language Models |
9 Oct 2023 |
the-anonymous-bs/favor/video_llama/common/dist_utils.py 98589643273920ba |
ran
|
no licence file found · pointer only |
| Sentence-level Prompts Benefit Composed Image Retrieval |
9 Oct 2023 |
chunmeifeng/sprc/src/lavis/common/dist_utils.py 98589643273920ba |
ran
|
no licence file found · pointer only |
| Expedited Training of Visual Conditioned Language Generation via Redundancy Reduction |
5 Oct 2023 |
yiren-jian/evlgen/lavis/common/dist_utils.py 98589643273920ba |
ran
|
BSD-3-Clause (permissive) |
| Making LLaMA SEE and Draw with SEED Tokenizer |
2 Oct 2023 |
ailab-cvc/seed/models/seed_qformer/utils.py 98589643273920ba |
ran
|
no licence file found · pointer only |
| Sparkles: Unlocking Chats Across Multiple Images for Multimodal Instruction-Following Models |
31 Aug 2023 |
HYPJUDY/Sparkles/sparkles/common/dist_utils.py 98589643273920ba |
ran
|
BSD-3-Clause (permissive) |
| VIGC: Visual Instruction Generation and Correction |
24 Aug 2023 |
opendatalab/vigc/vigc/common/dist_utils.py 98589643273920ba |
ran
|
Apache-2.0 (permissive) |
| Fine-tuning Multimodal LLMs to Follow Zero-shot Demonstrative Instructions |
8 Aug 2023 |
DCDmllm/Cheetah/Cheetah/cheetah/common/dist_utils.py 98589643273920ba |
ran
|
BSD-3-Clause (permissive) |
| BuboGPT: Enabling Visual Grounding in Multi-Modal LLMs |
17 Jul 2023 |
magic-research/bubogpt/bubogpt/common/dist_utils.py 98589643273920ba |
ran
|
BSD-3-Clause (permissive) |
| Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding |
5 Jun 2023 |
damo-nlp-sg/video-llama/video_llama/common/dist_utils.py 98589643273920ba |
ran
|
BSD-3-Clause recorded; this copy not marked cleared · pointer only |
| ChatBridge: Bridging Modalities with Large Language Model as a Language Catalyst |
25 May 2023 |
joez17/chatbridge/chatbridge/common/dist_utils.py 98589643273920ba |
ran
|
BSD-3-Clause (permissive) |
| DetGPT: Detect What You Need via Reasoning |
23 May 2023 |
optimalscale/detgpt/detgpt/common/dist_utils.py 98589643273920ba |
ran
|
BSD-3-Clause (permissive) |
| X-LLM: Bootstrapping Advanced Large Language Models by Treating Multi-Modalities as Foreign Languages |
7 May 2023 |
phellonchen/x-llm/xllm/common/dist_utils.py 98589643273920ba |
ran
|
Apache-2.0 (permissive) |
| CodeTalker: Speech-Driven 3D Facial Animation with Discrete Motion Prior |
6 Jan 2023 |
Doubiiu/CodeTalker/base/utilities.py 06fca4c705811abe |
ran · violated contract
|
MIT (permissive) |
| TransNet V2: An effective deep network architecture for fast shot transition detection |
11 Aug 2020 |
shallwe999/TransNetV2-SBD-Visualize/transnetv2.py e6c126c38a88b3c8 |
unverified |
MIT (permissive) |
| HITNet: Hierarchical Iterative Tile Refinement Network for Real-time Stereo Matching |
23 Jul 2020 |
meteorshowers/X-StereoLab/tools/train_net_disp.py 06fca4c705811abe |
ran · violated contract
|
MIT (permissive) |
| arXiv:Xu_Adaptive_Multi-Modal_Cross-Entropy_Loss_for_Stereo_Matching_CVPR_2024_paper |
|
xxxupeng/ADL/train_DDP.py 06fca4c705811abe |
ran · violated contract
|
MIT (permissive) |
| arXiv:Hajimiri_A_Strong_Baseline_for_Generalized_Few-Shot_Semantic_Segmentation_CVPR_2023_paper |
|
sinahmr/DIaM/src/util.py 3baca7953e57a8e0 |
unverified |
MIT (permissive) |
| arXiv:Chen_DViN_Dynamic_Visual_Routing_Network_for_Weakly_Supervised_Referring_Expression_CVPR_2025_paper |
|
XxFChen/DViN/utils/distributed.py 5a7f17ee3741438b |
unverified |
Apache-2.0 (permissive) |
| arXiv:2024.findings-emnlp.803 |
|
PandragonXIII/CIDER/code/models/minigpt4/common/dist_utils.py 98589643273920ba |
ran
|
Apache-2.0 (permissive) |
| arXiv:2024.acl-long.19 |
|
yiren-jian/EVLGen/lavis/common/dist_utils.py 98589643273920ba |
ran
|
BSD-3-Clause recorded; this copy not marked cleared · pointer only |