| Generative Retrieval for Unsupervised Text-Based Person Search added by Syntology |
2026-09 (from id) |
Flame-Chasers/GTR/models/vit.py c6ec173f19f5c34d |
ran · our draft was wrong
|
no licence file found · pointer only |
| C²Prompt: Class-aware Client Knowledge Interaction for Federated Continual Learning added by Syntology |
2025-09 (from id) |
zhoujiahuan1991/NeurIPS2025-C2Prompt/models_Cprompt/vision_transformer.py c6ec173f19f5c34d |
ran · our draft was wrong
|
no licence file found · pointer only |
| Towards Anytime Retrieval: A Benchmark for Anytime Person Re-Identification added by Syntology |
2025-09 (from id) |
kw66/AT-ReID/AT-ReID-fast/model/moae.py fb7b918bdc68a83c |
unverified |
no licence file found · pointer only |
| CMRAG: Co-modality-based visual document retrieval and question answering added by Syntology |
2025-09 (from id) |
ChenWangHKU/CMRAG/retriever/model/encoder.py e136986f60c19028 |
unverified |
no licence file found · pointer only |
| Beyond Simple Edits: Composed Video Retrieval with Dense Modifications added by Syntology |
2025-08 (from id) |
OmkarThawakar/BSE-CoVR/src/model/vit.py c6ec173f19f5c34d |
ran · our draft was wrong
|
no licence file found · pointer only |
| Seedream 2.0: A Native Chinese-English Bilingual Image Generation Foundation Model |
10 Mar 2025 |
DYEvaLab/EvalMuse/lavis/models/vit.py c6ec173f19f5c34d |
ran · our draft was wrong
|
Apache-2.0 (permissive) |
| When and How Does CLIP Enable Domain and Compositional Generalization? |
13 Feb 2025 |
salesforce/LAVIS/lavis/models/vit.py c6ec173f19f5c34d |
ran · our draft was wrong
|
BSD-3-Clause (permissive) |
| Semantic-Aligned Adversarial Evolution Triangle for High-Transferability Vision-Language Attack |
4 Nov 2024 |
jiaxiaojunqaq/sa-aet/models/vit.py c6ec173f19f5c34d |
ran · our draft was wrong
|
MIT (permissive) |
| Vector Quantization Prompting for Continual Learning |
27 Oct 2024 |
jiaolifengmi/vq-prompt/models/vit.py c6ec173f19f5c34d |
ran · our draft was wrong
|
MIT (permissive) |
| Multimodal LLM Enhanced Cross-lingual Cross-modal Retrieval |
30 Sep 2024 |
lijiabei-7/leccr/LECCR/clip/model.py 076d66604810aa98 |
ran
fingerprinted |
no licence file found · pointer only |
| HERMES: temporal-coHERent long-forM understanding with Episodes and Semantics |
30 Aug 2024 |
joslefaure/HERMES/lavis/models/vit.py c6ec173f19f5c34d |
ran · our draft was wrong
|
MIT (permissive) |
| EgoCVR: An Egocentric Benchmark for Fine-Grained Composed Video Retrieval |
23 Jul 2024 |
explainableml/egocvr/model/blip/vit.py c6ec173f19f5c34d |
ran · our draft was wrong
|
MIT (permissive) |
| AGLA: Mitigating Object Hallucinations in Large Vision-Language Models with Assembly of Global and Local Attention |
18 Jun 2024 |
lackel/agla/lavis/models/vit.py c6ec173f19f5c34d |
ran · our draft was wrong
|
no licence file found · pointer only |
| One Perturbation is Enough: On Generating Universal Adversarial Perturbations against Vision-Language Pre-training Models |
8 Jun 2024 |
ffhibnese/cpgc_vlp_universal_attacks/models/vit.py c6ec173f19f5c34d |
ran · our draft was wrong
|
no licence file found · pointer only |
| SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension |
25 Apr 2024 |
ailab-cvc/seed-bench/lavis/models/vit.py c6ec173f19f5c34d |
ran · our draft was wrong
|
no licence file found · pointer only |
| Self-Bootstrapped Visual-Language Model for Knowledge Selection and Question Answering |
22 Apr 2024 |
haodongze/self-ksel-qans/lavis/models/vit.py c6ec173f19f5c34d |
ran · our draft was wrong
|
MIT (permissive) |
| Pre-trained Vision and Language Transformers Are Few-Shot Incremental Learners |
2 Apr 2024 |
KHU-AGI/PriViLege/models/vision_transformer.py c6ec173f19f5c34d |
ran · our draft was wrong
|
MIT (permissive) |
| Composed Video Retrieval via Enriched Context and Discriminative Embeddings |
25 Mar 2024 |
omkarthawakar/composed-video-retrieval/src/model/vit.py c6ec173f19f5c34d |
ran · our draft was wrong
|
Apache-2.0 (permissive) |
| Subjective-Aligned Dataset and Metric for Text-to-Video Quality Assessment |
18 Mar 2024 |
qmme/t2vqa/model/vit.py c6ec173f19f5c34d |
ran · our draft was wrong
|
no licence file found · pointer only |
| Can LLMs' Tuning Methods Work in Medical Multimodal Domain? |
11 Mar 2024 |
timmy-chan/miss/models/vit.py c6ec173f19f5c34d |
ran · our draft was wrong
|
no licence file found · pointer only |
| Embodied Understanding of Driving Scenarios |
7 Mar 2024 |
opendrivelab/elm/lavis/models/vit.py c6ec173f19f5c34d |
ran · our draft was wrong
|
no licence file found · pointer only |
| VisLingInstruct: Elevating Zero-Shot Learning in Multi-Modal Language Models with Autonomous Instruction Optimization |
12 Feb 2024 |
zhudongsheng75/vislinginstruct/vislinginstruct/models/vit.py c6ec173f19f5c34d |
ran · our draft was wrong
|
no licence file found · pointer only |
| CREMA: Generalizable and Efficient Video-Language Reasoning via Multimodal Modular Fusion |
8 Feb 2024 |
Yui010206/CREMA/lavis/models/vit.py c6ec173f19f5c34d |
ran · our draft was wrong
|
BSD-3-Clause (permissive) |
| Inducing High Energy-Latency of Large Vision-Language Models with Verbose Images |
20 Jan 2024 |
KuofengGao/Verbose_Images/lavis/models/vit.py c6ec173f19f5c34d |
ran · our draft was wrong
|
no licence file found · pointer only |
| Moonshot: Towards Controllable Video Generation and Editing with Multimodal Conditions |
3 Jan 2024 |
salesforce/lavis/lavis/models/vit.py c6ec173f19f5c34d |
ran · our draft was wrong
|
BSD-3-Clause (permissive) |
| When Parameter-efficient Tuning Meets General-purpose Vision-language Models |
16 Dec 2023 |
melonking32/petal/lavis/models/vit.py c6ec173f19f5c34d |
ran · our draft was wrong
|
no licence file found · pointer only |
| ToViLaG: Your Visual-Language Generative Model is Also An Evildoer |
13 Dec 2023 |
victorup/ToViLaG/method/BLIP/models/vit.py c6ec173f19f5c34d |
ran · our draft was wrong
|
no licence file found · pointer only |
| Image Content Generation with Causal Reasoning |
12 Dec 2023 |
ieit-agi/mix-shannon/lavis/models/vit.py c6ec173f19f5c34d |
ran · our draft was wrong
|
Apache-2.0 (permissive) |
| Recursive Visual Programming |
4 Dec 2023 |
para-lost/rvp/base_models/tcl/tcl_vit.py 4262ed39e3ce5c97 |
ran
|
licence not identified · pointer only |
| Synthesize, Diagnose, and Optimize: Towards Fine-Grained Vision-Language Understanding |
30 Nov 2023 |
wjpoom/spec/spec/models/blip_utils/vit.py c6ec173f19f5c34d |
ran · our draft was wrong
|
no licence file found · pointer only |
| X-InstructBLIP: A Framework for aligning X-Modal instruction-aware representations to LLMs and Emergent Cross-modal Reasoning |
30 Nov 2023 |
artemisp/lavis-xinstructblip/lavis/models/vit.py c6ec173f19f5c34d |
ran · our draft was wrong
|
BSD-3-Clause (permissive) |
| MultiDelete for Multimodal Machine Unlearning |
18 Nov 2023 |
chengjiali/Multimodal-Machine-Unlearning/lavis/models/vit.py c6ec173f19f5c34d |
ran · our draft was wrong
|
no licence file found · pointer only |
| Rephrase, Augment, Reason: Visual Grounding of Questions for Vision-Language Models |
9 Oct 2023 |
archiki/repare/Lavis/lavis/models/vit.py c6ec173f19f5c34d |
ran · our draft was wrong
|
MIT (permissive) |
| Expedited Training of Visual Conditioned Language Generation via Redundancy Reduction |
5 Oct 2023 |
yiren-jian/evlgen/lavis/models/vit.py c6ec173f19f5c34d |
ran · our draft was wrong
|
BSD-3-Clause (permissive) |
| Making LLaMA SEE and Draw with SEED Tokenizer |
2 Oct 2023 |
ailab-cvc/seed/models/seed_qformer/vit.py c6ec173f19f5c34d |
ran · our draft was wrong
|
no licence file found · pointer only |
| Detecting and Grounding Multi-Modal Media Manipulation and Beyond |
25 Sep 2023 |
rshaojimmy/multimodal-deepfake/models/vit.py c6ec173f19f5c34d |
ran · our draft was wrong
|
no licence file found · pointer only |
| VIGC: Visual Instruction Generation and Correction |
24 Aug 2023 |
opendatalab/vigc/vigc/models/vit.py c6ec173f19f5c34d |
ran · our draft was wrong
|
Apache-2.0 (permissive) |
| Simple Baselines for Interactive Video Retrieval with Questions and Answers |
21 Aug 2023 |
kevinliang888/IVR-QA-baselines/models/vit.py c6ec173f19f5c34d |
ran · our draft was wrong
|
no licence file found · pointer only |
| CTP: Towards Vision-Language Continual Pretraining via Compatible Momentum Contrast and Topology Preservation |
14 Aug 2023 |
kevinlight831/ctp/models/vit.py c6ec173f19f5c34d |
ran · our draft was wrong
|
MIT (permissive) |
| Set-level Guidance Attack: Boosting Adversarial Transferability of Vision-Language Pre-training Models |
26 Jul 2023 |
Zoky-2020/Set-level_Guidance_Attack/models/vit.py c6ec173f19f5c34d |
ran · our draft was wrong
|
MIT (permissive) |
| 3D-LLM: Injecting the 3D World into Large Language Models |
24 Jul 2023 |
umass-foundation-model/3d-llm/3DLLM_BLIP2-base/lavis/models/vit.py c6ec173f19f5c34d |
ran · our draft was wrong
|
MIT recorded; this copy not marked cleared · pointer only |
| FigCaps-HF: A Figure-to-Caption Generative Framework and Benchmark with Human Feedback |
20 Jul 2023 |
figcapshf/figcapshf/models/vit.py c6ec173f19f5c34d |
ran · our draft was wrong
|
no licence file found · pointer only |
| RaSa: Relation and Sensitivity Aware Representation Learning for Text-based Person Search |
23 May 2023 |
Flame-Chasers/RaSa/models/vit.py c6ec173f19f5c34d |
ran · our draft was wrong
|
MIT (permissive) |
| An Empirical Study of Multimodal Model Merging |
28 Apr 2023 |
identical code first harvested elsewhere c6ec173f19f5c34d |
ran · our draft was wrong
|
licence of this copy not recorded |
| Unmasked Teacher: Towards Training-Efficient Video Foundation Models |
28 Mar 2023 |
opengvlab/unmasked_teacher/multi_modality/models/utils.py dbd049b02a63c3a1 |
unverified |
MIT (permissive) |
| Dynamic Graph Enhanced Contrastive Learning for Chest X-ray Report Generation |
18 Mar 2023 |
mlii0117/dcl/models/vit.py c6ec173f19f5c34d |
ran · our draft was wrong
|
MIT (permissive) |
| UniAdapter: Unified Parameter-Efficient Transfer Learning for Cross-modal Modeling |
13 Feb 2023 |
rerv/uniadapter/models/vit.py c6ec173f19f5c34d |
ran · our draft was wrong
|
BSD-3-Clause recorded; this copy not marked cleared · pointer only |
| mPLUG-2: A Modularized Multi-modal Foundation Model Across Text, Image and Video |
1 Feb 2023 |
X-PLUG/mPLUG-2/models/vit.py c6ec173f19f5c34d |
ran · our draft was wrong
|
Apache-2.0 (permissive) |
| Self-supervised vision-language pretraining for Medical visual question answering |
24 Nov 2022 |
pengfeiliheu/m2i2/models/mae.py c6ec173f19f5c34d |
ran · our draft was wrong
|
MIT (permissive) |
| ConStruct-VL: Data-Free Continual Structured VL Concepts Learning |
17 Nov 2022 |
jamessealesmith/construct-vl/models/vit.py c6ec173f19f5c34d |
ran · our draft was wrong
|
MIT recorded; this copy not marked cleared · pointer only |
| Plug-and-Play VQA: Zero-shot VQA by Conjoining Large Pretrained Models with Zero Training |
17 Oct 2022 |
Tzoulio/Large_Models_Dialogue_for_Active_Perception/llm-vqa_dialogue/lavis/models/vit.py c6ec173f19f5c34d |
ran · our draft was wrong
|
MIT recorded; this copy not marked cleared · pointer only |
| IDEA: Increasing Text Diversity via Online Multi-Label Recognition for Vision-Language Pre-training |
12 Jul 2022 |
xinyu1205/idea-pytorch/models/vit.py c6ec173f19f5c34d |
ran · our draft was wrong
|
MIT (permissive) |
| PEVL: Position-enhanced Pre-training and Prompt Tuning for Vision-language Models |
23 May 2022 |
thunlp/PEVL/models/vit.py c6ec173f19f5c34d |
ran · our draft was wrong
|
MIT (permissive) |
| Open Vocabulary Object Detection with Pseudo Bounding-Box Labels |
18 Nov 2021 |
salesforce/pb-ovd/ALBEF/models/vit.py c6ec173f19f5c34d |
ran · our draft was wrong
|
BSD-3-Clause (permissive) |
| Fashion IQ: A New Dataset Towards Retrieving Images by Natural Language Feedback |
30 May 2019 |
hssip/fashionsap/models/vit.py c6ec173f19f5c34d |
ran · our draft was wrong
|
BSD-3-Clause (permissive) |
| arXiv:aaai_35495 |
|
salesforce/BLIP/models/vit.py c6ec173f19f5c34d |
ran · our draft was wrong
|
BSD-3-Clause (permissive) |
| arXiv:aaai_28415 |
|
wuhy68/p-Adapter/eff_blip/vit.py c6ec173f19f5c34d |
ran · our draft was wrong
|
Apache-2.0 (permissive) |
| arXiv:Cheng_VindLU_A_Recipe_for_Effective_Video-and-Language_Pretraining_CVPR_2023_paper |
|
klauscc/VindLU/models/utils.py dbd049b02a63c3a1 |
unverified |
MIT (permissive) |
| arXiv:2024.acl-long.19 |
|
yiren-jian/EVLGen/lavis/models/vit.py c6ec173f19f5c34d |
ran · our draft was wrong
|
BSD-3-Clause recorded; this copy not marked cleared · pointer only |