| Bridging Vision Foundation Model Priors with CLIP for Spatial-aware Few-shot Anomaly Detection in Medical Images added by Syntology |
2026-09 (from id) |
JuzhengMiao/Spatial-FAD/CLIP/model.py dcd422d66b0581d8 |
ran · our draft was wrong
|
MIT (permissive) |
| Expert Knowledge-Guided Decision Calibration for Accurate Fine-Grained Tree Species Classification added by Syntology |
2026-01 (from id) |
WHU-USI3DV/TreeCLS/models/open_clip/model.py dcd422d66b0581d8 |
ran · our draft was wrong
|
no licence file found · pointer only |
| Streamline Without Sacrifice -- Squeeze out Computation Redundancy in LMM |
21 May 2025 |
identical code first harvested elsewhere 61d9ad9efd32efee |
ran · our draft was wrong
|
licence of this copy not recorded |
| Streamline Without Sacrifice -- Squeeze out Computation Redundancy in LMM |
21 May 2025 |
identical code first harvested elsewhere dcd422d66b0581d8 |
ran · our draft was wrong
|
licence of this copy not recorded |
| CAFe: Unifying Representation and Generation with Contrastive-Autoregressive Finetuning |
25 Mar 2025 |
identical code first harvested elsewhere dcd422d66b0581d8 |
ran · our draft was wrong
|
licence of this copy not recorded |
| LongProLIP: A Probabilistic Vision-Language Model with Long Context Text |
11 Mar 2025 |
identical code first harvested elsewhere dcd422d66b0581d8 |
ran · our draft was wrong
|
licence of this copy not recorded |
| LongProLIP: A Probabilistic Vision-Language Model with Long Context Text |
11 Mar 2025 |
identical code first harvested elsewhere 0529ec3d12d83d1c |
unverified |
licence of this copy not recorded |
| GeoLangBind: Unifying Earth Observation with Agglomerative Vision-Language Foundation Models |
8 Mar 2025 |
xiong-zhitong/geolb-siglip/open_clip/src/open_clip/model.py dcd422d66b0581d8 |
ran · our draft was wrong
|
Apache-2.0 (permissive) |
| Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step |
23 Jan 2025 |
identical code first harvested elsewhere dcd422d66b0581d8 |
ran · our draft was wrong
|
licence of this copy not recorded |
| Post-hoc Probabilistic Vision-Language Models |
8 Dec 2024 |
naver-ai/prolip/src/prolip/model.py 0529ec3d12d83d1c |
unverified |
licence not identified · pointer only |
| Hidden in the Noise: Two-Stage Robust Watermarking for Images |
5 Dec 2024 |
Kasraarabi/Hidden-in-the-Noise/open_clip/model.py dcd422d66b0581d8 |
ran · our draft was wrong
|
no licence file found · pointer only |
| FLAIR: VLM with Fine-grained Language-informed Image Representations |
4 Dec 2024 |
explainableml/flair/src/flair/model.py dcd422d66b0581d8 |
ran · our draft was wrong
|
no licence file found · pointer only |
| DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models |
22 Nov 2024 |
identical code first harvested elsewhere dcd422d66b0581d8 |
ran · our draft was wrong
|
licence of this copy not recorded |
| On Erroneous Agreements of CLIP Image Embeddings |
7 Nov 2024 |
lst627/CLIP-Embeds/open_clip/src/open_clip/model.py dcd422d66b0581d8 |
ran · our draft was wrong
|
no licence file found · pointer only |
| ROBIN: Robust and Invisible Watermarks for Diffusion Models with Adversarial Optimization |
6 Nov 2024 |
Hannah1102/ROBIN/open_clip/model.py dcd422d66b0581d8 |
ran · our draft was wrong
|
no licence file found · pointer only |
| Benchmarking Vision Language Model Unlearning via Fictitious Facial Identity Dataset |
5 Nov 2024 |
safolab-wisc/fiubench/utils.py dcd422d66b0581d8 |
ran · our draft was wrong
|
no licence file found · pointer only |
| Classification Done Right for Vision-Language Pre-Training |
5 Nov 2024 |
x-cls/superclass/opencls/open_clip/model.py 1f22cc0143fbf7d1 |
ran · our draft was wrong
|
Apache-2.0 (permissive) |
| Beyond Accuracy: Ensuring Correct Predictions With Correct Rationales |
31 Oct 2024 |
deep-real/DCP/utils/model.py dcd422d66b0581d8 |
ran · our draft was wrong
|
no licence file found · pointer only |
| Probabilistic Language-Image Pre-Training |
24 Oct 2024 |
identical code first harvested elsewhere dcd422d66b0581d8 |
ran · our draft was wrong
|
licence of this copy not recorded |
| What If the Input is Expanded in OOD Detection? |
24 Oct 2024 |
tmlr-group/CoVer/open_clip/model.py dcd422d66b0581d8 |
ran · our draft was wrong
|
no licence file found · pointer only |
| Personalized Image Generation with Large Multimodal Models |
18 Oct 2024 |
yiyanxu/pigeon/Pigeon/models/modeling_visual_encoder.py dcd422d66b0581d8 |
ran · our draft was wrong
|
no licence file found · pointer only |
| TULIP: Token-length Upgraded CLIP |
13 Oct 2024 |
ivonajdenkoska/tulip/open_clip/model.py dcd422d66b0581d8 |
ran · our draft was wrong
|
Apache-2.0 (permissive) |
| SePPO: Semi-Policy Preference Optimization for Diffusion Alignment |
7 Oct 2024 |
dwanzhang-ai/seppo/utils/open_clip/model.py dcd422d66b0581d8 |
ran · our draft was wrong
|
no licence file found · pointer only |
| SegEarth-OV: Towards Training-Free Open-Vocabulary Segmentation for Remote Sensing Images |
2 Oct 2024 |
likyoo/SegEarth-OV/open_clip/model.py dcd422d66b0581d8 |
ran · our draft was wrong
|
no licence file found · pointer only |
| PointAD: Comprehending 3D Anomalies from Points and Pixels for Zero-shot 3D Anomaly Detection |
1 Oct 2024 |
zqhang/accurate-winclip-pytorch/src/open_clip/model.py 1bc68318bf26bdf0 |
ran · our draft was wrong
|
MIT (permissive) |
| Embedding Geometries of Contrastive Language-Image Pre-Training |
19 Sep 2024 |
eify/open_clip/src/open_clip/model.py dcd422d66b0581d8 |
ran · our draft was wrong
|
no licence file found · pointer only |
| ECG-Chat: A Large ECG-Language Model for Cardiac Disease Diagnosis |
16 Aug 2024 |
YubaoZhao/ECG-Chat/open_clip/open_clip/model.py dcd422d66b0581d8 |
ran · our draft was wrong
|
no licence file found · pointer only |
| ProxyCLIP: Proxy Attention Improves CLIP for Open-Vocabulary Segmentation |
9 Aug 2024 |
mc-lan/proxyclip/open_clip/model.py dcd422d66b0581d8 |
ran · our draft was wrong
|
no licence file found · pointer only |
| VSSD: Vision Mamba with Non-Causal State Space Duality |
26 Jul 2024 |
YuHengsss/Trident/open_clip/model.py dcd422d66b0581d8 |
ran · our draft was wrong
|
Apache-2.0 (permissive) |
| AdaCLIP: Adapting CLIP with Hybrid Learnable Prompts for Zero-Shot Anomaly Detection |
22 Jul 2024 |
caoyunkang/adaclip/method/clip_model.py dcd422d66b0581d8 |
ran · our draft was wrong
|
MIT (permissive) |
| VideoEval: Comprehensive Benchmark Suite for Low-Cost Evaluation of Video Foundation Model |
9 Jul 2024 |
leexinhao/VideoEval/VidTAB_Zeroshot/eva_clip/model.py dcd422d66b0581d8 |
ran · our draft was wrong
|
no licence file found · pointer only |
| Mitigate the Gap: Investigating Approaches for Improving Cross-Modal Alignment in CLIP |
25 Jun 2024 |
sarahesl/alignclip/align_clip/model.py dcd422d66b0581d8 |
ran · our draft was wrong
|
no licence file found · pointer only |
| ClawMachine: Learning to Fetch Visual Tokens for Referential Comprehension |
17 Jun 2024 |
martian422/ClawMachine/ClawMachine/model/multimodal_encoder/add_modeling_visual_encoder.py dcd422d66b0581d8 |
ran · our draft was wrong
|
no licence file found · pointer only |
| Vision Model Pre-training on Interleaved Image-Text Data via Latent Compression Learning |
11 Jun 2024 |
opengvlab/lcl/src/open_clip/model.py dcd422d66b0581d8 |
ran · our draft was wrong
|
MIT (permissive) |
| VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs |
11 Jun 2024 |
damo-nlp-sg/inf-clip/inf_clip/models/clip_arch.py dcd422d66b0581d8 |
ran · our draft was wrong
|
Apache-2.0 (permissive) |
| PuLID: Pure and Lightning ID Customization via Contrastive Alignment |
24 Apr 2024 |
ToTheBeginning/PuLID/eva_clip/model.py 61d9ad9efd32efee |
ran · our draft was wrong
|
Apache-2.0 (permissive) |
| PuLID: Pure and Lightning ID Customization via Contrastive Alignment |
24 Apr 2024 |
zsxkib/PuLID/eva_clip/model.py dcd422d66b0581d8 |
ran · our draft was wrong
|
Apache-2.0 (permissive) |
| RingID: Rethinking Tree-Ring Watermarking for Enhanced Multi-Key Identification |
22 Apr 2024 |
showlab/ringid/open_clip/model.py dcd422d66b0581d8 |
ran · our draft was wrong
|
no licence file found · pointer only |
| FiLo: Zero-Shot Anomaly Detection by Fine-Grained Description and High-Quality Localization |
21 Apr 2024 |
casia-iva-lab/filo/models/vv_open_clip/model.py dcd422d66b0581d8 |
ran · our draft was wrong
|
Apache-2.0 (permissive) |
| PromptAD: Learning Prompts with only Normal Samples for Few-Shot Anomaly Detection |
8 Apr 2024 |
funz-0/promptad/PromptAD/CLIPAD/model.py dcd422d66b0581d8 |
ran · our draft was wrong
|
Unlicense (permissive) |
| Gaussian Shading: Provable Performance-Lossless Image Watermarking for Diffusion Models |
7 Apr 2024 |
bsmhmmlf/Gaussian-Shading/open_clip/model.py dcd422d66b0581d8 |
ran · our draft was wrong
|
MIT (permissive) |
| ViTamin: Designing Scalable Vision Models in the Vision-Language Era |
2 Apr 2024 |
beckschen/vitamin/ViTamin/open_clip/model.py dcd422d66b0581d8 |
ran · our draft was wrong
|
Apache-2.0 (permissive) |
| DreamLIP: Language-Image Pre-training with Long Captions |
25 Mar 2024 |
zyf0619sjtu/DreamLIP/open_clip/model.py dcd422d66b0581d8 |
ran · our draft was wrong
|
CC-BY-4.0 · pointer only |
| UrbanVLP: Multi-Granularity Vision-Language Pretraining for Urban Socioeconomic Indicator Prediction |
25 Mar 2024 |
citymind-lab/urbanvlp/open_clip_mine/model.py dcd422d66b0581d8 |
ran · our draft was wrong
|
no licence file found · pointer only |
| Long-CLIP: Unlocking the Long-Text Capability of CLIP |
22 Mar 2024 |
beichenzbc/long-clip/open_clip_long/model.py dcd422d66b0581d8 |
ran · our draft was wrong
|
Apache-2.0 (permissive) |
| Adapting Visual-Language Models for Generalizable Anomaly Detection in Medical Images |
19 Mar 2024 |
mediabrain-sjtu/mvfa-ad/CLIP/model.py dcd422d66b0581d8 |
ran · our draft was wrong
|
MIT (permissive) |
| Toward Generalist Anomaly Detection via In-context Residual Learning with Few-shot Sample Prompts |
11 Mar 2024 |
mala-lab/WinCLIP/open_clip/model.py dcd422d66b0581d8 |
ran · our draft was wrong
|
GPL-3.0 (copyleft) · pointer only |
| Video-LaVIT: Unified Video-Language Pre-training with Decoupled Visual-Motional Tokenization |
5 Feb 2024 |
jy0205/lavit/LaVIT/models/modeling_visual_encoder.py dcd422d66b0581d8 |
ran · our draft was wrong
|
no licence file found · pointer only |
| Open-Vocabulary SAM: Segment and Recognize Twenty-thousand Classes Interactively |
5 Jan 2024 |
harboryuan/ovsam/ext/open_clip/model.py dcd422d66b0581d8 |
ran · our draft was wrong
|
no licence file found · pointer only |
| SkyScript: A Large and Semantically Diverse Vision-Language Dataset for Remote Sensing |
20 Dec 2023 |
wangzhecheng/skyscript/src/open_clip/model.py dcd422d66b0581d8 |
ran · our draft was wrong
|
MIT (permissive) |
| InstructTA: Instruction-Tuned Targeted Attack for Large Vision-Language Models |
4 Dec 2023 |
xunguangwang/instructta/EVA-CLIP/rei/eva_clip/model.py dcd422d66b0581d8 |
ran · our draft was wrong
|
no licence file found · pointer only |
| BioCLIP: A Vision Foundation Model for the Tree of Life |
30 Nov 2023 |
identical code first harvested elsewhere dcd422d66b0581d8 |
ran · our draft was wrong
|
licence of this copy not recorded |
| Diffusion Model Alignment Using Direct Preference Optimization |
21 Nov 2023 |
SalesforceAIResearch/DiffusionDPO/utils/open_clip/model.py dcd422d66b0581d8 |
ran · our draft was wrong
|
Apache-2.0 recorded; this copy not marked cleared · pointer only |
| OtterHD: A High-Resolution Multi-modality Model |
7 Nov 2023 |
luodian/otter/pipeline/benchmarks/public_datasets_suite/models/otter.py dcd422d66b0581d8 |
ran · our draft was wrong
|
MIT (permissive) |
| AnomalyCLIP: Object-agnostic Prompt Learning for Zero-shot Anomaly Detection |
29 Oct 2023 |
zqhang/WinCLIP-pytorch/src/open_clip/model.py dcd422d66b0581d8 |
ran · our draft was wrong
|
MIT (permissive) |
| Improved Baselines with Visual Instruction Tuning |
5 Oct 2023 |
identical code first harvested elsewhere dcd422d66b0581d8 |
ran · our draft was wrong
|
licence of this copy not recorded |
| Understanding and Mitigating the Label Noise in Pre-training on Downstream Tasks |
29 Sep 2023 |
Hhhhhhao/Noisy-Model-Learning/open_clip/open_clip/model.py dcd422d66b0581d8 |
ran · our draft was wrong
|
no licence file found · pointer only |
| Bootstrap Fine-Grained Vision-Language Alignment for Unified Zero-Shot Anomaly Localization |
30 Aug 2023 |
hq-deng/AnoVL/open_clip/model.py dcd422d66b0581d8 |
ran · our draft was wrong
|
MIT (permissive) |
| CLIP-KD: An Empirical Study of CLIP Model Distillation |
24 Jul 2023 |
winycg/clip-kd/src/open_clip/model.py dcd422d66b0581d8 |
ran · our draft was wrong
|
no licence file found · pointer only |
| Diff-Foley: Synchronized Video-to-Audio Synthesis with Latent Diffusion Models |
29 Jun 2023 |
identical code first harvested elsewhere dcd422d66b0581d8 |
ran · our draft was wrong
|
licence of this copy not recorded |
| APRIL-GAN: A Zero-/Few-Shot Anomaly Classification and Segmentation Method for CVPR 2023 VAND Workshop Challenge Tracks 1&2: 1st Place on Zero-shot AD and 4th Place on Few-shot AD |
27 May 2023 |
bychelsea/vand-april-gan/open_clip/model.py dcd422d66b0581d8 |
ran · our draft was wrong
|
MIT (permissive) |
| Unicom: Universal and Compact Representation Learning for Image Retrieval |
12 Apr 2023 |
identical code first harvested elsewhere dcd422d66b0581d8 |
ran · our draft was wrong
|
licence of this copy not recorded |
| WinCLIP: Zero-/Few-Shot Anomaly Classification and Segmentation |
26 Mar 2023 |
caoyunkang/WinClip/WinCLIP/CLIPAD/model.py dcd422d66b0581d8 |
ran · our draft was wrong
|
MIT (permissive) |
| Hard Prompts Made Easy: Gradient-Based Discrete Optimization for Prompt Tuning and Discovery |
7 Feb 2023 |
YuxinWenRick/hard-prompts-made-easy/open_clip/model.py dcd422d66b0581d8 |
ran · our draft was wrong
|
MIT (permissive) |
| Learning Customized Visual Models with Retrieval-Augmented Knowledge |
17 Jan 2023 |
microsoft/react/react_customization/src/open_clip/model.py dcd422d66b0581d8 |
ran · our draft was wrong
|
MIT (permissive) |
| Reproducible scaling laws for contrastive language-image learning |
14 Dec 2022 |
identical code first harvested elsewhere dcd422d66b0581d8 |
ran · our draft was wrong
|
licence of this copy not recorded |
| LAION-5B: An open large-scale dataset for training next generation image-text models |
16 Oct 2022 |
identical code first harvested elsewhere dcd422d66b0581d8 |
ran · our draft was wrong
|
licence of this copy not recorded |
| Learning Transferable Visual Models From Natural Language Supervision |
26 Feb 2021 |
identical code first harvested elsewhere dcd422d66b0581d8 |
ran · our draft was wrong
|
licence of this copy not recorded |
| arXiv:Ma_ReMP-AD_Retrieval-enhanced_Multi-modal_Prompt_Fusion_for_Few-Shot_Industrial_Visual_Anomaly_ICCV_2025_paper |
|
cshcma/ReMP-AD/open_clip/model.py dcd422d66b0581d8 |
ran · our draft was wrong
|
MIT (permissive) |
| arXiv:Hertz_Style_Aligned_Image_Generation_via_Shared_Attention_CVPR_2024_paper |
|
aim-uofa/StyleDrop-PyTorch/open_clip/model.py dcd422d66b0581d8 |
ran · our draft was wrong
|
MIT (permissive) |