| M3R: Localized Rainfall Nowcasting with Meteorology-Informed MultiModal Attention added by Syntology |
2026-04 (from id) |
Sanjeev97/M3Rain/models/m3.py a7371f8ce642b182 |
ran · metamorphic tier: deterministic
|
no licence file found · pointer only |
| Scaling Continual Learning to 300+ Tasks with Bi-Level Routing Mixture-of-Experts added by Syntology |
2026-02 (from id) |
LMMMEng/CaRE/backbone/vit_brmoe.py afa43a669cda3d4f |
unverified |
Apache-2.0 (permissive) |
| CovMatch: Cross-Covariance Guided Multimodal Dataset Distillation with Trainable Text Encoder added by Syntology |
2025-10 (from id) |
Yongalls/CovMatch/src/model.py dee6bfa940b2b6a4 |
ran · metamorphic tier: deterministic
|
no licence file found · pointer only |
| Logits DeConfusion with CLIP for Few-Shot Learning |
16 Apr 2025 |
LiShuo1001/LDC/clip_ldc/model.py 1c4c241386ee3fe2 |
ran
|
no licence file found · pointer only |
| EMPLACE: Self-Supervised Urban Scene Change Detection |
22 Mar 2025 |
Timalph/EMPLACE/Rectangular_VIT.py 4cf9be3ee681cf7a |
unverified |
no licence file found · pointer only |
| arXiv:2503.15141 |
2025-03 (from id) |
djukicn/ocebo/models/ocebo.py 03a1d18dda4dfaf6 |
unverified |
Apache-2.0 (permissive) |
| MOS: Modeling Object-Scene Associations in Generalized Category Discovery |
15 Mar 2025 |
jethropeng/mos/network.py ad4dd58fd97d3768 |
ran
|
MIT (permissive) |
| SVIP: Semantically Contextualized Visual Patches for Zero-Shot Learning |
13 Mar 2025 |
uqzhichen/SVIP/models/vit_model.py 4c53eb4910649344 |
unverified |
no licence file found · pointer only |
| LongProLIP: A Probabilistic Vision-Language Model with Long Context Text |
11 Mar 2025 |
naver-ai/prolip/src/prolip/model.py 84aa74aff57a00bc |
unverified |
licence not identified · pointer only |
| LVFace: Large Vision model for Face Recogniton |
23 Jan 2025 |
bytedance/LVFace/backbones/vit.py 6d673e028dd93c24 |
unverified |
MIT (permissive) |
| Autoregressive Video Generation without Vector Quantization |
18 Dec 2024 |
baaivision/nova/diffnext/models/transformers/transformer_nova.py 3a44328c8bf72fe3 |
unverified |
Apache-2.0 (permissive) |
| On the Surprising Effectiveness of Attention Transfer for Vision Transformers |
14 Nov 2024 |
alexlioralexli/attention-transfer/models_dual_vit.py 02bbece47468309a |
unverified |
licence not identified · pointer only |
| Classification Done Right for Vision-Language Pre-Training |
5 Nov 2024 |
x-cls/superclass/opencls/open_clip/cls_model.py 554c03cc48b8d760 |
unverified |
Apache-2.0 (permissive) |
| SAFE: Slow and Fast Parameter-Efficient Tuning for Continual Learning with Pre-Trained Models |
4 Nov 2024 |
MIFA-Lab/SAFE/petl/vision_transformer_adapter.py 94ecb2b330290cca |
unverified |
MIT (permissive) |
| PointAD: Comprehending 3D Anomalies from Points and Pixels for Zero-shot 3D Anomaly Detection |
1 Oct 2024 |
zqhang/pointad/AnomalyCLIP_lib/AnomalyCLIP.py f8c143f839756ff3 |
unverified |
no licence file found · pointer only |
| AdaCLIP: Adapting CLIP with Hybrid Learnable Prompts for Zero-Shot Anomaly Detection |
22 Jul 2024 |
caoyunkang/AdaCLIP/method/adaclip.py 51acd62f8b9ee4da |
unverified |
MIT (permissive) |
| Recurrent Early Exits for Federated Learning with Heterogeneous Clients |
23 May 2024 |
royson/reefl/src/models/reefl_vit.py dbbfaa54d7e3f160 |
unverified |
MIT (permissive) |
| LookHere: Vision Transformers with Directed Attention Generalize and Extrapolate |
22 May 2024 |
greencubic/lookhere/lookhere.py 9b6612fb34289faa |
ran · metamorphic tier: deterministic
|
MIT (permissive) |
| Exploring Learngene via Stage-wise Weight Sharing for Initializing Variable-sized Models |
25 Apr 2024 |
AlphaXia/SWS/SWS_vit.py be0e5172207a25d4 |
unverified |
Apache-2.0 (permissive) |
| LUM-ViT: Learnable Under-sampling Mask Vision Transformer for Bandwidth Limited Optical Signal Acquisition |
3 Mar 2024 |
maxllf/lum-vit/LUM-ViT.py d2434fa3227589c0 |
ran · metamorphic tier: deterministic
|
no licence file found · pointer only |
| Generalizable Whole Slide Image Classification with Fine-Grained Visual-Semantic Interaction |
29 Feb 2024 |
ls1rius/wsi_five/models/FiVE.py fe38f5b5e9b6ed8c |
ran
|
no licence file found · pointer only |
| Hierarchical Vision Transformers for Context-Aware Prostate Cancer Grading in Whole Slide Images |
19 Dec 2023 |
computationalpathologygroup/hvit/source/models.py 21750327371e6568 |
ran · metamorphic tier: deterministic
|
no licence file found · pointer only |
| Learning Human Action Recognition Representations Without Real Humans |
10 Nov 2023 |
howardzh01/ppma/code/omnivision/models/vision_transformer.py 3e2d901955f04262 |
unverified |
licence not identified · pointer only |
| AnomalyCLIP: Object-agnostic Prompt Learning for Zero-shot Anomaly Detection |
29 Oct 2023 |
zqhang/anomalyclip/AnomalyCLIP_lib/AnomalyCLIP.py 175f83b44a42759d |
unverified |
MIT (permissive) |
| Accurate and Fast Compressed Video Captioning |
22 Sep 2023 |
acherstyx/CoCap/cocap/modules/compressed_video/compressed_video_transformer.py ead6c7478f4685ab |
ran · metamorphic tier: deterministic
|
MIT (permissive) |
| Shatter and Gather: Learning Referring Image Segmentation with Text Supervision |
29 Aug 2023 |
kdwonn/SaG/model/cross_modal_attention.py 19f4e674b83797e4 |
unverified |
no licence file found · pointer only |
| Masked Autoencoders are Efficient Class Incremental Learners |
24 Aug 2023 |
scok30/MAE-CIL/continual/vit.py e1d62aeb525afd2a |
ran
|
no licence file found · pointer only |
| Semantics Meets Temporal Correspondence: Self-supervised Object-centric Learning in Videos |
19 Aug 2023 |
shvdiwnkozbw/SMTC/src/model/model_action.py 932f488e9f83940b |
unverified |
no licence file found · pointer only |
| Prune Spatio-temporal Tokens by Semantic-aware Temporal Accumulation |
8 Aug 2023 |
Mark12Ding/STA/model_vit.py ae539267a5c6d2d3 |
unverified |
BSD-2-Clause (permissive) |
| Segment Any Point Cloud Sequences by Distilling Vision Foundation Models |
15 Jun 2023 |
valeoai/SLidR/model/modules/dino/vision_transformer.py 684c765a98595395 |
ran
fingerprinted |
licence not identified · pointer only |
| Text Promptable Surgical Instrument Segmentation with Vision-Language Models |
15 Jun 2023 |
franciszzj/tp-sis/model/segmenter.py d1bb444eef3c0759 |
ran
fingerprinted |
no licence file found · pointer only |
| ChessGPT: Bridging Policy Learning and Language Modeling |
15 Jun 2023 |
waterhorse1/chessgpt/chessclip/src/open_clip/model.py 13b07e13a2877644 |
ran
|
Apache-2.0 (permissive) |
| Improving Visual Prompt Tuning for Self-supervised Vision Transformers |
8 Jun 2023 |
ryongithub/gatedprompttuning/src/models/vit_prompt/vit_mae.py 33fb7be835804856 |
ran · metamorphic tier: deterministic
|
MIT (permissive) |
| Micron-BERT: BERT-based Facial Micro-Expression Recognition |
6 Apr 2023 |
uark-cviu/Micron-BERT/models/vision_transformer.py 1338feaa3f18f57c |
unverified |
no licence file found · pointer only |
| Cross-Modal Implicit Relation Reasoning and Aligning for Text-to-Image Person Retrieval |
22 Mar 2023 |
anosorae/irra/model/build.py 8d1b78ed1b645333 |
ran · metamorphic tier: deterministic
|
MIT (permissive) |
| Spatial-Aware Token for Weakly Supervised Object Localization |
18 Mar 2023 |
wpy1999/SAT/Model/SAT.py 55ee94b26d520b2d |
ran
|
no licence file found · pointer only |
| Where We Are and What We're Looking At: Query Based Worldwide Image Geo-localization Using Hierarchies and Scenes |
7 Mar 2023 |
AHKerrigan/GeoGuessNet/networks.py a53662872764fd3e |
unverified |
no licence file found · pointer only |
| RangeViT: Towards Vision Transformers for 3D Semantic Segmentation in Autonomous Driving |
24 Jan 2023 |
valeoai/rangevit/models/rangevit.py 6b3bd6eb6760b684 |
unverified |
Apache-2.0 (permissive) |
| Weakly Supervised Object Localization via Transformer with Implicit Spatial Calibration |
21 Jul 2022 |
164140757/SCM/lib/models/deit.py c56f1df1064bfab1 |
unverified |
MIT (permissive) |
| Peripheral Vision Transformer |
14 Jun 2022 |
juhongm999/pervit/model/pervit.py 65961eec77ac4145 |
unverified |
Apache-2.0 (permissive) |
| Safe Self-Refinement for Transformer-based Domain Adaptation |
16 Apr 2022 |
tsun/SSRT/model/SSRT.py 2a0271883efece1e |
unverified |
MIT (permissive) |
| VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-Training |
23 Mar 2022 |
MCG-NJU/VideoMAE-Action-Detection/modeling_finetune.py 15190c9e8134c543 |
unverified |
licence not identified · pointer only |
| Visual Prompt Tuning |
23 Mar 2022 |
Yiming-M/CLIP-EBC/models/encoder/vit.py 232688cf9e59bf54 |
unverified |
MIT (permissive) |
| Unified Visual Transformer Compression |
15 Mar 2022 |
VITA-Group/UVC/UVC/models/modeling.py 4900a340f86d38d4 |
unverified |
MIT (permissive) |
| Multi-class Token Transformer for Weakly Supervised Semantic Segmentation |
6 Mar 2022 |
xulianuwa/mctformer/models.py 4a3e09668a5f4b9f |
ran · metamorphic tier: deterministic
|
no licence file found · pointer only |
| Masked Autoencoders Are Scalable Vision Learners |
11 Nov 2021 |
FlyEgle/MAE-pytorch/model/Transformers/VIT/mae.py 19ec788aecbdfc22 |
unverified |
no licence file found · pointer only |
| IconQA: A New Benchmark for Abstract Diagram Understanding and Visual Language Reasoning |
25 Oct 2021 |
lupantech/iconqa/models/patch_transformer.py 81aa9d1ddc833ef9 |
ran
fingerprinted |
no licence file found · pointer only |
| TransMatcher: Deep Image Matching Through Transformers for Generalizable Person Re-identification |
30 May 2021 |
JDAI-CV/fast-reid/fastreid/modeling/backbones/vision_transformer.py 9368e7f2101c9a30 |
unverified |
Apache-2.0 (permissive) |
| CrossViT: Cross-Attention Multi-Scale Vision Transformer for Image Classification |
27 Mar 2021 |
IBM/CrossViT/models/crossvit.py b05cff1db97d5001 |
unverified |
Apache-2.0 (permissive) |
| Transformer-Based Attention Networks for Continuous Pixel-Wise Prediction |
22 Mar 2021 |
ygjwd12345/TransDepth/pytorch/TransUNet/networks/vit_seg_modeling.py 079c4f5086884f75 |
unverified |
MIT (permissive) |
| Training data-efficient image transformers & distillation through attention |
23 Dec 2020 |
alibaba/EasyCV/easycv/models/backbones/vision_transformer.py a4ea24ce126312f8 |
unverified |
Apache-2.0 (permissive) |
| An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale |
22 Oct 2020 |
YousefGamal220/Vision-Transformers/vision_transformer.py b186094acc970ace |
ran · metamorphic tier: deterministic
|
no licence file found · pointer only |
| An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale |
22 Oct 2020 |
mdmhriday/vision-transformers/models/vit.py 42dd7dea8413b92f |
ran · metamorphic tier: deterministic
|
MIT (permissive) |
| An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale |
22 Oct 2020 |
nateraw/lightning-vision-transformer/vit.py 78aca8814a1f8d16 |
ran · metamorphic tier: deterministic
|
MIT (permissive) |
| An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale |
22 Oct 2020 |
wangguanan/light-reid/lightreid/models/backbones/transformers/vit_timm.py d1e7240963528e63 |
ran
|
no licence file found · pointer only |
| An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale |
22 Oct 2020 |
smu-ivpl/DeepfakeDetection/facebook_deit.py cf2cdfee402d4881 |
ran
|
no licence file found · pointer only |
| An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale |
22 Oct 2020 |
asyml/vision-transformer-pytorch/src/model.py 3b44acd70cadcf44 |
ran
|
Apache-2.0 (permissive) |
| An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale |
22 Oct 2020 |
Burf/VisionTransformer-Tensorflow2/vit/vit.py 18eba67fe6d3f33b |
unverified |
MIT (permissive) |
| An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale |
22 Oct 2020 |
pytorch/vision/torchvision/models/vision_transformer.py 546b54df7bf1a2a7 |
unverified |
BSD-3-Clause (permissive) |
| An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale |
22 Oct 2020 |
mahmoodlab/hipt/HIPT_4K/vision_transformer.py 958ea05c3c1ae4ba |
unverified |
licence not identified · pointer only |
| An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale |
22 Oct 2020 |
nachiket273/VisTrans/vistrans/models/vit.py 2d96b3b67dc31418 |
unverified |
MIT (permissive) |
| An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale |
22 Oct 2020 |
gupta-abhay/ViT/vit/vit.py d022f0d80d7dc24f |
unverified |
MIT (permissive) |
| An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale |
22 Oct 2020 |
facebookresearch/ClassyVision/classy_vision/models/vision_transformer.py 443bd35acbf7f0b2 |
unverified |
MIT (permissive) |
| An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale |
22 Oct 2020 |
Abdulrahman-Adel/Real-Life-Violence-Detection/src/models/model_01.py e37137b1becc1165 |
unverified |
no licence file found · pointer only |
| An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale |
22 Oct 2020 |
DominikBatic/EndoViT/pretraining/mae/models_vit.py d3f0ee5f79046d1b |
unverified |
no licence file found · pointer only |
| An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale |
22 Oct 2020 |
junyongyou/triq/src/vit_iqa/ViT_pytorch/models/modeling.py 110bd6632560d038 |
unverified |
MIT (permissive) |
| An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale |
22 Oct 2020 |
ludics/ViT-Retri/vit_retri/models/modeling.py f50ac01b2aea5cee |
unverified |
MIT (permissive) |
| An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale |
22 Oct 2020 |
google-research/vision_transformer/vit_jax/models_vit.py bd276b1048e360eb |
unverified |
Apache-2.0 (permissive) |
| An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale |
22 Oct 2020 |
jankrepl/mildlyoverfitted/github_adventures/vision_transformer/custom.py f1b7003e1837294f |
unverified |
MIT (permissive) |
| An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale |
22 Oct 2020 |
skchen1993/TrangFG/models/modeling.py 951da7c89ffaf2ef |
unverified |
MIT (permissive) |
| An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale |
22 Oct 2020 |
Ugenteraan/Vanilla-ViT/ViT/ViT.py 953f99d43181a4e6 |
unverified |
MIT (permissive) |
| An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale |
22 Oct 2020 |
AlifAshrafee/ViT-pytorch-for-Cooking-State-Recognition/models/modeling.py 7fc8f8faa831f7f2 |
unverified |
MIT (permissive) |
| An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale |
22 Oct 2020 |
ttt496/VisionTransformer/vit_jax/models.py 51d3b9873cfbdfba |
unverified |
Apache-2.0 (permissive) |
| An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale |
22 Oct 2020 |
bshantam97/Attention_Based_Networks/vision_transformer.py fca973fc1fa56318 |
unverified |
no licence file found · pointer only |
| An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale |
22 Oct 2020 |
s-chh/pytorch-scratch-vision-transformer-vit/model.py c5610b768e912553 |
unverified |
MIT (permissive) |
| arXiv:aaai_29533 |
|
AlphaXia/TLEG/TLEG_vit.py 012078f556e5d24a |
unverified |
no licence file found · pointer only |