| UDT: Reconciling U-Nets and Diffusion Transformers with Data-Adaptive Token Reduction added by Syntology |
2026-08 (from id) |
JN-Yun/UDT/models/UDT.py 7b75bf9ed4e14ba7 |
ran · fixture could not drive it
fingerprinted |
no licence file found · pointer only |
| PACT: Self-Evolving Physical Safety Alignment for Diffusion Policies in Embodied Manipulation added by Syntology |
2026-06 (from id) |
thu-ml/RDT2/models/rdt/model.py 06a688b14731d363 |
ran · fixture could not drive it
fingerprinted |
Apache-2.0 (permissive) |
| Robust-U1: Can MLLMs Self-Recover Corrupted Visual Content for Robust Understanding? added by Syntology |
2026-06 (from id) |
jqtangust/Robust-U1/modeling/modeling/bagel/modeling_utils.py e5947aba1d10885f |
ran · fixture could not drive it
fingerprinted |
no licence file found · pointer only |
| On the training of physics-informed neural operators for solving parametric partial differential equations added by Syntology |
2026-06 (from id) |
NanxiiChen/PI-CViT/models/cvit.py 4814373f5d9ccbcc |
ran · fixture could not drive it
fingerprinted |
no licence file found · pointer only |
| Tadpole: Autoencoders as Foundation Models for 3D PDEs with Online Learning added by Syntology |
2026-05 (from id) |
tum-pbs/tadpole/tadpole/architecture/downstream/llm.py 7387c0fc16222bac |
ran
fingerprinted |
Apache-2.0 (permissive) |
| SynerMedGen: Synergizing Medical Multimodal Understanding with Generation via Task Alignment added by Syntology |
2026-05 (from id) |
piooip/SynerMedGen/modeling/bagel/modeling_utils.py e5947aba1d10885f |
ran · fixture could not drive it
fingerprinted |
Apache-2.0 (permissive) |
| Meta-Learning Hyperparameters for Parameter Efficient Fine-Tuning added by Syntology |
2026-03 (from id) |
doem97/metalora/models/satmae_vit.py 4bdd36eab04c3f62 |
ran · fixture could not drive it
fingerprinted |
Apache-2.0 (permissive) |
| EchoJEPA: A Latent Predictive Foundation Model for Echocardiography added by Syntology |
2026-02 (from id) |
bowang-lab/EchoJEPA/src/models/predictor.py f6dd6f9161d38af6 |
ran · fixture could not drive it
fingerprinted |
Apache-2.0 (permissive) |
| UniX: Unifying Autoregression and Diffusion for Chest X-Ray Understanding and Generation added by Syntology |
2026-01 (from id) |
ZrH42/UniX/modeling/unix/modeling_utils.py e5947aba1d10885f |
ran · fixture could not drive it
fingerprinted |
MIT (permissive) |
| AR-VRM: Imitating Human Motions for Visual Robot Manipulation with Analogical Reasoning added by Syntology |
2025-08 (from id) |
idejie/ar/models/ar/transformer_utils.py 77d61c6dcc6c0625 |
unverified |
no licence file found · pointer only |
| Uni-cot: Towards Unified Chain-of-Thought Reasoning Across Text and Vision added by Syntology |
2025-08 (from id) |
Fr0zenCrane/UniCoT/modeling/bagel/modeling_utils.py e5947aba1d10885f |
ran · fixture could not drive it
fingerprinted |
Apache-2.0 (permissive) |
| DreamVLA: A Vision-Language-Action Model Dreamed with Comprehensive World Knowledge |
6 Jul 2025 |
Zhangwenyao1/DreamVLA/models/dreamvla_model.py 5c8ea92d43ead77d |
ran · fixture could not drive it
fingerprinted |
no licence file found · pointer only |
| arXiv:2506.21803 |
2025-06 (from id) |
HKU-MedAI/MELP/src/melp/backbone/pos_embed.py 4bdd36eab04c3f62 |
ran · fixture could not drive it
fingerprinted |
MIT (permissive) |
| Hunyuan3D 2.5: Towards High-Fidelity 3D Assets Generation with Ultimate Details |
19 Jun 2025 |
tencent/hunyuan3d-2/hy3dgen/shapegen/models/conditioner.py 7700f3428f58afbe |
unverified |
licence not identified · pointer only |
| Test3R: Learning to Reconstruct 3D at Test Time |
16 Jun 2025 |
nopqaq/test3r/croco/models/pos_embed.py 4bdd36eab04c3f62 |
ran · fixture could not drive it
fingerprinted |
no licence file found · pointer only |
| Self-supervised Learning of Echocardiographic Video Representations via Online Cluster Distillation |
13 Jun 2025 |
mdivyanshu97/discovr/models/modeling_pretrain.py 44aacaf688b398e0 |
ran · fixture could not drive it
fingerprinted |
no licence file found · pointer only |
| Emerging Properties in Unified Multimodal Pretraining |
20 May 2025 |
neverbiasu/ComfyUI-BAGEL/modeling/bagel/modeling_utils.py e5947aba1d10885f |
ran · fixture could not drive it
fingerprinted |
Apache-2.0 (permissive) |
| CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework |
17 Apr 2025 |
event-ahu/cm3ae/pos_embed.py 4bdd36eab04c3f62 |
ran · fixture could not drive it
fingerprinted |
MIT (permissive) |
| Linguistics-aware Masked Image Modeling for Self-supervised Scene Text Recognition |
24 Mar 2025 |
zhangyifei01/LMIM/lmim_pretrain/models_lmim.py 12035a2f77d8016c |
unverified |
no licence file found · pointer only |
| Cosmos World Foundation Model Platform for Physical AI |
7 Jan 2025 |
nvidia-cosmos/cosmos-predict1/cosmos_predict1/autoregressive/modules/embedding.py e5947aba1d10885f |
ran · fixture could not drive it
fingerprinted |
Apache-2.0 (permissive) |
| Predictive Inverse Dynamics Models are Scalable Learners for Robotic Manipulation |
2024-12 (from id) |
openrobotlab/seer/models/seer_model.py 5c8ea92d43ead77d |
ran · fixture could not drive it
fingerprinted |
Apache-2.0 (permissive) |
| Causal Diffusion Transformers for Generative Modeling |
16 Dec 2024 |
identical code first harvested elsewhere e5947aba1d10885f |
ran · fixture could not drive it
fingerprinted |
licence of this copy not recorded |
| ZoomLDM: Latent Diffusion Model for multi-scale image generation |
25 Nov 2024 |
cvlab-stonybrook/ZoomLDM/cdm_dit/models.py e5947aba1d10885f |
ran · fixture could not drive it
fingerprinted |
MIT (permissive) |
| UrbanDiT: A Foundation Model for Open-World Urban Spatio-Temporal Learning |
19 Nov 2024 |
tsinghua-fib-lab/UrbanDiT/src/Embed.py 5c8ea92d43ead77d |
ran · fixture could not drive it
fingerprinted |
MIT (permissive) |
| IKEA Manuals at Work: 4D Grounding of Assembly Instructions on Internet Videos |
18 Nov 2024 |
yunongLiu1/IKEA-Manuals-at-Work/src/IKEAVideo/featurizers/MAE.py 12035a2f77d8016c |
unverified |
no licence file found · pointer only |
| ElasTST: Towards Robust Varied-Horizon Forecasting with Elastic Time-Series Transformer |
4 Nov 2024 |
microsoft/ProbTS/probts/utils/position_emb.py e5947aba1d10885f |
ran · fixture could not drive it
fingerprinted |
MIT (permissive) |
| Reconstructive Visual Instruction Tuning |
12 Oct 2024 |
haochen-wang409/ross/ross/model/utils.py 5c8ea92d43ead77d |
ran · fixture could not drive it
fingerprinted |
Apache-2.0 (permissive) |
| Generative Modeling of Molecular Dynamics Trajectories |
26 Sep 2024 |
bjing2016/mdgen/mdgen/model/latent_model.py e5947aba1d10885f |
ran · fixture could not drive it
fingerprinted |
MIT (permissive) |
| MonoFormer: One Transformer for Both Diffusion and Autoregression |
24 Sep 2024 |
MonoFormer/MonoFormer/models/modeling.py e5947aba1d10885f |
ran · fixture could not drive it
fingerprinted |
no licence file found · pointer only |
| Embedding Geometries of Contrastive Language-Image Pre-Training |
19 Sep 2024 |
eify/open_clip/src/open_clip/pos_embed.py 4bdd36eab04c3f62 |
ran · fixture could not drive it
fingerprinted |
no licence file found · pointer only |
| GR-MG: Leveraging Partially Annotated Data via Multi-Modal Goal-Conditioned Policy |
26 Aug 2024 |
bytedance/GR-MG/policy/model/vision_transformer.py 5c8ea92d43ead77d |
ran · fixture could not drive it
fingerprinted |
Apache-2.0 (permissive) |
| MegaFusion: Extend Diffusion Models towards Higher-resolution Image Generation without Further Tuning |
20 Aug 2024 |
haoningwu3639/MegaFusion/SD3-MegaFusion/model/embedding.py 63384d1943c486c1 |
ran
fingerprinted |
no licence file found · pointer only |
| Audio Prompt Adapter: Unleashing Music Editing Abilities for Text-to-Music with Lightweight Finetuning |
23 Jul 2024 |
fundwotsai2001/ap-adapter/audio_encoder/models_mae.py 4bdd36eab04c3f62 |
ran · fixture could not drive it
fingerprinted |
no licence file found · pointer only |
| ViTime: A Visual Intelligence-Based Foundation Model for Time Series Forecasting |
10 Jul 2024 |
ikeyang/vitime/model/ViTimeAutoencoder.py e5947aba1d10885f |
ran · fixture could not drive it
fingerprinted |
no licence file found · pointer only |
| AFBench: A Large-scale Benchmark for Airfoil Design |
2024-06 (from id) |
hitcslj/AFBench/models/dit.py 0cecc649cd728624 |
ran
fingerprinted |
MIT (permissive) |
| VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs |
11 Jun 2024 |
damo-nlp-sg/inf-clip/inf_clip/models/pos_embed.py 4bdd36eab04c3f62 |
ran · fixture could not drive it
fingerprinted |
Apache-2.0 (permissive) |
| A Closer Look at Time Steps is Worthy of Triple Speed-Up for Diffusion Model Training |
27 May 2024 |
1zeryu/speed/speed/networks/dit/mdt.py e5947aba1d10885f |
ran · fixture could not drive it
fingerprinted |
Apache-2.0 (permissive) |
| A Closer Look at Time Steps is Worthy of Triple Speed-Up for Diffusion Model Training |
27 May 2024 |
1zeryu/speed/speed/networks/pixart/PixArt.py 7700f3428f58afbe |
unverified |
Apache-2.0 (permissive) |
| CViT: Continuous Vision Transformer for Operator Learning |
22 May 2024 |
predictiveintelligencelab/cvit/src/model.py 4814373f5d9ccbcc |
ran · fixture could not drive it
fingerprinted |
MIT (permissive) |
| Inf-DiT: Upsampling Any-Resolution Image with Memory-Efficient Diffusion Transformer |
7 May 2024 |
thudm/inf-dit/dit/embeddings.py e5947aba1d10885f |
ran · fixture could not drive it
fingerprinted |
Apache-2.0 (permissive) |
| Continual Learning on a Diet: Learning from Sparsely Labeled Streams Under Constrained Computation |
19 Apr 2024 |
wx-zhang/continual-learning-on-a-diet/model/pose_embed.py 12035a2f77d8016c |
unverified |
no licence file found · pointer only |
| Diffusion-RWKV: Scaling RWKV-Like Architectures for Diffusion Models |
6 Apr 2024 |
feizc/diffusion-rwkv/models_drwkv.py e5947aba1d10885f |
ran · fixture could not drive it
fingerprinted |
no licence file found · pointer only |
| AdaGlimpse: Active Visual Exploration with Arbitrary Glimpse Position and Scale |
4 Apr 2024 |
apardyl/adaglimpse/architectures/mae_utils.py 5c8ea92d43ead77d |
ran · fixture could not drive it
fingerprinted |
no licence file found · pointer only |
| Long-CLIP: Unlocking the Long-Text Capability of CLIP |
22 Mar 2024 |
beichenzbc/long-clip/open_clip_long/pos_embed.py 4bdd36eab04c3f62 |
ran · fixture could not drive it
fingerprinted |
Apache-2.0 (permissive) |
| EquiAV: Leveraging Equivariance for Audio-Visual Contrastive Learning |
14 Mar 2024 |
JongSuk1/EquiAV/models/pos_embed.py 12035a2f77d8016c |
unverified |
MIT (permissive) |
| Chronos: Learning the Language of Time Series |
12 Mar 2024 |
mobile-sensing-and-ubicomp-laboratory/normwear/modules/layers.py 5c8ea92d43ead77d |
ran · fixture could not drive it
fingerprinted |
Apache-2.0 (permissive) |
| PixArt-Σ: Weak-to-Strong Training of Diffusion Transformer for 4K Text-to-Image Generation |
7 Mar 2024 |
identical code first harvested elsewhere e5947aba1d10885f |
ran · fixture could not drive it
fingerprinted |
licence of this copy not recorded |
| Multi-HMR: Multi-Person Whole-Body Human Mesh Recovery in a Single Shot |
22 Feb 2024 |
naver/multi-hmr/multi_hmr_anny/pos_embed.py 4bdd36eab04c3f62 |
ran · fixture could not drive it
fingerprinted |
no licence file found · pointer only |
| Key Patch Proposer: Key Patches Contain Rich Information |
18 Feb 2024 |
CA-TT-AC/key-patch-proposer/util/pos_embed.py 4bdd36eab04c3f62 |
ran · fixture could not drive it
fingerprinted |
no licence file found · pointer only |
| EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters |
6 Feb 2024 |
baaivision/EVA/EVA-01/eva/modeling_mae_pretrain.py 12035a2f77d8016c |
unverified |
MIT (permissive) |
| Spy-Watermark: Robust Invisible Watermarking for Backdoor Attack |
4 Jan 2024 |
rfww/spy-watermark/models/pos_embed.py 12035a2f77d8016c |
unverified |
no licence file found · pointer only |
| LEAP-VO: Long-term Effective Any Point Tracking for Visual Odometry |
3 Jan 2024 |
wrchen530/leapvo/main/leap/core/embeddings.py e5947aba1d10885f |
ran · fixture could not drive it
fingerprinted |
no licence file found · pointer only |
| One-Step Diffusion Distillation via Deep Equilibrium Models |
12 Dec 2023 |
locuslab/get/models/get.py e5947aba1d10885f |
ran · fixture could not drive it
fingerprinted |
MIT (permissive) |
| DiffiT: Diffusion Vision Transformers for Image Generation |
4 Dec 2023 |
nvlabs/diffit/diffit/diffit.py aa58f30251e59c74 |
ran
fingerprinted |
licence not identified · pointer only |
| Continual Self-supervised Learning: Towards Universal Multi-modal Medical Data Representation Learning |
29 Nov 2023 |
yeerwen/medcoss/model/Unimodel.py 12035a2f77d8016c |
unverified |
no licence file found · pointer only |
| CROMA: Remote Sensing Representations with Contrastive Radar-Optical Masked Autoencoders |
1 Nov 2023 |
antofuller/croma/pretrain_croma.py 2fae31fd41fc4c69 |
unverified |
MIT (permissive) |
| PixArt-$α$: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis |
30 Sep 2023 |
identical code first harvested elsewhere e5947aba1d10885f |
ran · fixture could not drive it
fingerprinted |
licence of this copy not recorded |
| MindGPT: Interpreting What You See with Non-invasive Brain Recordings |
27 Sep 2023 |
jxuanc/mindgpt/modules/pos_embed.py 12035a2f77d8016c |
unverified |
no licence file found · pointer only |
| Learning Tri-modal Embeddings for Zero-Shot Soundscape Mapping |
19 Sep 2023 |
mvrl/geoclap/geoclap/models/SATMAE.py 4bdd36eab04c3f62 |
ran · fixture could not drive it
fingerprinted |
no licence file found · pointer only |
| Stochastic positional embeddings improve masked image modeling |
31 Jul 2023 |
amirbar/stop/src/deit.py 4bdd36eab04c3f62 |
ran · fixture could not drive it
fingerprinted |
no licence file found · pointer only |
| MOCA: Self-supervised Representation Learning by Predicting Masked Online Codebook Assignments |
18 Jul 2023 |
valeoai/moca/util/model_utils.py 5c8ea92d43ead77d |
ran · fixture could not drive it
fingerprinted |
no licence file found · pointer only |
| Masked Trajectory Models for Prediction, Representation, and Control |
4 May 2023 |
facebookresearch/mtm/research/mtm/models/mtm_model.py efa8e5ed19b69df0 |
ran · honoured contract
fingerprinted |
MIT (permissive) |
| Diffusion Models as Masked Autoencoders |
6 Apr 2023 |
kimdanni/DiffMAE/models_cross.py 6a53acdcb72e78dd |
ran · fixture could not drive it
fingerprinted |
licence not identified · pointer only |
| Mask and Restore: Blind Backdoor Defense at Test Time with Masked Autoencoder |
27 Mar 2023 |
tsun/bdmae/mae/pos_embed.py 12035a2f77d8016c |
unverified |
MIT (permissive) |
| Object-Centric Slot Diffusion |
20 Mar 2023 |
jindongjiang/latent-slot-diffusion/src/models/utils.py 4bdd36eab04c3f62 |
ran · fixture could not drive it
fingerprinted |
MIT (permissive) |
| Advancing Radiograph Representation Learning with Masked Record Modeling |
30 Jan 2023 |
rl4m/mrm-pytorch/model_mrm.py ddbefd1ad8482d18 |
unverified |
MIT (permissive) |
| Masked Event Modeling: Self-Supervised Pretraining for Event Cameras |
20 Dec 2022 |
tum-vision/mem/mem/modeling_mae.py 12035a2f77d8016c |
unverified |
Apache-2.0 (permissive) |
| Scalable Diffusion Models with Transformers |
19 Dec 2022 |
hustvl/dig/models_dig.py e5947aba1d10885f |
ran · fixture could not drive it
fingerprinted |
MIT (permissive) |
| CroCo: Self-Supervised Pre-training for 3D Vision Tasks by Cross-View Completion |
19 Oct 2022 |
naver/croco/models/croco.py 4bdd36eab04c3f62 |
ran · fixture could not drive it
fingerprinted |
licence not identified · pointer only |
| Hephaestus: A large scale multitask dataset towards InSAR understanding |
20 Apr 2022 |
orion-ai-lab/hephaestus/self_supervised/mae/mae_model.py da34f99c07cf5566 |
unverified |
MIT (permissive) |
| Masked Autoencoders Are Scalable Vision Learners |
11 Nov 2021 |
0jason000/mae_vit/src/mae_vit.py 12035a2f77d8016c |
unverified |
Apache-2.0 (permissive) |
| Masked Autoencoders Are Scalable Vision Learners |
11 Nov 2021 |
nasa-impact/hls-foundation-os/geospatial_fm/geospatial_fm.py f6ed4fbd11e09648 |
unverified |
Apache-2.0 (permissive) |
| Nested Hierarchical Transformer: Towards Accurate, Data-Efficient and Interpretable Visual Understanding |
26 May 2021 |
freder-chen/vitp/util/pos_embed.py 12035a2f77d8016c |
unverified |
Apache-2.0 (permissive) |
| An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale |
22 Oct 2020 |
identical code first harvested elsewhere 5c8ea92d43ead77d |
ran · fixture could not drive it
fingerprinted |
licence of this copy not recorded |
| Unified Perceptual Parsing for Scene Understanding |
26 Jul 2018 |
ESA-PhiLab/PhilEO-MajorTOM/model/phileo_vit.py 4bdd36eab04c3f62 |
ran · fixture could not drive it
fingerprinted |
Apache-2.0 (permissive) |
| arXiv:Zhou_Towards_Effective_Foundation_Model_Adaptation_for_Extreme_Cross-Domain_Few-Shot_Learning_ICCV_2025_paper |
|
NWPUZhoufei/FMA/MAE_decoder.py 5c8ea92d43ead77d |
ran · fixture could not drive it
fingerprinted |
Apache-2.0 (permissive) |
| arXiv:Fang_Unleashing_Vanilla_Vision_Transformer_with_Masked_Image_Modeling_for_Object_ICCV_2023_paper |
|
hustvl/MIMDet/utils/pos_embed.py 12035a2f77d8016c |
unverified |
MIT (permissive) |