| Deep Spatially-Regularized and Superpixel-Based Diffusion Learning for Unsupervised Hyperspectral Image Clustering added by Syntology |
2026-04 (from id) |
vburan01/DS2DL/pretrain_models.py 3d9050b0171c08f6 |
ran · fixture could not drive it
|
no licence file found · pointer only |
| AutoSSVH: Exploring Automated Frame Sampling for Efficient Self-Supervised Video Hashing |
4 Apr 2025 |
EliSpectre/CVPR25-AutoSSVH/model/AutoSSVH.py 5299444dd5349043 |
ran · honoured contract
fingerprinted |
no licence file found · pointer only |
| Uni-Sign: Toward Unified Sign Language Understanding at Scale |
25 Jan 2025 |
zechengli19/uni-sign/models.py 672ebe50dee3907a |
ran · honoured contract
fingerprinted |
no licence file found · pointer only |
| VideoICL: Confidence-based Iterative In-context Learning for Out-of-Distribution Video Understanding |
3 Dec 2024 |
kangsankim07/videoicl/InternVideo/InternVideo1/Downstream/Spatial-Temporal-Action-Localization/modeling_finetune.py 9b2250a3d13ed688 |
ran · honoured contract
|
no licence file found · pointer only |
| Enhancing Temporal Modeling of Video LLMs via Time Gating |
8 Oct 2024 |
lavi-lab/tg-vid/stllm/models/utils.py da651e3979a18f84 |
ran · honoured contract
fingerprinted |
no licence file found · pointer only |
| Scaling Proprioceptive-Visual Learning with Heterogeneous Pre-trained Transformers |
30 Sep 2024 |
liruiw/lerobot/lerobot/common/policies/hpt/modeling_hpt.py 9b7d923de2c72b94 |
ran · honoured contract
fingerprinted |
Apache-2.0 (permissive) |
| TinyVLA: Towards Fast, Data-Efficient Vision-Language-Action Models for Robotic Manipulation |
19 Sep 2024 |
liyaxuanliyaxuan/TinyVLA/policy_heads/models/detr_vae.py 35f9adf6c1ae4bbd |
unverified |
MIT (permissive) |
| VisionUnite: A Vision-Language Foundation Model for Ophthalmology Enhanced with Clinical Knowledge |
5 Aug 2024 |
HUANGLIZI/VisionUnite/ImageBind/models/multimodal_preprocessors.py 35a743e50e17e66d |
ran · honoured contract
fingerprinted |
MIT (permissive) |
| EgoVideo: Exploring Egocentric Foundation Model and Downstream Adaptation |
26 Jun 2024 |
opengvlab/egovideo/eccv-2022/modeling_finetune.py 9b2250a3d13ed688 |
ran · honoured contract
|
no licence file found · pointer only |
| OphNet: A Large-Scale Video Benchmark for Ophthalmic Surgical Workflow Understanding |
11 Jun 2024 |
minghu0830/ophnet-benchmark/baselines/task2/backbone/videomaev2/models/modeling_finetune.py da651e3979a18f84 |
ran · honoured contract
fingerprinted |
MIT (permissive) |
| Sparse-Tuning: Adapting Vision Transformers with Efficient Fine-tuning and Inference |
23 May 2024 |
liuting20/sparse-tuning/models/vit_video.py 9b2250a3d13ed688 |
ran · honoured contract
|
no licence file found · pointer only |
| MovieChat+: Question-aware Sparse Memory for Long Video Question Answering |
26 Apr 2024 |
rese1f/MovieChat/MovieChat/models/multimodal_preprocessors.py 35a743e50e17e66d |
ran · honoured contract
fingerprinted |
BSD-3-Clause (permissive) |
| TIM: A Time Interval Machine for Audio-Visual Action Recognition |
8 Apr 2024 |
JacobChalk/TIM/feature_extractors/VideoMAE/modeling_finetune.py da651e3979a18f84 |
ran · honoured contract
fingerprinted |
no licence file found · pointer only |
| Bridging the Gap Between End-to-End and Two-Step Text Spotting |
6 Apr 2024 |
mxin262/bridging-text-spotting/adet/modeling/bridge.py e270b1e0375e984c |
ran · honoured contract
fingerprinted |
licence not identified · pointer only |
| ST-LLM: Large Language Models Are Effective Temporal Learners |
30 Mar 2024 |
TencentARC/ST-LLM/stllm/models/utils.py da651e3979a18f84 |
ran · honoured contract
fingerprinted |
Apache-2.0 (permissive) |
| Benchmarking the Robustness of Temporal Action Detection Models Against Temporal Corruptions |
29 Mar 2024 |
Alvin-Zeng/temporal-robustness-benchmark/extract_corrupted_feature_code/videomae_v2/models/modeling_finetune.py da651e3979a18f84 |
ran · honoured contract
fingerprinted |
no licence file found · pointer only |
| Video Mamba Suite: State Space Model as a Versatile Alternative for Video Understanding |
14 Mar 2024 |
opengvlab/video-mamba-suite/video-mamba-suite/action-recognition/models/modeling_finetune.py f9c6fc8c1ddfbfac |
ran
|
MIT (permissive) |
| CAST: Cross-Attention in Space and Time for Video Action Recognition |
30 Nov 2023 |
khu-vll/cast/models/bidir_modeling_crossattn.py 9b2250a3d13ed688 |
ran · honoured contract
|
licence not identified · pointer only |
| Learning Human Action Recognition Representations Without Real Humans |
10 Nov 2023 |
howardzh01/ppma/code/omnivision/models/vision_transformer.py 35a743e50e17e66d |
ran · honoured contract
fingerprinted |
licence not identified · pointer only |
| Frozen Transformers in Language Models Are Effective Visual Encoder Layers |
19 Oct 2023 |
ziqipang/lm4visualencoding/video_understanding/modeling_finetune.py da651e3979a18f84 |
ran · honoured contract
fingerprinted |
MIT (permissive) |
| LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic Alignment |
3 Oct 2023 |
zhihaozhang97/ru-ai/imagebind/models/multimodal_preprocessors.py 35a743e50e17e66d |
ran · honoured contract
fingerprinted |
MIT (permissive) |
| Unsupervised Open-Vocabulary Object Localization in Videos |
18 Sep 2023 |
amazon-science/object-centric-vol/models/videomae.py da651e3979a18f84 |
ran · honoured contract
fingerprinted |
Apache-2.0 (permissive) |
| FactoFormer: Factorized Hyperspectral Transformers with Self-Supervised Pretraining |
18 Sep 2023 |
csiro-robotics/factoformer/pretraining/utils.py 28fcf01c3871bad4 |
ran
fingerprinted |
licence not identified · pointer only |
| Disentangling Spatial and Temporal Learning for Efficient Image-to-Video Transfer Learning |
14 Sep 2023 |
alibaba-mmai-research/dist/models/base/vit_video.py 790ff3bdfa86f045 |
ran
fingerprinted |
no licence file found · pointer only |
| Point-Bind & Point-LLM: Aligning Point Cloud with Multi-modality for 3D Understanding, Generation, and Instruction Following |
1 Sep 2023 |
ziyuguo99/point-bind_point-llm/Point-LLM/ImageBind/models/multimodal_preprocessors.py 35a743e50e17e66d |
ran · honoured contract
fingerprinted |
MIT (permissive) |
| Prune Spatio-temporal Tokens by Semantic-aware Temporal Accumulation |
8 Aug 2023 |
mark12ding/sta/model_vit.py 9b2250a3d13ed688 |
ran · honoured contract
|
BSD-2-Clause (permissive) |
| MAE-DFER: Efficient Masked Autoencoder for Self-supervised Dynamic Facial Expression Recognition |
5 Jul 2023 |
sunlicai/mae-dfer/modeling_finetune.py 9b2250a3d13ed688 |
ran · honoured contract
|
MIT (permissive) |
| Masked Diffusion Models Are Fast Distribution Learners |
20 Jun 2023 |
jiachenlei/maskdm/models/mask_uvit.py 7dff50b3dd446750 |
unverified |
MIT (permissive) |
| PandaGPT: One Model To Instruction-Follow Them All |
25 May 2023 |
yxuansu/pandagpt/code/model/ImageBind/models/multimodal_preprocessors.py 35a743e50e17e66d |
ran · honoured contract
fingerprinted |
Apache-2.0 (permissive) |
| ImageBind: One Embedding Space To Bind Them All |
9 May 2023 |
facebookresearch/imagebind/imagebind/models/imagebind_model.py d4f1ab1a5a3b3c4d |
ran · honoured contract
fingerprinted |
licence not identified · pointer only |
| Lightweight, Pre-trained Transformers for Remote Sensing Timeseries |
27 Apr 2023 |
nasaharvest/presto/single_file_presto.py 68ff931ff8b48cce |
unverified |
MIT (permissive) |
| Lightweight, Pre-trained Transformers for Remote Sensing Timeseries |
27 Apr 2023 |
nasaharvest/presto/presto/presto.py 5e50178bfd9083d2 |
unverified |
MIT (permissive) |
| Multi-modal learning for geospatial vegetation forecasting |
28 Mar 2023 |
earthnet2021/earthnet-models-pytorch/earthnet_models_pytorch/model/contextformer.py 3e8438e23b2df6a9 |
unverified |
MIT (permissive) |
| Unmasked Teacher: Towards Training-Efficient Video Foundation Models |
28 Mar 2023 |
opengvlab/unmasked_teacher/single_modality/action_detection/modeling_finetune.py da651e3979a18f84 |
ran · honoured contract
fingerprinted |
MIT (permissive) |
| Knowledge-enhanced Visual-Language Pre-training on Chest Radiology Images |
27 Feb 2023 |
xiaoman-zhang/kad/A3_CLIP/models/vit.py 9b2250a3d13ed688 |
ran · honoured contract
|
MIT (permissive) |
| Ego-Body Pose Estimation via Ego-Head Pose Estimation |
9 Dec 2022 |
lijiaman/egoego_release/egoego/model/transformer_module.py 38165c64626f844a |
unverified |
MIT (permissive) |
| Self-Supervised Vision Transformers for Malware Detection |
15 Aug 2022 |
sachith500/sherlock/modeling_finetune.py 9b2250a3d13ed688 |
ran · honoured contract
|
MIT (permissive) |
| Mugs: A Multi-Granular Self-Supervised Learning Framework |
27 Mar 2022 |
sail-sg/mugs/eval/eval_finetuning/model_for_finetuning.py a45ecd9b366b453e |
unverified |
Apache-2.0 (permissive) |
| VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-Training |
23 Mar 2022 |
MCG-NJU/VideoMAE-Action-Detection/modeling_finetune.py da651e3979a18f84 |
ran · honoured contract
fingerprinted |
licence not identified · pointer only |
| PortaSpeech: Portable and High-Quality Generative Text-to-Speech |
30 Sep 2021 |
keonlee9420/PortaSpeech/model/linguistic_encoder.py 0f8f22937463d3c5 |
ran · honoured contract
fingerprinted |
MIT (permissive) |
| Meta-StyleSpeech : Multi-Speaker Adaptive Text-to-Speech Generation |
6 Jun 2021 |
keonlee9420/StyleSpeech/model/modules.py 0f8f22937463d3c5 |
ran · honoured contract
fingerprinted |
MIT (permissive) |
| Parallel Tacotron 2: A Non-Autoregressive Neural TTS Model with Differentiable Duration Modeling |
2021-03 (from id) |
keonlee9420/Cross-Speaker-Emotion-Transfer/model/modules.py 0f8f22937463d3c5 |
ran · honoured contract
fingerprinted |
MIT (permissive) |
| FastPitch: Parallel Text-to-speech with Pitch Prediction |
11 Jun 2020 |
keonlee9420/FastPitchFormant/model/modules.py 0f8f22937463d3c5 |
ran · honoured contract
fingerprinted |
MIT (permissive) |
| FastSpeech 2: Fast and High-Quality End-to-End Text to Speech |
8 Jun 2020 |
mtresearcher/FastSpeech2/transformer/Models.py 0f8f22937463d3c5 |
ran · honoured contract
fingerprinted |
MIT (permissive) |
| Blank Language Models |
8 Feb 2020 |
Varal7/blank_language_model/transformer/Models.py 0f8f22937463d3c5 |
ran · honoured contract
fingerprinted |
Apache-2.0 (permissive) |
| Satellite Image Time Series Classification with Pixel-Set Encoders and Temporal Self-Attention |
18 Nov 2019 |
maja601/pytorch-psetae/models/tae.py 49a272e908e62940 |
unverified |
MIT (permissive) |
| Learn to Explain Efficiently via Neural Logic Inductive Learning |
6 Oct 2019 |
gblackout/NLIL/model/Models.py 0f8f22937463d3c5 |
ran · honoured contract
fingerprinted |
MIT (permissive) |
| BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding |
11 Oct 2018 |
huanghonggit/Mask-Language-Model/model/bert.py dcca3917b85d56dc |
ran · honoured contract
fingerprinted |
Apache-2.0 (permissive) |
| Attention Is All You Need |
12 Jun 2017 |
Matthewdowney18/Transformer_Dialogue/src/transformer/Models.py 0f8f22937463d3c5 |
ran · honoured contract
fingerprinted |
no licence file found · pointer only |
| Attention Is All You Need |
12 Jun 2017 |
graykode/nlp-tutorial/5-1.Transformer/Transformer.py 138cf0fec8960156 |
ran · honoured contract
|
MIT (permissive) |
| arXiv:Zheng_CO2-Net_A_Physics-Informed_Spatio-Temporal_Model_for_Global_Surface_CO2_Reconstruction_ICCV_2025_paper |
|
Leamonz/CORE/models/spatial_expert.py 35a743e50e17e66d |
ran · honoured contract
fingerprinted |
MIT (permissive) |
| arXiv:2023.findings-emnlp.672 |
|
guihuzhang/FactSpotter/g2t_gen_code/code/modeling_kgpt.py 38165c64626f844a |
unverified |
MIT (permissive) |
| arXiv:2020.emnlp-main.697 |
|
wenhuchen/KGPT/code/Model.py 0f8f22937463d3c5 |
ran · honoured contract
fingerprinted |
MIT (permissive) |