| EmoLat: Text-driven Image Sentiment Transfer via Emotion Latent Space added by Syntology |
2026-01 (from id) |
JingVIPLab/EmoLat/model/transformer.py 4223bbb30a11ebd5 |
unverified |
no licence file found · pointer only |
| Rank-based Geographical Regularization: Revisiting Contrastive Self-Supervised Learning for Multispectral Remote Sensing Imagery added by Syntology |
2026-01 (from id) |
tomburgert/georank/models/scalemae_backbone.py c22ce1de0a795380 |
unverified |
no licence file found · pointer only |
| Human Motion Synthesis in 3D Scenes via Unified Scene Semantic Occupancy added by Syntology |
2025-11 (from id) |
jingyugong/SSOMotion/model/transformer.py 5aad82953ecc92f5 |
unverified |
no licence file found · pointer only |
| Disambiguating Reference in Visually Grounded Dialogues through Joint Modeling of Textual and Multimodal Semantic Structures |
16 May 2025 |
ashkamath/mdetr/models/transformer.py 0796c9e7b3ba5e23 |
unverified |
Apache-2.0 (permissive) |
| VLog: Video-Language Models by Generative Retrieval of Narration Vocabulary |
12 Mar 2025 |
showlab/VLog/VLog/model/models.py c8d145f15e26a93d |
unverified |
no licence file found · pointer only |
| The Devil is in the Spurious Correlation: Boosting Moment Retrieval via Temporal Dynamic Learning |
13 Jan 2025 |
xyangzhou/TD-DETR/td_detr/transformer.py 595b7178ce4c0616 |
unverified |
MIT (permissive) |
| Interpretable Enzyme Function Prediction via Residue-Level Detection |
10 Jan 2025 |
yangzhao1230/protdetr/models/transformer.py d8c3057c1f1df693 |
unverified |
MIT (permissive) |
| FlashVTG: Feature Layering and Adaptive Score Handling Network for Video Temporal Grounding |
18 Dec 2024 |
zhuo-cao/flashvtg/FlashVTG/transformer.py 2359afb511e0c62c |
unverified |
no licence file found · pointer only |
| Mamba-ST: State Space Model for Efficient Style Transfer |
16 Sep 2024 |
filippobotti/mambast/models/mamba.py 163a3014ed2dbf50 |
unverified |
no licence file found · pointer only |
| Visual Grounding with Multi-modal Conditional Adaptation |
8 Sep 2024 |
mr-bigworth/mmca/models_mmca_vector_based/visual_model/transformer.py dd06382c371fad2d |
unverified |
MIT (permissive) |
| Cross-Platform Video Person ReID: A New Benchmark Dataset and Adaptation Approach |
14 Aug 2024 |
FHR-L/VSLA-CLIP/model/make_model_clipvideoreid_reidadapter_pbp.py a8f5824659641940 |
unverified |
no licence file found · pointer only |
| Boosting Gaze Object Prediction via Pixel-level Supervision from Vision Foundation Model |
2 Aug 2024 |
jinyang06/SamGOP/maskGOP/transformer.py f52bf9a389e50bbc |
unverified |
Apache-2.0 (permissive) |
| NavGPT-2: Unleashing Navigational Reasoning Capability for Large Vision-Language Models |
17 Jul 2024 |
GengzeZhou/NavGPT-2/map_nav_src/models/transformer.py df2919fd1dad1d43 |
unverified |
MIT (permissive) |
| Nonverbal Interaction Detection |
11 Jul 2024 |
weijianan1/nvi/NVI-DEHR/models/transformer.py 5c4a9bab338ee34e |
unverified |
MIT (permissive) |
| SegVG: Transferring Object Bounding Box to Segmentation for Visual Grounding |
3 Jul 2024 |
WeitaiKang/SegVG/models/SegVG.py 9008729eeca3ae5e |
unverified |
no licence file found · pointer only |
| LW-DETR: A Transformer Replacement to YOLO for Real-Time Detection |
5 Jun 2024 |
atten4vis/lw-detr/models/transformer.py ec7a90cc24d6630d |
unverified |
Apache-2.0 (permissive) |
| Disentangled Pre-training for Human-Object Interaction Detection |
2 Apr 2024 |
xingaoli/DP-HOI/models/transformer.py fdc5e45115a1f362 |
unverified |
Apache-2.0 (permissive) |
| Beyond MOT: Semantic Multi-Object Tracking |
8 Mar 2024 |
Nathan-Li123/SMOTer/smoter/modeling/roi_heads/transformer.py 88dddc90d2044da0 |
unverified |
Apache-2.0 (permissive) |
| HOISDF: Constraining 3D Hand-Object Pose Estimation with Global Signed Distance Fields |
26 Feb 2024 |
amathislab/hoisdf/common/nets/transformer.py 47b8b020f555d270 |
unverified |
no licence file found · pointer only |
| Glance and Focus: Memory Prompting for Multi-Event Video Question Answering |
3 Jan 2024 |
ByZ0e/Glance-Focus/model/transformer_gf.py e66b5e8b24dfea32 |
unverified |
MIT (permissive) |
| Towards Learning a Generalist Model for Embodied Navigation |
4 Dec 2023 |
lavi-lab/navillm/models/detr_transformer.py df2919fd1dad1d43 |
unverified |
MIT (permissive) |
| Mask-Attention-Free Transformer for 3D Instance Segmentation |
4 Sep 2023 |
dvlab-research/mask-attention-free-transformer/maft/model/transformer.py 31be13e9daaa9da9 |
unverified |
no licence file found · pointer only |
| PromptMRG: Diagnosis-Driven Prompts for Medical Report Generation |
24 Aug 2023 |
jhb86253817/promptmrg/models/transformer.py 6c667e1b7257c930 |
unverified |
MIT (permissive) |
| RLIPv2: Fast Scaling of Relational Language-Image Pre-training |
18 Aug 2023 |
jacobyuan7/ocn-hoi-benchmark/models/transformer.py 9a524c433c47ed10 |
unverified |
Apache-2.0 (permissive) |
| Lip2Vec: Efficient and Robust Visual Speech Recognition via Latent-to-Latent Visual to Audio Representation Mapping |
11 Aug 2023 |
YasserdahouML/Lip2Vec/models/transformer.py 5c86d34cae9ef6a1 |
unverified |
no licence file found · pointer only |
| Bird's-Eye-View Scene Graph for Vision-Language Navigation |
9 Aug 2023 |
defaultrui/bev-scene-graph/bsg_vln/map_nav_src/models/transformer.py df2919fd1dad1d43 |
unverified |
no licence file found · pointer only |
| UniVTG: Towards Unified Video-Language Temporal Grounding |
31 Jul 2023 |
showlab/univtg/model/transformer.py a03d2e1a023a5c63 |
unverified |
MIT (permissive) |
| UniVTG: Towards Unified Video-Language Temporal Grounding |
31 Jul 2023 |
showlab/univtg/model/transformer_encoder_droppath.py 141ae5fb5b4041f0 |
unverified |
MIT (permissive) |
| Learning Dynamic Query Combinations for Transformer-based Object Detection and Segmentation |
23 Jul 2023 |
bytedance/DQ-Det/Cond-DETR-DQ/models/transformer.py 86f65ed94e4398da |
unverified |
Apache-2.0 (permissive) |
| Compositional Text-to-Image Synthesis with Attention Map Control of Diffusion Models |
23 May 2023 |
OPPO-Mente-Lab/attention-mask-control/boxnet_models/transformer.py 9083e7dd88fe1570 |
unverified |
MIT (permissive) |
| EVREAL: Towards a Comprehensive Benchmark and Analysis Suite for Event-based Video Reconstruction |
30 Apr 2023 |
ercanburak/EVREAL/model/eitr/transformer.py 858bdc32c51ebdef |
unverified |
MIT (permissive) |
| Disentangling Writer and Character Styles for Handwriting Generation |
26 Mar 2023 |
dailenson/SDT/models/transformer.py d3c210804250a1dc |
unverified |
MIT (permissive) |
| Understanding Embodied Reference with Touch-Line Transformer |
11 Oct 2022 |
yang-li-2000/understanding-embodied-reference-with-touch-line-transformer/models/transformer.py 1e7c78d96968a055 |
unverified |
Apache-2.0 (permissive) |
| Understanding Embodied Reference with Touch-Line Transformer |
11 Oct 2022 |
yang-li-2000/understanding-embodied-reference-with-touch-line-transformer/models/transformer_ori.py b0beffbe8deaa71a |
unverified |
Apache-2.0 (permissive) |
| Fine-Grained Image Style Transfer with Visual Transformers |
11 Oct 2022 |
researchmm/sttr/models_istt/transformer_nonorm_flx.py a7bb87542c72ca8e |
unverified |
MIT (permissive) |
| ECO-TR: Efficient Correspondences Finding Via Coarse-to-Fine Refinement |
25 Sep 2022 |
dltan7/ECO-TR/src/models/ecotr_modules/transformer.py 8ee33b336f178e2f |
unverified |
Apache-2.0 (permissive) |
| RLIP: Relational Language-Image Pre-training for Human-Object Interaction Detection |
5 Sep 2022 |
JacobYuan7/RLIP/models/ParSetransformer.py 1db67394391c79fa |
unverified |
Apache-2.0 (permissive) |
| Cross-Attention of Disentangled Modalities for 3D Human Mesh Recovery with Transformers |
27 Jul 2022 |
postech-ami/fastmetro/src/modeling/model/modeling_fastmetro.py c6caa9a5fb3f06a7 |
unverified |
MIT (permissive) |
| SiRi: A Simple Selective Retraining Mechanism for Transformer-based Visual Grounding |
27 Jul 2022 |
qumengxue/siri-vg/models/transformer.py f8cea2a98b3268dd |
unverified |
Apache-2.0 (permissive) |
| Rethinking the Two-Stage Framework for Grounded Situation Recognition |
10 Dec 2021 |
kellyiss/situformer/models/transformer.py da89ce279ae764b6 |
unverified |
no licence file found · pointer only |
| UniTAB: Unifying Text and Box Outputs for Grounded Vision-Language Modeling |
23 Nov 2021 |
microsoft/UniTAB/models/transformer_unitab.py 0b02672adf92e20a |
unverified |
MIT (permissive) |
| Grounded Situation Recognition with Transformers |
19 Nov 2021 |
jhcho99/gsrtr/models/transformer.py 9888762109072b2a |
unverified |
Apache-2.0 recorded; this copy not marked cleared · pointer only |
| Transformer-based Dual Relation Graph for Multi-label Image Recognition |
10 Oct 2021 |
iCVTEAM/TDRG/models/TDRG.py 2a6ad2bf42c596fb |
unverified |
Apache-2.0 (permissive) |
| Boundary-aware Transformers for Skin Lesion Segmentation |
8 Oct 2021 |
jcwang123/BA-Transformer/src/transformer.py 1b1d552ee72944ff |
unverified |
no licence file found · pointer only |
| PubTables-1M: Towards comprehensive table extraction from unstructured documents |
30 Sep 2021 |
phamquiluan/table-transformer/detr/models/transformer.py 302ec916d91b12a8 |
unverified |
MIT recorded; this copy not marked cleared · pointer only |
| Pix2seq: A Language Modeling Framework for Object Detection |
22 Sep 2021 |
gaopengcuhk/Unofficial-Pix2Seq/models/transformer.py cd17dcaf812b820d |
unverified |
Apache-2.0 recorded; this copy not marked cleared · pointer only |
| GroupFormer: Group Activity Recognition with Clustered Spatial-Temporal Transformer |
28 Aug 2021 |
xueyee/groupformer/group/models/transformer.py 302ec916d91b12a8 |
unverified |
Apache-2.0 (permissive) |
| Improving 3D Object Detection with Channel-wise Transformer |
23 Aug 2021 |
hlsheng1/ct3d/pcdet/models/roi_heads/ct3d_head.py 2b5d92421f68b895 |
unverified |
MIT (permissive) |
| Conditional DETR for Fast Training Convergence |
13 Aug 2021 |
atten4vis/conditionaldetr/models/transformer.py bf149877606cbfba |
unverified |
Apache-2.0 (permissive) |
| DETReg: Unsupervised Pretraining with Region Priors for Object Detection |
8 Jun 2021 |
amirbar/detreg/models/transformer.py ccd0edc0cca7b92e |
unverified |
Apache-2.0 (permissive) |
| MDETR -- Modulated Detection for End-to-End Multi-Modal Understanding |
26 Apr 2021 |
b-faye/lightmdetr/models/transformer.py 09e538653a528ac8 |
unverified |
Apache-2.0 (permissive) |
| MDETR -- Modulated Detection for End-to-End Multi-Modal Understanding |
26 Apr 2021 |
b-faye/lightmdetr/models/transformer_plus.py 751cd5726698116b |
unverified |
Apache-2.0 (permissive) |
| M3DeTR: Multi-representation, Multi-scale, Mutual-relation 3D Object Detection with Transformers |
24 Apr 2021 |
rayguan97/M3DeTR/pcdet/models/backbones_2d/transformer.py 87f4c841962a967b |
unverified |
Apache-2.0 (permissive) |
| TransVG: End-to-End Visual Grounding with Transformers |
17 Apr 2021 |
nku-shengzheliu/Pytorch-TransVG/models/transformer.py 87992961bac26fc7 |
unverified |
MIT (permissive) |
| Video-aided Unsupervised Grammar Induction |
9 Apr 2021 |
Sy-Zhang/MMC-PCFG/lib/model/vpcfg/transformer.py 302ec916d91b12a8 |
unverified |
MIT (permissive) |
| TubeR: Tubelet Transformer for Video Action Detection |
2 Apr 2021 |
amazon-science/tubelet-transformer/models/transformer/transformer.py 7b94aa61b4d5ecac |
unverified |
Apache-2.0 (permissive) |
| End-to-End Trainable Multi-Instance Pose Estimation with Transformers |
22 Mar 2021 |
pranoyr/pose-estimation-with-transformers/models/transformer.py 302ec916d91b12a8 |
unverified |
Apache-2.0 recorded; this copy not marked cleared · pointer only |
| End-to-End Trainable Multi-Instance Pose Estimation with Transformers |
22 Mar 2021 |
amathislab/poet/models/transformer.py de542ddb0a808b05 |
unverified |
Apache-2.0 (permissive) |
| End-to-End Object Detection with Transformers |
26 May 2020 |
facebookresearch/detr/models/transformer.py 302ec916d91b12a8 |
unverified |
Apache-2.0 recorded; this copy not marked cleared · pointer only |
| Enhancing the Transformer with Explicit Relational Encoding for Math Problem Solving |
15 Oct 2019 |
ischlag/TP-Transformer/models/transformer.py ce9444c4e1d8e4a6 |
unverified |
MIT (permissive) |
| arXiv:aaai_28500 |
|
AsuradaYuci/TF-CLIP/model/make_model_clipreid.py b14287d0108cb653 |
unverified |
MIT (permissive) |
| arXiv:Zheng_Towards_Learning_a_Generalist_Model_for_Embodied_Navigation_CVPR_2024_paper |
|
LaVi-Lab/NaviLLM/models/detr_transformer.py df2919fd1dad1d43 |
unverified |
MIT (permissive) |
| arXiv:Ye_Hierarchical_Modular_Network_for_Video_Captioning_CVPR_2022_paper |
|
MarcusNerva/HMN/models/encoders/transformer.py 3adf1f93bdafc63e |
unverified |
MIT (permissive) |
| arXiv:Tang_Progressive_Attention_on_Multi-Level_Dense_Difference_Maps_for_Generic_Event_CVPR_2022_paper |
|
MCG-NJU/DDM/DDM-Net/modeling/transformer.py 4341942277ed876c |
unverified |
MIT (permissive) |
| arXiv:Dang_FASTer_Focal_token_Acquiring-and-Scaling_Transformer_for_Long-term_3D_Objection_Detection_CVPR_2025_paper |
|
MSunDYY/FASTer/pcdet/models/model_utils/faster_utils.py 5939165c82dcb087 |
unverified |
Apache-2.0 (permissive) |
| arXiv:Chen_Recurrent_Glimpse-Based_Decoder_for_Detection_With_Transformer_CVPR_2022_paper |
|
zhechen/Deformable-DETR-REGO/models/transformer.py e4cf3178d4a1e4c0 |
unverified |
Apache-2.0 recorded; this copy not marked cleared · pointer only |