Methods › General › Attention Mechanisms › Attention › Papers, page 219
Attention Is All You Need
Attention
Papers archive 2025-07-28
archive papers tagged: 31,583 · with a code link: 13,473 · where Syntology ran a sample: 3,998 (3,366 with a run with no instrument failure, 632 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (3,998 of 31,583 tagged: 3,366 with a run with no instrument failure, 632 where every run was a failure of Syntology's instrument)
Page 219 of 316: papers 21,801 to 21,900 of 31,583, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Most Important Person-guided Dual-branch Cross-Patch Attention for Group Affect Recognition 14 Dec 2022 · 0 repositories · arXiv:2212.07055
-
Efficient Self-supervised Learning with Contextualized Target Representations for Vision, Speech and Language 14 Dec 2022 · 5 repositories · arXiv:2212.07525Syntology community repositories only · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 8 harvested samples) · 2 pointer-only (licence)
-
Evaluating Byte and Wordpiece Level Models for Massively Multilingual Semantic Parsing 14 Dec 2022 · 0 repositories · arXiv:2212.07223
-
Explainability of Text Processing and Retrieval Methods: A Critical Survey 14 Dec 2022 · 0 repositories · arXiv:2212.07126
-
PAC-MAN: Multi-Relation Network in Social Community for Personalized Hashtag Recommendation 14 Dec 2022 · 1 repository
-
VTCC-NLP at NL4Opt competition subtask 1: An Ensemble Pre-trained language models for Named Entity Recognition 14 Dec 2022 · 0 repositories · arXiv:2212.07219
-
CREPE: Can Vision-Language Foundation Models Reason Compositionally? 13 Dec 2022 · 1 repository · arXiv:2212.07796Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
Generative artificial intelligence-enabled dynamic detection of nicotine-related circuits 13 Dec 2022 · 0 repositories · arXiv:2212.06330
-
GPViT: A High Resolution Non-Hierarchical Vision Transformer with Group Propagation 13 Dec 2022 · 2 repositories · arXiv:2212.06795
-
Paraphrase Identification with Deep Learning: A Review of Datasets and Methods 13 Dec 2022 · 0 repositories · arXiv:2212.06933
-
RT-1: Robotics Transformer for Real-World Control at Scale 13 Dec 2022 · 1 repository · arXiv:2212.06817Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples)
-
Smart Journey in Istanbul: A Mobile Application in Smart Cities for Traffic Estimation by Harnessing Time Series 13 Dec 2022 · 0 repositories · arXiv:2212.09448
-
What do Vision Transformers Learn? A Visual Exploration 13 Dec 2022 · 1 repository · arXiv:2212.06727
-
DexBERT: Effective, Task-Agnostic and Fine-grained Representation Learning of Android Bytecode 12 Dec 2022 · 1 repository · arXiv:2212.05976
-
Automated ICD Coding using Extreme Multi-label Long Text Transformer-based Models 12 Dec 2022 · 1 repository · arXiv:2212.05857
-
BeautyREC: Robust, Efficient, and Content-preserving Makeup Transfer 12 Dec 2022 · 0 repositories · arXiv:2212.05855
-
Classifying the Ideological Orientation of User-Submitted Texts in Social Media 12 Dec 2022 · 1 repository
-
CTT-Net: A Multi-view Cross-token Transformer for Cataract Postoperative Visual Acuity Prediction 12 Dec 2022 · 1 repository · arXiv:2212.05794
-
Deep learning approaches to building rooftop thermal bridge detection from aerial images 12 Dec 2022 · 1 repository
-
NMS Strikes Back 12 Dec 2022 · 1 repository · arXiv:2212.06137
-
P-Transformer: Towards Better Document-to-Document Neural Machine Translation 12 Dec 2022 · 1 repository · arXiv:2212.05830
-
ROIFormer: Semantic-Aware Region of Interest Transformer for Efficient Self-Supervised Monocular Depth Estimation 12 Dec 2022 · 0 repositories · arXiv:2212.05729
-
T5Score: Discriminative Fine-tuning of Generative Evaluation Metrics 12 Dec 2022 · 2 repositories · arXiv:2212.05726
-
Video Prediction by Efficient Transformers 12 Dec 2022 · 1 repository · arXiv:2212.06026Syntology official: harvested, nothing ran · 0 ran · 7 unverified (of 7 harvested samples)
-
Extending TrOCR for Text Localization-Free OCR of Full-Page Scanned Receipt Images 11 Dec 2022 · 0 repositories · arXiv:2212.05525
-
PromptCAL: Contrastive Affinity Learning via Auxiliary Prompts for Generalized Novel Category Discovery 11 Dec 2022 · 1 repository · arXiv:2212.05590Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 1 where Syntology's instrument failed) · 5 unverified (of 7 harvested samples) · 1 pointer-only (licence)
-
Elixir: Train a Large Language Model on a Small GPU Cluster 10 Dec 2022 · 2 repositories · arXiv:2212.05339
-
Joint Spatio-Temporal Modeling for the Semantic Change Detection in Remote Sensing Images 10 Dec 2022 · 3 repositories · arXiv:2212.05245Syntology official: no sample here; runs from other or unrecorded repositories · 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
Thinking Fast and Slow in Large Language Models 10 Dec 2022 · 0 repositories · arXiv:2212.05206
-
MAGVIT: Masked Generative Video Transformer 10 Dec 2022 · 1 repository · arXiv:2212.05199Syntology official (archive's flag): 10 ran · 10 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 0 violated, 9 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 10 harvested samples) · 1 pointer-only (licence)
-
Position Embedding Needs an Independent Layer Normalization 10 Dec 2022 · 1 repository · arXiv:2212.05262
-
Punctuation Restoration for Singaporean Spoken Languages: English, Malay, and Mandarin 10 Dec 2022 · 1 repository · arXiv:2212.05356
-
SMILE: Scaling Mixture-of-Experts with Efficient Bi-level Routing 10 Dec 2022 · 0 repositories · arXiv:2212.05191
-
Structured information extraction from complex scientific text with fine-tuned large language models 10 Dec 2022 · 0 repositories · arXiv:2212.05238
-
Dynamic Test-Time Augmentation via Differentiable Functions 9 Dec 2022 · 1 repository · arXiv:2212.04681
-
Incorporating Emotions into Health Mention Classification Task on Social Media 9 Dec 2022 · 1 repository · arXiv:2212.05039
-
Masked Lip-Sync Prediction by Audio-Visual Contextual Exploitation in Transformers 9 Dec 2022 · 0 repositories · arXiv:2212.04970
-
MIMO Is All You Need : A Strong Multi-In-Multi-Out Baseline for Video Prediction 9 Dec 2022 · 1 repository · arXiv:2212.04655
-
Cross-Domain Synthetic-to-Real In-the-Wild Depth and Normal Estimation for 3D Scene Understanding 9 Dec 2022 · 0 repositories · arXiv:2212.05040
-
RCDT: Relational Remote Sensing Change Detection with Transformer 9 Dec 2022 · 1 repository · arXiv:2212.04869
-
Sparse Upcycling: Training Mixture-of-Experts from Dense Checkpoints 9 Dec 2022 · 1 repository · arXiv:2212.05055Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
The Turing Deception 9 Dec 2022 · 0 repositories · arXiv:2212.06721
-
TRBLLmaker -- Transformer Reads Between Lyrics Lines maker 9 Dec 2022 · 0 repositories · arXiv:2212.04917
-
Explain to me like I am five -- Sentence Simplification Using Transformers 8 Dec 2022 · 1 repository · arXiv:2212.04595
-
Federated Learning for Inference at Anytime and Anywhere 8 Dec 2022 · 0 repositories · arXiv:2212.04084
-
Group Generalized Mean Pooling for Vision Transformer 8 Dec 2022 · 0 repositories · arXiv:2212.04114
-
Harnessing the Power of Multi-Task Pretraining for Ground-Truth Level Natural Language Explanations 8 Dec 2022 · 1 repository · arXiv:2212.04231Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 1 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 8 harvested samples) · 6 pointer-only (licence)
-
LLM-Planner: Few-Shot Grounded Planning for Embodied Agents with Large Language Models 8 Dec 2022 · 1 repository · arXiv:2212.04088
-
NP4G : Network Programming for Generalization 8 Dec 2022 · 1 repository · arXiv:2212.11118
-
NRTR: Neuron Reconstruction with Transformer from 3D Optical Microscopy Images 8 Dec 2022 · 0 repositories · arXiv:2212.04163
-
The Role of AI in Drug Discovery: Challenges, Opportunities, and Strategies 8 Dec 2022 · 0 repositories · arXiv:2212.08104
-
Memorization of Named Entities in Fine-tuned BERT Models 7 Dec 2022 · 1 repository · arXiv:2212.03749
-
DeepSpeed Data Efficiency: Improving Deep Learning Model Quality and Training Efficiency via Efficient Data Sampling and Routing 7 Dec 2022 · 1 repository · arXiv:2212.03597
-
Gaussian Radar Transformer for Semantic Segmentation in Noisy Radar Data 7 Dec 2022 · 0 repositories · arXiv:2212.03690
-
Hierarchical multimodal transformers for Multi-Page DocVQA 7 Dec 2022 · 1 repository · arXiv:2212.05935
-
Learning-To-Embed: Adopting Transformer based models for E-commerce Products Representation Learning 7 Dec 2022 · 0 repositories · arXiv:2212.03725
-
Multimodal Vision Transformers with Forced Attention for Behavior Analysis 7 Dec 2022 · 0 repositories · arXiv:2212.03968
-
SimVTP: Simple Video Text Pre-training with Masked Autoencoders 7 Dec 2022 · 0 repositories · arXiv:2212.03490
-
TweetDrought: A Deep-Learning Drought Impacts Recognizer based on Twitter Data 7 Dec 2022 · 0 repositories · arXiv:2212.04001
-
ViTPose++: Vision Transformer for Generic Body Pose Estimation 7 Dec 2022 · 2 repositories · arXiv:2212.04246Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
A K-variate Time Series Is Worth K Words: Evolution of the Vanilla Transformer Architecture for Long-term Multivariate Time Series Forecasting 6 Dec 2022 · 0 repositories · arXiv:2212.02789
-
AbHE: All Attention-based Homography Estimation 6 Dec 2022 · 0 repositories · arXiv:2212.03029
-
Adaptive Testing of Computer Vision Models 6 Dec 2022 · 1 repository · arXiv:2212.02774Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 6 harvested samples) · 6 pointer-only (licence)
-
Controlled Text Generation using T5 based Encoder-Decoder Soft Prompt Tuning and Analysis of the Utility of Generated Text in AI 6 Dec 2022 · 0 repositories · arXiv:2212.02924
-
Counterfactual reasoning: Do language models need world knowledge for causal understanding? 6 Dec 2022 · 1 repository · arXiv:2212.03278
-
CySecBERT: A Domain-Adapted Language Model for the Cybersecurity Domain 6 Dec 2022 · 0 repositories · arXiv:2212.02974
-
Document-Level Abstractive Summarization 6 Dec 2022 · 1 repository · arXiv:2212.03013
-
Vision Transformer Computation and Resilience for Dynamic Inference 6 Dec 2022 · 0 repositories · arXiv:2212.02687
-
Event-based Monocular Dense Depth Estimation with Recurrent Transformers 6 Dec 2022 · 0 repositories · arXiv:2212.02791
-
FacT: Factor-Tuning for Lightweight Adaptation on Vision Transformer 6 Dec 2022 · 1 repository · arXiv:2212.03145Syntology official: harvested, nothing ran · 0 ran · 1 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
IncepFormer: Efficient Inception Transformer with Pyramid Pooling for Semantic Segmentation 6 Dec 2022 · 1 repository · arXiv:2212.03035
-
LUNA: Language Understanding with Number Augmentations on Transformers via Number Plugins and Pre-training 6 Dec 2022 · 1 repository · arXiv:2212.02691Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 1 honoured, 0 violated, 1 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
Modern French Poetry Generation with RoBERTa and GPT-2 6 Dec 2022 · 0 repositories · arXiv:2212.02911
-
Open World DETR: Transformer based Open World Object Detection 6 Dec 2022 · 0 repositories · arXiv:2212.02969
-
Pretrained Diffusion Models for Unified Human Motion Synthesis 6 Dec 2022 · 0 repositories · arXiv:2212.02837
-
Semantic-aware Message Broadcasting for Efficient Unsupervised Domain Adaptation 6 Dec 2022 · 1 repository · arXiv:2212.02739
-
Semantic-Conditional Diffusion Networks for Image Captioning 6 Dec 2022 · 2 repositories · arXiv:2212.03099
-
Simple Baseline for Weather Forecasting Using Spatiotemporal Context Aggregation Network 6 Dec 2022 · 1 repository · arXiv:2212.02952Syntology official (archive's flag): 5 ran · 5 ran (of which 4 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 5 unverified (of 10 harvested samples) · 10 pointer-only (licence)
-
Style transfer and classification in hebrew news items 6 Dec 2022 · 0 repositories · arXiv:2212.03019
-
UniGeo: Unifying Geometry Logical Reasoning via Reformulating Mathematical Expression 6 Dec 2022 · 2 repositories · arXiv:2212.02746Syntology official (archive's flag): 2 ran · 2 ran (of which 2 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified; every one of the 2 samples that ran constructed an object rather than computing a result (of 4 harvested samples) · 4 pointer-only (licence)
-
Video Object of Interest Segmentation 6 Dec 2022 · 0 repositories · arXiv:2212.02871
-
3D-LatentMapper: View Agnostic Single-View Reconstruction of 3D Shapes 5 Dec 2022 · 0 repositories · arXiv:2212.02184
-
Audio-Driven Co-Speech Gesture Video Generation 5 Dec 2022 · 0 repositories · arXiv:2212.02350
-
Automatic Generation of Factual News Headlines in Finnish 5 Dec 2022 · 0 repositories · arXiv:2212.02170
-
FBLNet: FeedBack Loop Network for Driver Attention Prediction 5 Dec 2022 · 0 repositories · arXiv:2212.02096
-
Mask Matching Transformer for Few-Shot Segmentation 5 Dec 2022 · 1 repository · arXiv:2301.01208
-
Retrieval as Attention: End-to-end Learning of Retrieval and Reading within a Single Transformer 5 Dec 2022 · 1 repository · arXiv:2212.02027
-
Unifying Vision, Text, and Layout for Universal Document Processing 5 Dec 2022 · 5 repositories · arXiv:2212.02623Syntology official: no sample here; runs from other or unrecorded repositories · 15 ran (of which 0 constructed an object rather than computing a result; 14 with no instrument failure: 0 honoured, 2 violated, 12 with no contract checked; 1 where Syntology's instrument failed) · 2 unverified (of 17 harvested samples) · 4 pointer-only (licence)
-
Video Games as a Corpus: Sentiment Analysis using Fallout New Vegas Dialog 5 Dec 2022 · 0 repositories · arXiv:2212.02168
-
Joint Self-Supervised Image-Volume Representation Learning with Intra-Inter Contrastive Clustering 4 Dec 2022 · 0 repositories · arXiv:2212.01893
-
Languages You Know Influence Those You Learn: Impact of Language Characteristics on Multi-Lingual Text-to-Text Transfer 4 Dec 2022 · 0 repositories · arXiv:2212.01757
-
Exploring Stochastic Autoregressive Image Modeling for Visual Representation 3 Dec 2022 · 1 repository · arXiv:2212.01610
-
Exploring the Limits of Differentially Private Deep Learning with Group-wise Clipping 3 Dec 2022 · 0 repositories · arXiv:2212.01539
-
Global memory transformer for processing long documents 3 Dec 2022 · 0 repositories · arXiv:2212.01650
-
Recognition and Prediction of Surgical Gestures and Trajectories Using Transformer Models in Robot-Assisted Surgery 3 Dec 2022 · 0 repositories · arXiv:2212.01683
-
ColD Fusion: Collaborative Descent for Distributed Multitask Finetuning 2 Dec 2022 · 0 repositories · arXiv:2212.01378
-
Event knowledge in large language models: the gap between the impossible and the unlikely 2 Dec 2022 · 1 repository · arXiv:2212.01488
-
FECAM: Frequency Enhanced Channel Attention Mechanism for Time Series Forecasting 2 Dec 2022 · 1 repository · arXiv:2212.01209
-
Relation-Aware Language-Graph Transformer for Question Answering 2 Dec 2022 · 1 repository · arXiv:2212.00975
-
Multi-scale Transformer Network with Edge-aware Pre-training for Cross-Modality MR Image Synthesis 2 Dec 2022 · 2 repositories · arXiv:2212.01108