Methods › General › Attention Mechanisms › Attention › Papers, page 222
Attention Is All You Need
Attention
Papers archive 2025-07-28
archive papers tagged: 31,583 · with a code link: 13,473 · where Syntology ran a sample: 3,998 (3,366 with a run with no instrument failure, 632 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (3,998 of 31,583 tagged: 3,366 with a run with no instrument failure, 632 where every run was a failure of Syntology's instrument)
Page 222 of 316: papers 22,101 to 22,200 of 31,583, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Breakpoint Transformers for Modeling and Tracking Intermediate Beliefs 15 Nov 2022 · 1 repository · arXiv:2211.07950Syntology official: no sample here; runs from other or unrecorded repositories · 3 ran (of which 2 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 1 where Syntology's instrument failed) · 3 unverified (of 6 harvested samples) · 1 pointer-only (licence)
-
Contextual Transformer for Offline Meta Reinforcement Learning 15 Nov 2022 · 0 repositories · arXiv:2211.08016
-
ConvFormer: Combining CNN and Transformer for Medical Image Segmentation 15 Nov 2022 · 0 repositories · arXiv:2211.08564
-
Cross-Modality Transformer for Visible-Infrared Person Re-Identification 15 Nov 2022 · 0 repositories
-
Dynamic Temporal Filtering in Video Models 15 Nov 2022 · 1 repository · arXiv:2211.08252
-
Empowering Language Models with Knowledge Graph Reasoning for Question Answering 15 Nov 2022 · 0 repositories · arXiv:2211.08380
-
FedTune: A Deep Dive into Efficient Federated Fine-Tuning with Pre-trained Transformers 15 Nov 2022 · 0 repositories · arXiv:2211.08025
-
GLUE-X: Evaluating Natural Language Understanding Models from an Out-of-distribution Generalization Perspective 15 Nov 2022 · 1 repository · arXiv:2211.08073
-
Hybrid Transformers for Music Source Separation 15 Nov 2022 · 2 repositories · arXiv:2211.08553Syntology community repositories only · 9 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 0 violated, 9 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 10 harvested samples) · 10 pointer-only (licence)
-
Knowledge Distillation for Detection Transformer with Consistent Distillation Points Sampling 15 Nov 2022 · 2 repositories · arXiv:2211.08071
-
Latent Bottlenecked Attentive Neural Processes 15 Nov 2022 · 1 repository · arXiv:2211.08458
-
PromptCap: Prompt-Guided Task-Aware Image Captioning 15 Nov 2022 · 1 repository · arXiv:2211.09699
-
RobBERT-2022: Updating a Dutch Language Model to Account for Evolving Language Use 15 Nov 2022 · 0 repositories · arXiv:2211.08192
-
DeS3: Adaptive Attention-driven Self and Soft Shadow Removal using ViT Similarity 15 Nov 2022 · 1 repository · arXiv:2211.08089
-
Are Hard Examples also Harder to Explain? A Study with Human and Model-Generated Explanations 14 Nov 2022 · 1 repository · arXiv:2211.07517Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
Fcaformer: Forward Cross Attention in Hybrid Vision Transformer 14 Nov 2022 · 2 repositories · arXiv:2211.07198
-
CST5: Data Augmentation for Code-Switched Semantic Parsing 14 Nov 2022 · 1 repository · arXiv:2211.07514
-
Evade the Trap of Mediocrity: Promoting Diversity and Novelty in Text Generation via Concentrating Attention 14 Nov 2022 · 1 repository · arXiv:2211.07164
-
On Analyzing the Role of Image for Visual-enhanced Relation Extraction 14 Nov 2022 · 2 repositories · arXiv:2211.07504
-
QueryForm: A Simple Zero-shot Form Entity Query Framework 14 Nov 2022 · 0 repositories · arXiv:2211.07730
-
Technological taxonomies for hypernym and hyponym retrieval in patent texts 14 Nov 2022 · 1 repository · arXiv:2212.06039
-
Towards Robust Numerical Question Answering: Diagnosing Numerical Capabilities of NLP Systems 14 Nov 2022 · 0 repositories · arXiv:2211.07455
-
UGIF: UI Grounded Instruction Following 14 Nov 2022 · 0 repositories · arXiv:2211.07615
-
WSC-Trans: A 3D network model for automatic multi-structural segmentation of temporal bone CT 14 Nov 2022 · 0 repositories · arXiv:2211.07143
-
Demystify Self-Attention in Vision Transformers from a Semantic Perspective: Analysis and Application 13 Nov 2022 · 0 repositories · arXiv:2211.08543
-
Enhancing Few-shot Image Classification with Cosine Transformer 13 Nov 2022 · 1 repository · arXiv:2211.06828
-
GreenPLM: Cross-Lingual Transfer of Monolingual Pre-Trained Language Models at Almost No Cost 13 Nov 2022 · 1 repository · arXiv:2211.06993
-
Learning from partially labeled data for multi-organ and tumor segmentation 13 Nov 2022 · 1 repository · arXiv:2211.06894
-
Residual Degradation Learning Unfolding Framework with Mixing Priors across Spectral and Spatial for Compressive Spectral Imaging 13 Nov 2022 · 1 repository · arXiv:2211.06891
-
SSL4EO-S12: A Large-Scale Multi-Modal, Multi-Temporal Dataset for Self-Supervised Learning in Earth Observation 13 Nov 2022 · 4 repositories · arXiv:2211.07044
-
Textual Data Augmentation for Patient Outcomes Prediction 13 Nov 2022 · 0 repositories · arXiv:2211.06778
-
Large Language Models Meet Harry Potter: A Bilingual Dataset for Aligning Dialogue Agents with Characters 13 Nov 2022 · 1 repository · arXiv:2211.06869
-
Xu at SemEval-2022 Task 4: Pre-BERT Neural Network Methods vs Post-BERT RoBERTa Approach for Patronizing and Condescending Language Detection 13 Nov 2022 · 1 repository · arXiv:2211.06874
-
AU-Aware Vision Transformers for Biased Facial Expression Recognition 12 Nov 2022 · 0 repositories · arXiv:2211.06609
-
Dark patterns in e-commerce: a dataset and its baseline evaluations 12 Nov 2022 · 1 repository · arXiv:2211.06543
-
DEYO: DETR with YOLO for Step-by-Step Object Detection 12 Nov 2022 · 0 repositories · arXiv:2211.06588
-
End-to-End Machine Learning Framework for Facial AU Detection in Intensive Care Units 12 Nov 2022 · 0 repositories · arXiv:2211.06570
-
Kinematics Transformer: Solving The Inverse Modeling Problem of Soft Robots using Transformers 12 Nov 2022 · 0 repositories · arXiv:2211.06643
-
MultiCrossViT: Multimodal Vision Transformer for Schizophrenia Prediction using Structural MRI and Functional Network Connectivity Data 12 Nov 2022 · 0 repositories · arXiv:2211.06726
-
An Adapter based Multi-label Pre-training for Speech Separation and Enhancement 11 Nov 2022 · 0 repositories · arXiv:2211.06041
-
Control Transformer: Robot Navigation in Unknown Environments through PRM-Guided Return-Conditioned Sequence Modeling 11 Nov 2022 · 0 repositories · arXiv:2211.06407
-
DocuT5: Seq2seq SQL Generation with Table Documentation 11 Nov 2022 · 0 repositories · arXiv:2211.06193
-
Efficient HLA imputation from sequential SNPs data by Transformer 11 Nov 2022 · 1 repository · arXiv:2211.06430
-
Using Persuasive Writing Strategies to Explain and Detect Health Misinformation 11 Nov 2022 · 1 repository · arXiv:2211.05985
-
PatchBlender: A Motion Prior for Video Transformers 11 Nov 2022 · 0 repositories · arXiv:2211.14449
-
SSGVS: Semantic Scene Graph-to-Video Synthesis 11 Nov 2022 · 0 repositories · arXiv:2211.06119
-
The Architectural Bottleneck Principle 11 Nov 2022 · 0 repositories · arXiv:2211.06420
-
Token Transformer: Can class token help window-based transformer build better long-range interactions? 11 Nov 2022 · 0 repositories · arXiv:2211.06083
-
Assistive Completion of Agrammatic Aphasic Sentences: A Transfer Learning Approach using Neurolinguistics-based Synthetic Dataset 10 Nov 2022 · 0 repositories · arXiv:2211.05557
-
BERT-Based Combination of Convolutional and Recurrent Neural Network for Indonesian Sentiment Analysis 10 Nov 2022 · 0 repositories · arXiv:2211.05273
-
BERT in Plutarch's Shadows 10 Nov 2022 · 0 repositories · arXiv:2211.05673
-
Biomedical Multi-hop Question Answering Using Knowledge Graph Embeddings and Language Models 10 Nov 2022 · 0 repositories · arXiv:2211.05351
-
PAD-Net: An Efficient Framework for Dynamic Networks 10 Nov 2022 · 1 repository · arXiv:2211.05528
-
Demystify Transformers & Convolutions in Modern Image Deep Networks 10 Nov 2022 · 1 repository · arXiv:2211.05781Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 1 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 5 harvested samples) · 5 pointer-only (licence)
-
High-Quality Entity Segmentation 10 Nov 2022 · 1 repository · arXiv:2211.05776
-
Hyperbolic Cosine Transformer for LiDAR 3D Object Detection 10 Nov 2022 · 0 repositories · arXiv:2211.05580
-
On Optimizing the Communication of Model Parallelism 10 Nov 2022 · 0 repositories · arXiv:2211.05322
-
OneFormer: One Transformer to Rule Universal Image Segmentation 10 Nov 2022 · 4 repositories · arXiv:2211.06220Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 5 harvested samples)
-
Syntax-Guided Domain Adaptation for Aspect-based Sentiment Analysis 10 Nov 2022 · 0 repositories · arXiv:2211.05457
-
Unifying Flow, Stereo and Depth Estimation 10 Nov 2022 · 1 repository · arXiv:2211.05783
-
VieCap4H-VLSP 2021: ObjectAoA-Enhancing performance of Object Relation Transformer with Attention on Attention for Vietnamese image captioning 10 Nov 2022 · 0 repositories · arXiv:2211.05405
-
BLOOM: A 176B-Parameter Open-Access Multilingual Language Model 9 Nov 2022 · 7 repositories · arXiv:2211.05100Syntology 2 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 9 unverified (of 11 harvested samples)
-
Collateral facilitation in humans and language models 9 Nov 2022 · 1 repository · arXiv:2211.05198Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
Cross-lingual Transfer Learning for Check-worthy Claim Identification over Twitter 9 Nov 2022 · 0 repositories · arXiv:2211.05087
-
Distribution-Aligned Fine-Tuning for Efficient Neural Retrieval 9 Nov 2022 · 0 repositories · arXiv:2211.04942
-
Efficient Large-scale Audio Tagging via Transformer-to-CNN Knowledge Distillation 9 Nov 2022 · 2 repositories · arXiv:2211.04772Syntology community repositories only · 2 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples) · 1 pointer-only (licence)
-
Efficiently Scaling Transformer Inference 9 Nov 2022 · 0 repositories · arXiv:2211.05102
-
Large Language Models with Controllable Working Memory 9 Nov 2022 · 0 repositories · arXiv:2211.05110
-
Mask More and Mask Later: Efficient Pre-training of Masked Language Models by Disentangling the [MASK] Token 9 Nov 2022 · 1 repository · arXiv:2211.04898
-
Masked Vision-Language Transformers for Scene Text Recognition 9 Nov 2022 · 1 repository · arXiv:2211.04785
-
Pure Transformer with Integrated Experts for Scene Text Recognition 9 Nov 2022 · 0 repositories · arXiv:2211.04963
-
Sentiment Analysis of Persian Language: Review of Algorithms, Approaches and Datasets 9 Nov 2022 · 0 repositories · arXiv:2212.06041
-
SG-Shuffle: Multi-aspect Shuffle Transformer for Scene Graph Generation 9 Nov 2022 · 0 repositories · arXiv:2211.04773
-
Towards Reasoning-Aware Explainable VQA 9 Nov 2022 · 0 repositories · arXiv:2211.05190
-
Training a Vision Transformer from scratch in less than 24 hours with 1 GPU 9 Nov 2022 · 1 repository · arXiv:2211.05187Syntology 4 ran (of which 4 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified; every one of the 4 samples that ran constructed an object rather than computing a result (of 6 harvested samples) · 6 pointer-only (licence)
-
Transformers Meet Small Datasets 9 Nov 2022 · 0 repositories
-
ViTALiTy: Unifying Low-rank and Sparse Approximation for Vision Transformer Acceleration with a Linear Taylor Attention 9 Nov 2022 · 1 repository · arXiv:2211.05109Syntology official (archive's flag): 13 ran · 13 ran (of which 0 constructed an object rather than computing a result; 11 with no instrument failure: 1 honoured, 0 violated, 10 with no contract checked; 2 where Syntology's instrument failed) · 1 unverified (of 14 harvested samples) · 14 pointer-only (licence)
-
A Multimodal Approach for Dementia Detection from Spontaneous Speech with Tensor Fusion Layer 8 Nov 2022 · 0 repositories · arXiv:2211.04368
-
flexBART: Flexible Bayesian regression trees with categorical predictors 8 Nov 2022 · 1 repository · arXiv:2211.04459
-
Active Example Selection for In-Context Learning 8 Nov 2022 · 1 repository · arXiv:2211.04486Syntology official (archive's flag): 10 ran · 10 ran (of which 3 constructed an object rather than computing a result; 4 with no instrument failure: 1 honoured, 0 violated, 3 with no contract checked; 6 where Syntology's instrument failed) · 6 unverified (of 16 harvested samples)
-
Conciseness: An Overlooked Language Task 8 Nov 2022 · 0 repositories · arXiv:2211.04126
-
Discover, Explanation, Improvement: An Automatic Slice Detection Framework for Natural Language Processing 8 Nov 2022 · 0 repositories · arXiv:2211.04476
-
Linear Self-Attention Approximation via Trainable Feedforward Kernel 8 Nov 2022 · 0 repositories · arXiv:2211.04076
-
Pushing the limits of self-supervised speaker verification using regularized distillation framework 8 Nov 2022 · 1 repository · arXiv:2211.04168
-
SimOn: A Simple Framework for Online Temporal Action Localization 8 Nov 2022 · 1 repository · arXiv:2211.04905
-
Splitting expands the application range of Vision Transformer -- variable Vision Transformer (vViT) 8 Nov 2022 · 0 repositories · arXiv:2211.03992
-
AD-BERT: Using Pre-trained contextualized embeddings to Predict the Progression from Mild Cognitive Impairment to Alzheimer's Disease 7 Nov 2022 · 0 repositories · arXiv:2212.06042
-
Retrieval augmentation of large language models for lay language generation 7 Nov 2022 · 1 repository · arXiv:2211.03818Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 7 harvested samples)
-
CoNMix for Source-free Single and Multi-target Domain Adaptation 7 Nov 2022 · 1 repository · arXiv:2211.03876Syntology official: harvested, nothing ran · 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 2 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
MogaNet: Multi-order Gated Aggregation Network 7 Nov 2022 · 7 repositories · arXiv:2211.03295Syntology official (archive's flag): 12 ran · 12 ran (of which 7 constructed an object rather than computing a result; 10 with no instrument failure: 0 honoured, 0 violated, 10 with no contract checked; 2 where Syntology's instrument failed) · 3 unverified (of 15 harvested samples)
-
Group DETR v2: Strong Object Detector with Encoder-Decoder Pretraining 7 Nov 2022 · 0 repositories · arXiv:2211.03594
-
How Much Does Attention Actually Attend? Questioning the Importance of Attention in Pretrained Transformers 7 Nov 2022 · 1 repository · arXiv:2211.03495
-
Sequential Transformer for End-to-End Person Search 6 Nov 2022 · 0 repositories · arXiv:2211.04323
-
Suffix Retrieval-Augmented Language Modeling 6 Nov 2022 · 1 repository · arXiv:2211.03053
-
Wall Street Tree Search: Risk-Aware Planning for Offline Reinforcement Learning 6 Nov 2022 · 0 repositories · arXiv:2211.04583
-
Inductive Graph Transformer for Delivery Time Estimation 5 Nov 2022 · 1 repository · arXiv:2211.02863
-
Learning to Infer from Unlabeled Data: A Semi-supervised Learning Approach for Robust Natural Language Inference 5 Nov 2022 · 1 repository · arXiv:2211.02971
-
A Transformer Architecture for Online Gesture Recognition of Mathematical Expressions 4 Nov 2022 · 0 repositories · arXiv:2211.02643
-
A Weakly-Supervised Streaming Multilingual Speech Model with Truly Zero-Shot Capability 4 Nov 2022 · 0 repositories · arXiv:2211.02499
-
BERT-Deep CNN: State-of-the-Art for Sentiment Analysis of COVID-19 Tweets 4 Nov 2022 · 0 repositories · arXiv:2211.09733