Methods › General › Attention Mechanisms › Attention › Papers, page 238
Attention Is All You Need
Attention
Papers archive 2025-07-28
archive papers tagged: 31,583 · with a code link: 13,473 · where Syntology ran a sample: 3,998 (3,366 with a run with no instrument failure, 632 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (3,998 of 31,583 tagged: 3,366 with a run with no instrument failure, 632 where every run was a failure of Syntology's instrument)
Page 238 of 316: papers 23,701 to 23,800 of 31,583, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
NaturalProver: Grounded Mathematical Proof Generation with Language Models 25 May 2022 · 1 repository · arXiv:2205.12910Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
ORCA: Interpreting Prompted Language Models via Locating Supporting Data Evidence in the Ocean of Pretraining Data 25 May 2022 · 0 repositories · arXiv:2205.12600
-
RobustLR: Evaluating Robustness to Logical Perturbation in Deductive Reasoning 25 May 2022 · 1 repository · arXiv:2205.12598Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 5 harvested samples)
-
TaCube: Pre-computing Data Cubes for Answering Numerical-Reasoning Questions over Tabular Data 25 May 2022 · 1 repository · arXiv:2205.12682
-
Text-to-Face Generation with StyleGAN2 25 May 2022 · 0 repositories · arXiv:2205.12512
-
Do we need Label Regularization to Fine-tune Pre-trained Language Models? 25 May 2022 · 0 repositories · arXiv:2205.12428
-
Train Flat, Then Compress: Sharpness-Aware Minimization Learns More Compressible Models 25 May 2022 · 0 repositories · arXiv:2205.12694
-
Transcormer: Transformer for Sentence Scoring with Sliding Language Modeling 25 May 2022 · 1 repository · arXiv:2205.12986Syntology 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
VTP: Volumetric Transformer for Multi-view Multi-person 3D Pose Estimation 25 May 2022 · 0 repositories · arXiv:2205.12602
-
VulBERTa: Simplified Source Code Pre-Training for Vulnerability Detection 25 May 2022 · 1 repository · arXiv:2205.12424
-
AdaMix: Mixture-of-Adaptations for Parameter-efficient Model Tuning 24 May 2022 · 1 repository · arXiv:2205.12410
-
EdiT5: Semi-Autoregressive Text-Editing with T5 Warm-Start 24 May 2022 · 0 repositories · arXiv:2205.12209
-
FLUTE: Figurative Language Understanding through Textual Explanations 24 May 2022 · 1 repository · arXiv:2205.12404
-
Formulating Few-shot Fine-tuning Towards Language Model Pre-training: A Pilot Study on Named Entity Recognition 24 May 2022 · 1 repository · arXiv:2205.11799
-
Garden-Path Traversal in GPT-2 24 May 2022 · 1 repository · arXiv:2205.12302Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
GeoMLAMA: Geo-Diverse Commonsense Probing on Multilingual Pre-Trained Language Models 24 May 2022 · 1 repository · arXiv:2205.12247
-
History Compression via Language Models in Reinforcement Learning 24 May 2022 · 2 repositories · arXiv:2205.12258Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 1 where Syntology's instrument failed) · 2 unverified (of 6 harvested samples) · 6 pointer-only (licence)
-
The Authenticity Gap in Human Evaluation 24 May 2022 · 0 repositories · arXiv:2205.11930
-
K-12BERT: BERT for K-12 education 24 May 2022 · 1 repository · arXiv:2205.12335
-
Medical Scientific Table-to-Text Generation with Human-in-the-Loop under the Data Sparsity Constraint 24 May 2022 · 0 repositories · arXiv:2205.12368
-
Multi-Level Modeling Units for End-to-End Mandarin Speech Recognition 24 May 2022 · 0 repositories · arXiv:2205.11998
-
On the Role of Bidirectionality in Language Model Pre-Training 24 May 2022 · 0 repositories · arXiv:2205.11726
-
Partial-input baselines show that NLI models can ignore context, but they don't 24 May 2022 · 1 repository · arXiv:2205.12181
-
Privacy-Preserving Image Classification Using Vision Transformer 24 May 2022 · 0 repositories · arXiv:2205.12041
-
RetroMAE: Pre-Training Retrieval-oriented Language Models Via Masked Auto-Encoder 24 May 2022 · 1 repository · arXiv:2205.12035
-
Sparse Mixers: Combining MoE and Mixing to build a more efficient BERT 24 May 2022 · 1 repository · arXiv:2205.12399
-
Symbolic Expression Transformer: A Computer Vision Approach for Symbolic Regression 24 May 2022 · 0 repositories · arXiv:2205.11798
-
UMSNet: An Universal Multi-sensor Network for Human Activity Recognition 24 May 2022 · 0 repositories · arXiv:2205.11756
-
Word-order typology in Multilingual BERT: A case study in subordinate-clause detection 24 May 2022 · 0 repositories · arXiv:2205.11987
-
Workflow Discovery from Dialogues in the Low Data Regime 24 May 2022 · 1 repository · arXiv:2205.11690
-
A Question-Answer Driven Approach to Reveal Affirmative Interpretations from Verbal Negations 23 May 2022 · 1 repository · arXiv:2205.11467
-
Accurate and Resource-Efficient Lipreading with Efficientnetv2 and Transformers 23 May 2022 · 0 repositories
-
Artificial intelligence for topic modelling in Hindu philosophy: mapping themes between the Upanishads and the Bhagavad Gita 23 May 2022 · 1 repository · arXiv:2205.11020
-
BanglaNLG and BanglaT5: Benchmarks and Resources for Evaluating Low-Resource Natural Language Generation in Bangla 23 May 2022 · 2 repositories · arXiv:2205.11081
-
BolT: Fused Window Transformers for fMRI Time Series Analysis 23 May 2022 · 1 repository · arXiv:2205.11578
-
DistilCamemBERT: a distillation of the French model CamemBERT 23 May 2022 · 0 repositories · arXiv:2205.11111
-
Improving Short Text Classification With Augmented Data Using GPT-3 23 May 2022 · 0 repositories · arXiv:2205.10981
-
KOLD: Korean Offensive Language Dataset 23 May 2022 · 1 repository · arXiv:2205.11315
-
Learning to Ignore Adversarial Attacks 23 May 2022 · 0 repositories · arXiv:2205.11551
-
Looking for a Handsome Carpenter! Debiasing GPT-3 Job Advertisements 23 May 2022 · 1 repository · arXiv:2205.11374
-
On the Paradox of Learning to Reason from Data 23 May 2022 · 1 repository · arXiv:2205.11502Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 10 harvested samples)
-
Outliers Dimensions that Disrupt Transformers Are Driven by Frequency 23 May 2022 · 1 repository · arXiv:2205.11380Syntology official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 8 with no instrument failure: 0 honoured, 0 violated, 8 with no contract checked; 0 where Syntology's instrument failed) · 8 unverified (of 16 harvested samples)
-
Parameter-Efficient Sparsity for Large Language Models Fine-Tuning 23 May 2022 · 2 repositories · arXiv:2205.11005Syntology official (archive's flag): 2 ran · 2 ran (of which 2 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified; every one of the 2 samples that ran constructed an object rather than computing a result (of 2 harvested samples)
-
Penguins Don't Fly: Reasoning about Generics through Instantiations and Exceptions 23 May 2022 · 0 repositories · arXiv:2205.11658
-
Prompt Tuning for Discriminative Pre-trained Language Models 23 May 2022 · 1 repository · arXiv:2205.11166Syntology official (archive's flag): 7 ran · 7 ran (of which 4 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 2 where Syntology's instrument failed) · 5 unverified (of 12 harvested samples)
-
RL with KL penalties is better viewed as Bayesian inference 23 May 2022 · 0 repositories · arXiv:2205.11275
-
The Diminishing Returns of Masked Language Models to Science 23 May 2022 · 0 repositories · arXiv:2205.11342
-
SelfReformer: Self-Refined Network with Transformer for Salient Object Detection 23 May 2022 · 1 repository · arXiv:2205.11283Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 1 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
Simple Recurrence Improves Masked Language Models 23 May 2022 · 0 repositories · arXiv:2205.11588
-
Super Vision Transformer 23 May 2022 · 1 repository · arXiv:2205.11397
-
TempLM: Distilling Language Models into Template-Based Generators 23 May 2022 · 1 repository · arXiv:2205.11055
-
Time-series Transformer Generative Adversarial Networks 23 May 2022 · 5 repositories · arXiv:2205.11164
-
Use of Transformer-Based Models for Word-Level Transliteration of the Book of the Dean of Lismore 23 May 2022 · 0 repositories · arXiv:2205.11370
-
A Graph Enhanced BERT Model for Event Prediction 22 May 2022 · 0 repositories · arXiv:2205.10822
-
Dynamic Query Selection for Fast Visual Perceiver 22 May 2022 · 0 repositories · arXiv:2205.10873
-
GraphMAE: Self-Supervised Masked Graph Autoencoders 22 May 2022 · 3 repositories · arXiv:2205.10803Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
Instruction Induction: From Few Examples to Natural Language Task Descriptions 22 May 2022 · 1 repository · arXiv:2205.10782Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples)
-
Relphormer: Relational Graph Transformer for Knowledge Graph Representations 22 May 2022 · 1 repository · arXiv:2205.10852Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
A Study on Transformer Configuration and Training Objective 21 May 2022 · 0 repositories · arXiv:2205.10505
-
DProQ: A Gated-Graph Transformer for Protein Complex Structure Assessment 21 May 2022 · 1 repository · arXiv:2205.10627
-
HLATR: Enhance Multi-stage Text Retrieval with Hybrid List Aware Transformer Reranking 21 May 2022 · 1 repository · arXiv:2205.10569
-
Least-to-Most Prompting Enables Complex Reasoning in Large Language Models 21 May 2022 · 1 repository · arXiv:2205.10625
-
Life after BERT: What do Other Muppets Understand about Language? 21 May 2022 · 1 repository · arXiv:2205.10696
-
Pre-training Data Quality and Quantity for a Low-Resource Language: New Corpus and BERT Models for Maltese 21 May 2022 · 1 repository · arXiv:2205.10517
-
Transformer based Generative Adversarial Network for Liver Segmentation 21 May 2022 · 1 repository · arXiv:2205.10663
-
Visualizing CoAtNet Predictions for Aiding Melanoma Detection 21 May 2022 · 0 repositories · arXiv:2205.10515
-
Self-Supervised Time Series Representation Learning via Cross Reconstruction Transformer 20 May 2022 · 1 repository · arXiv:2205.09928
-
Current Trends and Approaches in Synonyms Extraction: Potential Adaptation to Arabic 20 May 2022 · 0 repositories · arXiv:2205.10412
-
Degradation-Aware Unfolding Half-Shuffle Transformer for Spectral Compressive Imaging 20 May 2022 · 1 repository · arXiv:2205.10102Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 3 where Syntology's instrument failed) · 4 unverified (of 8 harvested samples)
-
Exploring Extreme Parameter Compression for Pre-trained Language Models 20 May 2022 · 1 repository · arXiv:2205.10036
-
Learning to Count Anything: Reference-less Class-agnostic Counting with Weak Supervision 20 May 2022 · 2 repositories · arXiv:2205.10203Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 1 where Syntology's instrument failed) · 2 unverified (of 5 harvested samples) · 2 pointer-only (licence)
-
Lossless Acceleration for Seq2seq Generation with Aggressive Decoding 20 May 2022 · 2 repositories · arXiv:2205.10350
-
Mask-guided Vision Transformer (MG-ViT) for Few-Shot Learning 20 May 2022 · 0 repositories · arXiv:2205.09995
-
MSTRIQ: No Reference Image Quality Assessment Based on Swin Transformer with Multi-Stage Fusion 20 May 2022 · 0 repositories · arXiv:2205.10101
-
Pre-training Transformer Models with Sentence-Level Objectives for Answer Sentence Selection 20 May 2022 · 0 repositories · arXiv:2205.10455
-
Progressive Class Semantic Matching for Semi-supervised Text Classification 20 May 2022 · 1 repository · arXiv:2205.10189
-
Prototypical Calibration for Few-shot Learning of Language Models 20 May 2022 · 1 repository · arXiv:2205.10183
-
Temporally Precise Action Spotting in Soccer Videos Using Dense Detection Anchors 20 May 2022 · 1 repository · arXiv:2205.10450
-
Translating Hanja Historical Documents to Contemporary Korean and English 20 May 2022 · 0 repositories · arXiv:2205.10019
-
Uniform Masking: Enabling MAE Pre-training for Pyramid-based Vision Transformers with Locality 20 May 2022 · 1 repository · arXiv:2205.10063
-
A graph-transformer for whole slide image classification 19 May 2022 · 1 repository · arXiv:2205.09671Syntology official (archive's flag): 10 ran · 10 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 1 honoured, 0 violated, 8 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 10 harvested samples) · 1 pointer-only (licence)
-
Acceptability Judgements via Examining the Topology of Attention Maps 19 May 2022 · 1 repository · arXiv:2205.09630
-
ArabGlossBERT: Fine-Tuning BERT on Context-Gloss Pairs for WSD 19 May 2022 · 0 repositories · arXiv:2205.09685
-
Automated Scoring for Reading Comprehension via In-context BERT Tuning 19 May 2022 · 1 repository · arXiv:2205.09864
-
BabyNet: Residual Transformer Module for Birth Weight Prediction on Fetal Ultrasound Video 19 May 2022 · 1 repository · arXiv:2205.09382
-
Cross-Enhancement Transformer for Action Segmentation 19 May 2022 · 1 repository · arXiv:2205.09445
-
Insights on Neural Representations for End-to-End Speech Recognition 19 May 2022 · 0 repositories · arXiv:2205.09456
-
Overcoming Language Disparity in Online Content Classification with Multimodal Learning 19 May 2022 · 1 repository · arXiv:2205.09744
-
Psychiatric Scale Guided Risky Post Screening for Early Detection of Depression 19 May 2022 · 1 repository · arXiv:2205.09497
-
Towards Understanding Gender-Seniority Compound Bias in Natural Language Generation 19 May 2022 · 1 repository · arXiv:2205.09830
-
Transformer with Memory Replay 19 May 2022 · 0 repositories · arXiv:2205.09869
-
Transformers as Neural Augmentors: Class Conditional Sentence Generation via Variational Bayes 19 May 2022 · 1 repository · arXiv:2205.09391
-
TransTab: Learning Transferable Tabular Transformers Across Tables 19 May 2022 · 1 repository · arXiv:2205.09328Syntology official (archive's flag): 9 ran · 9 ran (of which 7 constructed an object rather than computing a result; 8 with no instrument failure: 0 honoured, 0 violated, 8 with no contract checked; 1 where Syntology's instrument failed) · 2 unverified (of 11 harvested samples)
-
TRT-ViT: TensorRT-oriented Vision Transformer 19 May 2022 · 0 repositories · arXiv:2205.09579
-
VNT-Net: Rotational Invariant Vector Neuron Transformers 19 May 2022 · 0 repositories · arXiv:2205.09690
-
AdaMCT: Adaptive Mixture of CNN-Transformer for Sequential Recommendation 18 May 2022 · 1 repository · arXiv:2205.08776Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
Entity Alignment with Reliable Path Reasoning and Relation-Aware Heterogeneous Graph Transformer 18 May 2022 · 0 repositories · arXiv:2205.08806
-
Evaluation of Transfer Learning for Polish with a Text-to-Text Model 18 May 2022 · 0 repositories · arXiv:2205.08808
-
Learning Rate Curriculum 18 May 2022 · 1 repository · arXiv:2205.09180
-
Leveraging Pseudo-labeled Data to Improve Direct Speech-to-Speech Translation 18 May 2022 · 1 repository · arXiv:2205.08993Syntology official (archive's flag): 5 ran · 5 ran (of which 1 constructed an object rather than computing a result; 5 with no instrument failure: 1 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 5 harvested samples) · 5 pointer-only (licence)