Methods › General › Attention Mechanisms › Attention › Papers, page 255
Attention Is All You Need
Attention
Papers archive 2025-07-28
archive papers tagged: 31,583 · with a code link: 13,473 · where Syntology ran a sample: 3,998 (3,366 with a run with no instrument failure, 632 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (3,998 of 31,583 tagged: 3,366 with a run with no instrument failure, 632 where every run was a failure of Syntology's instrument)
Page 255 of 316: papers 25,401 to 25,500 of 31,583, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
A Structured Semantic Reinforcement method for Task-Oriented Dialogue 16 Nov 2021 · 0 repositories
-
Active Dialogue Simulation in Conversational Systems 16 Nov 2021 · 0 repositories
-
AdapLeR: Speeding up Inference by Adaptive Length Reduction 16 Nov 2021 · 0 repositories
-
All Birds with One Stone: Multi-task Learning for Inference with One Forward Pass 16 Nov 2021 · 0 repositories
-
An Empirical Study of Document-to-document Neural Machine Translation 16 Nov 2021 · 0 repositories
-
An Information Theoretic Measurement of Topical Relevance in Learner Essays 16 Nov 2021 · 0 repositories
-
An Isotropy Analysis in the Multilingual BERT Embedding Space 16 Nov 2021 · 0 repositories
-
ANNA: Enhanced Language Representation for Question Answering 16 Nov 2021 · 0 repositories
-
Assessing the Coherence Modeling Capabilities of Pretrained Transformer-based Language Models 16 Nov 2021 · 0 repositories
-
Attention-based Multi-hypothesis Fusion for Speech Summarization 16 Nov 2021 · 2 repositories · arXiv:2111.08201
-
BARCOR: Towards A Unified Framework for Conversational Recommendation 16 Nov 2021 · 0 repositories
-
BERT got a Date: Introducing Transformers to Temporal Tagging 16 Nov 2021 · 0 repositories
-
BERT is Robust! A Case Against Synonym-Based Adversarial Examples in Text Classification 16 Nov 2021 · 0 repositories
-
BigFive: A Dataset of Coarse- and Fine-Grained Personality Characteristics 16 Nov 2021 · 0 repositories
-
BORT: Back and Denoising Reconstruction for End-to-End Task-Oriented Dialog 16 Nov 2021 · 0 repositories
-
Building Chinese Biomedical Language Models via Multi-Level Text Discrimination 16 Nov 2021 · 1 repository
-
CalBERT - Code-mixed Adaptive Language representations using BERT 16 Nov 2021 · 0 repositories
-
Causal Transformers: Improving the Robustness on Spurious Correlations 16 Nov 2021 · 0 repositories
-
Challenges for Open-domain Targeted Sentiment Analysis 16 Nov 2021 · 0 repositories
-
Chinese Word Attention based on Valid Division of Sentence 16 Nov 2021 · 0 repositories
-
Coloring the Blank Slate: Pre-training Imparts a Hierarchical Inductive Bias to Sequence-to-sequence Models 16 Nov 2021 · 0 repositories
-
Compressing Sentence Representation via Homomorphic Projective Distillation 16 Nov 2021 · 0 repositories
-
Constructing Phrase-level Semantic Labels to Form Multi-GrainedSupervision for Image-Text Retrieval 16 Nov 2021 · 0 repositories
-
Contextualized Sensorimotor Norms: multi-dimensional measures of sensorimotor strength for ambiguous English words, in context 16 Nov 2021 · 0 repositories
-
CREATE: A Benchmark for Chinese Short Video Retrieval and Title Generation 16 Nov 2021 · 0 repositories
-
Cross-domain Named Entity Recognition via Graph Matching 16 Nov 2021 · 0 repositories
-
CST5: Data augmentation for Code-Switched Semantic Parsing 16 Nov 2021 · 0 repositories
-
CVSS-BERT: Explainable Natural Language Processing to Determine the Severity of a Computer Security Vulnerability from its Description 16 Nov 2021 · 1 repository · arXiv:2111.08510
-
DAML-ST5: Low Resource Style Transfer via Domain Adaptive Meta Learning 16 Nov 2021 · 0 repositories
-
Data Augmentation and Learned Layer Aggregation for Improved Multilingual Language Understanding in Dialogue 16 Nov 2021 · 0 repositories
-
Data Augmentation for Intent Classification with Generic Large Language Models 16 Nov 2021 · 0 repositories
-
Data Contamination: From Memorization to Exploitation 16 Nov 2021 · 0 repositories
-
DAWSON: Data Augmentation using Weak Supervision On Natural Language 16 Nov 2021 · 0 repositories
-
Deep-to-bottom Weights Decay: A Systemic Knowledge Review Learning Technique for Transformer Layers in Knowledge Distillation 16 Nov 2021 · 0 repositories
-
Discontinuous Constituency and BERT: A Case Study of Dutch 16 Nov 2021 · 0 repositories
-
Efficient Long Sequence Encoding via Synchronization 16 Nov 2021 · 0 repositories
-
ElitePLM: An Empirical Study on General Language Ability Evaluation of Pretrained Language Models 16 Nov 2021 · 0 repositories
-
ELLE: Efficient Lifelong Pre-training for Emerging Data 16 Nov 2021 · 0 repositories
-
Empirical Analysis of Training Strategies of Transformer-based Japanese Chit-chat Systems 16 Nov 2021 · 0 repositories
-
End-To-End Sign Language Translation via Multitask Learning 16 Nov 2021 · 0 repositories
-
End-to-end Task-oriented Dialog Policy Learning based on Pre-trained Language Model 16 Nov 2021 · 0 repositories
-
Enhancing the Nonlinear Mutual Dependencies in Transformers with Mutual Information 16 Nov 2021 · 0 repositories
-
ERNIE-SPARSE: Learning Hierarchical Efficient Transformer Through Regularized Self-Attention 16 Nov 2021 · 0 repositories
-
Event Detection via Derangement Question Answering 16 Nov 2021 · 0 repositories
-
EventBERT 16 Nov 2021 · 0 repositories
-
Explicit Modeling the Context for Chinese NER 16 Nov 2021 · 0 repositories
-
Exploring and Adapting Chinese GPT to Pinyin Input Method 16 Nov 2021 · 0 repositories
-
Extreme Multi-label Text Classification with Multi-layer Experts 16 Nov 2021 · 0 repositories
-
Eye Gaze and Self-attention: How Humans and Transformers Attend Words in Sentences 16 Nov 2021 · 0 repositories
-
Fast and Accurate Transformer-based Translation with Character-Level Encoding and Subword-Level Decoding 16 Nov 2021 · 0 repositories
-
Feature-rich Open-vocabulary Interpretable Neural Representations for All of the World’s 7000 Languages 16 Nov 2021 · 0 repositories
-
Feature Structure Distillation for BERT Transferring 16 Nov 2021 · 0 repositories
-
Fix Bugs with Transformer through a Neural-Symbolic Edit Grammar 16 Nov 2021 · 0 repositories
-
FRSUM: Towards Faithful Abstractive Summarization via Enhancing Factual Robustness 16 Nov 2021 · 0 repositories
-
gaBERT — an Irish Language Model 16 Nov 2021 · 0 repositories
-
Generating Diverse and High-Quality Abstractive Summaries with Variational Transformers 16 Nov 2021 · 0 repositories
-
Generation of News Articles from Tweets : An Experiment 16 Nov 2021 · 0 repositories
-
Generative Pre-Trained Transformer for Design Concept Generation: An Exploration 16 Nov 2021 · 0 repositories · arXiv:2111.08489
-
Get the Point! Graph Enhanced Candidate Retrieval for Zero-shot Entity Linking 16 Nov 2021 · 0 repositories
-
GLM: General Language Model Pretraining with Autoregressive Blank Infilling 16 Nov 2021 · 0 repositories
-
Graph-based Fine-grained Multimodal Attention Mechanism for Sentiment Analysis 16 Nov 2021 · 0 repositories
-
How does the pre-training objective affect what large language models learn about linguistic properties? 16 Nov 2021 · 0 repositories
-
Impact of Tokenization on Language Models: An Analysis for Turkish 16 Nov 2021 · 0 repositories
-
Improving Compositional Generalization with Self-Training for Data-to-Text Generation 16 Nov 2021 · 0 repositories
-
Improving GPT-3 after deployment with a dynamic memory of feedback 16 Nov 2021 · 0 repositories
-
Improving Neural Models for Radiology Report Retrieval with Lexicon-based Automated Annotation 16 Nov 2021 · 0 repositories
-
Improving Unsupervised Sentence Simplification Using Fine-Tuned Masked Language Models 16 Nov 2021 · 0 repositories
-
Input-specific Attention Subnetworks for Adversarial Detection 16 Nov 2021 · 0 repositories
-
Integrated Semantic and Phonetic Post-correction for Chinese Speech Recognition 16 Nov 2021 · 1 repository · arXiv:2111.08400
-
Interpreting Language Models Through Knowledge Graph Extraction 16 Nov 2021 · 1 repository · arXiv:2111.08546
-
Interpreting the Robustness of Neural NLP Models to Textual Perturbations 16 Nov 2021 · 0 repositories
-
Investigating the Use of BERT Anchors for Bilingual Lexicon Induction with Minimal Supervision 16 Nov 2021 · 0 repositories
-
Is Neural Topic Modelling Better than Clustering? An Empirical Study on Clustering with Contextual Embeddings for Topics 16 Nov 2021 · 0 repositories
-
"Is Whole Word Masking Always Better for Chinese BERT?": Probing on Chinese Grammatical Error Correction 16 Nov 2021 · 0 repositories
-
KinyaBERT: a Morphology-aware Kinyarwanda Language Model 16 Nov 2021 · 0 repositories
-
Knowledge Enhanced Embedding: Improve Model Generalization Through Knowledge Graphs 16 Nov 2021 · 0 repositories
-
Knowledge Graph is in Rescue: Task Oriented Dialogue System for Response Generation without NLU and DM 16 Nov 2021 · 0 repositories
-
Knowledge-guided Transformer for Joint Theme and Emotion Classification of Chinese Classical Poetry 16 Nov 2021 · 0 repositories
-
Language Level Classification on German Texts using a Neural Approach 16 Nov 2021 · 0 repositories
-
Learn More from Less: Improving Conversational Recommender Systems via Contextual and Time-Aware Modeling 16 Nov 2021 · 0 repositories
-
Learning Methods for Solving Astronomy Course Problems 16 Nov 2021 · 0 repositories
-
Learning Non-Autoregressive Models from Search for Unsupervised Sentence Summarization 16 Nov 2021 · 0 repositories
-
Learning to Ignore Adversarial Attacks 16 Nov 2021 · 0 repositories
-
Life after BERT: What do Other Muppets Understand about Language? 16 Nov 2021 · 0 repositories
-
Listen to Both Sides and be Enlightened! -- Hierarchical Modality Fusion Network for Entity and Relation Extraction 16 Nov 2021 · 0 repositories
-
Looking Into the Black Box - How Are Idioms Processed in BERT? 16 Nov 2021 · 0 repositories
-
LordBERT: Embedding Long Text by Segment Ordering with BERT 16 Nov 2021 · 0 repositories
-
Making Transformers Solve Compositional Tasks 16 Nov 2021 · 0 repositories
-
MarCQAp: Effective Context Modeling for Conversational Question Answering 16 Nov 2021 · 0 repositories
-
MarkBERT: Marking Word Boundaries Improves Chinese BERT 16 Nov 2021 · 1 repository
-
MDERank: A Masked Document Embedding Rank Approach for Unsupervised Keyphrase Extraction 16 Nov 2021 · 0 repositories
-
Meeting Summarization with Pre-training and Clustering Methods 16 Nov 2021 · 1 repository · arXiv:2111.08210
-
Metadata Shaping: Natural Language Annotations for the Long Tail 16 Nov 2021 · 0 repositories
-
Modeling Hierarchical Syntax Structure with Triplet Position for Source Code Summarization 16 Nov 2021 · 1 repository
-
Moving the Eiffel Tower to ROME: Tracing and Editing Facts in GPT 16 Nov 2021 · 0 repositories
-
Multi-head or Single-head? An Empirical Comparison for Transformer Training 16 Nov 2021 · 0 repositories
-
N-grammer: Augmenting Transformers with latent n-grams 16 Nov 2021 · 0 repositories
-
Neural Keyphrase Generation: Analysis and Evaluation 16 Nov 2021 · 0 repositories
-
NSP-BERT: A Prompt-based Zero-Shot Learner Through an Original Pre-training Task —— Next Sentence Prediction 16 Nov 2021 · 0 repositories
-
On the Multilingual Capabilities of Very Large-Scale English Language Models 16 Nov 2021 · 0 repositories