Methods › General › Attention Modules › Multi-Head Attention › Papers, page 239
Multi-Head Attention
Papers archive 2025-07-28
archive papers tagged: 24,855 · with a code link: 11,214 · where Syntology ran a sample: 3,454 (2,916 with a run with no instrument failure, 538 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (3,454 of 24,855 tagged: 2,916 with a run with no instrument failure, 538 where every run was a failure of Syntology's instrument)
Page 239 of 249: papers 23,801 to 23,900 of 24,855, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Deepening Hidden Representations from Pre-trained Language Models 5 Nov 2019 · 0 repositories · arXiv:1911.01940
-
Efficient Multi-robot Exploration via Multi-head Attention-based Cooperation Strategy 5 Nov 2019 · 0 repositories · arXiv:1911.01774
-
Improving Bidirectional Decoding with Dynamic Target Semantics in Neural Machine Translation 5 Nov 2019 · 0 repositories · arXiv:1911.01597
-
Improving Slot Filling by Utilizing Contextual Information 5 Nov 2019 · 0 repositories · arXiv:1911.01680
-
Incremental Sense Weight Training for the Interpretation of Contextualized Word Embeddings 5 Nov 2019 · 0 repositories · arXiv:1911.01623
-
MML: Maximal Multiverse Learning for Robust Fine-Tuning of Language Models 5 Nov 2019 · 1 repository · arXiv:1911.06182
-
Unsupervised Cross-lingual Representation Learning at Scale 5 Nov 2019 · 35 repositories · arXiv:1911.02116Syntology official (archive's flag): 12 ran · 40 ran (of which 13 constructed an object rather than computing a result; 35 with no instrument failure: 4 honoured, 1 violated, 30 with no contract checked; 5 where Syntology's instrument failed) · 19 unverified (of 59 harvested samples) · 52 pointer-only (licence)
-
Assessing Social and Intersectional Biases in Contextualized Word Representations 4 Nov 2019 · 1 repository · arXiv:1911.01485
-
BAS: An Answer Selection Method Using BERT Language Model 4 Nov 2019 · 0 repositories · arXiv:1911.01528
-
An Algorithm for Routing Capsules in All Domains 2 Nov 2019 · 1 repository · arXiv:1911.00792
-
Machine Translation Evaluation using Bi-directional Entailment 2 Nov 2019 · 0 repositories · arXiv:1911.00681
-
Sentence-Level BERT and Multi-Task Learning of Age and Gender in Social Media 2 Nov 2019 · 0 repositories · arXiv:1911.00637
-
A Deep Learning-Based System for PharmaCoNER 1 Nov 2019 · 0 repositories
-
A Multi-Task Learning Framework for Extracting Bacteria Biotope Information 1 Nov 2019 · 0 repositories
-
A Recurrent BERT-based Model for Question Generation 1 Nov 2019 · 1 repository
-
Aggregating Bidirectional Encoder Representations Using MatchLSTM for Sequence Matching 1 Nov 2019 · 0 repositories
-
Applying BERT to Document Retrieval with Birch 1 Nov 2019 · 0 repositories
-
Automatically Extracting Challenge Sets for Non-Local Phenomena in Neural Machine Translation 1 Nov 2019 · 0 repositories
-
BERT Goes to Law School: Quantifying the Competitive Advantage of Access to Large Legal Corpora in Contract Understanding 1 Nov 2019 · 0 repositories · arXiv:1911.00473
-
BERT is Not an Interlingua and the Bias of Tokenization 1 Nov 2019 · 1 repository
-
Biomedical Named Entity Recognition with Multilingual BERT 1 Nov 2019 · 1 repository
-
BLCU-NLP at COIN-Shared Task1: Stagewise Fine-tuning BERT for Commonsense Inference in Everyday Narrations 1 Nov 2019 · 0 repositories
-
CAUnLP at NLP4IF 2019 Shared Task: Context-Dependent BERT for Sentence-Level Propaganda Detection 1 Nov 2019 · 0 repositories
-
CLER: Cross-task Learning with Expert Representation to Generalize Reading and Understanding 1 Nov 2019 · 0 repositories
-
Combining Global Sparse Gradients with Local Gradients in Distributed Neural Network Training 1 Nov 2019 · 0 repositories
-
Combining Unsupervised Pre-training and Annotator Rationales to Improve Low-shot Text Classification 1 Nov 2019 · 0 repositories
-
Contextualized Cross-Lingual Event Trigger Extraction with Minimal Resources 1 Nov 2019 · 0 repositories
-
Cost-Sensitive BERT for Generalisable Sentence Classification on Imbalanced Data 1 Nov 2019 · 0 repositories
-
Cross-Domain Modeling of Sentence-Level Evidence for Document Retrieval 1 Nov 2019 · 0 repositories
-
CVIT's submissions to WAT-2019 1 Nov 2019 · 0 repositories
-
Deep Bidirectional Transformers for Relation Extraction without Supervision 1 Nov 2019 · 0 repositories · arXiv:1911.00313
-
Detection of Propaganda Using Logistic Regression 1 Nov 2019 · 0 repositories
-
Dialect Text Normalization to Normative Standard Finnish 1 Nov 2019 · 1 repository
-
DialoGPT: Large-Scale Generative Pre-training for Conversational Response Generation 1 Nov 2019 · 6 repositories · arXiv:1911.00536Syntology community repositories only · 6 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 1 honoured, 0 violated, 4 with no contract checked; 1 where Syntology's instrument failed) · 6 unverified (of 12 harvested samples) · 12 pointer-only (licence)
-
Divisive Language and Propaganda Detection using Multi-head Attention Transformers with Deep Learning BERT-based Language Models for Binary Classification 1 Nov 2019 · 0 repositories
-
Domain Adaptation with BERT-based Domain Classification and Data Selection 1 Nov 2019 · 0 repositories
-
English to Hindi Multi-modal Neural Machine Translation and Hindi Image Captioning 1 Nov 2019 · 0 repositories
-
Enhanced Transformer Model for Data-to-Text Generation 1 Nov 2019 · 0 repositories
-
Enhancing BERT for Lexical Normalization 1 Nov 2019 · 0 repositories
-
Evaluating BERT for natural language inference: A case study on the CommitmentBank 1 Nov 2019 · 0 repositories
-
Event Causality Recognition Exploiting Multiple Annotators' Judgments and Background Knowledge 1 Nov 2019 · 0 repositories
-
Extract and Aggregate: A Novel Domain-Independent Approach to Factual Data Verification 1 Nov 2019 · 0 repositories
-
Extractive NarrativeQA with Heuristic Pre-Training 1 Nov 2019 · 0 repositories
-
FASPell: A Fast, Adaptable, Simple, Powerful Chinese Spell Checker Based On DAE-Decoder Paradigm 1 Nov 2019 · 1 repository
-
Fine-Grained Propaganda Detection with Fine-Tuned BERT 1 Nov 2019 · 0 repositories
-
Fine-tune BERT with Sparse Self-Attention Mechanism 1 Nov 2019 · 0 repositories
-
From Monolingual to Multilingual FAQ Assistant using Multilingual Co-training 1 Nov 2019 · 0 repositories
-
Fully Unsupervised Crosslingual Semantic Textual Similarity Metric Based on BERT for Identifying Parallel Data 1 Nov 2019 · 0 repositories
-
GEM: Generative Enhanced Model for adversarial attacks 1 Nov 2019 · 0 repositories
-
Generalizing Question Answering System with Pre-trained Language Model Fine-tuning 1 Nov 2019 · 0 repositories
-
HIT-SCIR at MRP 2019: A Unified Pipeline for Meaning Representation Parsing via Efficient Training and Effective Encoding 1 Nov 2019 · 0 repositories
-
How well do NLI models capture verb veridicality? 1 Nov 2019 · 0 repositories
-
Idiap NMT System for WAT 2019 Multimodal Translation Task 1 Nov 2019 · 0 repositories
-
IIT-KGP at COIN 2019: Using pre-trained Language Models for modeling Machine Comprehension 1 Nov 2019 · 0 repositories
-
Improving Answer Selection and Answer Triggering using Hard Negatives 1 Nov 2019 · 0 repositories
-
Improving Generalization of Transformer for Speech Recognition with Parallel Schedule Sampling and Relative Positional Embedding 1 Nov 2019 · 0 repositories · arXiv:1911.00203
-
Improving Natural Language Understanding by Reverse Mapping Bytepair Encoding 1 Nov 2019 · 0 repositories
-
Improving Pre-Trained Multilingual Model with Vocabulary Expansion 1 Nov 2019 · 0 repositories
-
Inspecting Unification of Encoding and Matching with Transformer: A Case Study of Machine Reading Comprehension 1 Nov 2019 · 0 repositories
-
JBNU at MRP 2019: Multi-level Biaffine Attention for Semantic Dependency Parsing 1 Nov 2019 · 0 repositories
-
Jeff Da at COIN - Shared Task: BIG MOOD: Relating Transformers to Explicit Commonsense Knowledge 1 Nov 2019 · 0 repositories
-
JUSTDeep at NLP4IF 2019 Task 1: Propaganda Detection using Ensemble Deep Learning Models 1 Nov 2019 · 0 repositories
-
LexicalAT: Lexical-Based Adversarial Reinforcement Training for Robust Sentiment Classification 1 Nov 2019 · 0 repositories
-
Long Warm-up and Self-Training: Training Strategies of NICT-2 NMT System at WAT-2019 1 Nov 2019 · 0 repositories
-
LTRC-MT Simple & Effective Hindi-English Neural Machine Translation Systems at WAT 2019 1 Nov 2019 · 0 repositories
-
Mixed Multi-Head Self-Attention for Neural Machine Translation 1 Nov 2019 · 0 repositories
-
MrMep: Joint Extraction of Multiple Relations and Multiple Entity Pairs Based on Triplet Attention 1 Nov 2019 · 0 repositories
-
Multi-View Domain Adapted Sentence Embeddings for Low-Resource Unsupervised Duplicate Question Detection 1 Nov 2019 · 0 repositories
-
Named Entity Recognition - Is There a Glass Ceiling? 1 Nov 2019 · 0 repositories
-
Natural Language Generation for Effective Knowledge Distillation 1 Nov 2019 · 1 repository
-
No, you're not alone: A better way to find people with similar experiences on Reddit 1 Nov 2019 · 0 repositories
-
NSIT@NLP4IF-2019: Propaganda Detection from News Articles using Transfer Learning 1 Nov 2019 · 0 repositories
-
On the Relation between Position Information and Sentence Length in Neural Machine Translation 1 Nov 2019 · 0 repositories
-
Our Neural Machine Translation Systems for WAT 2019 1 Nov 2019 · 0 repositories
-
Pingan Smart Health and SJTU at COIN - Shared Task: utilizing Pre-trained Language Models and Common-sense Knowledge in Machine Reading Tasks 1 Nov 2019 · 0 repositories
-
Policy Preference Detection in Parliamentary Debate Motions 1 Nov 2019 · 0 repositories
-
Pre-Training BERT on Domain Resources for Short Answer Grading 1 Nov 2019 · 0 repositories
-
Question Answering Using Hierarchical Attention on Top of BERT Features 1 Nov 2019 · 0 repositories
-
Recurrent Positional Embedding for Neural Machine Translation 1 Nov 2019 · 0 repositories
-
Recycling a Pre-trained BERT Encoder for Neural Machine Translation 1 Nov 2019 · 0 repositories
-
Reevaluating Argument Component Extraction in Low Resource Settings 1 Nov 2019 · 0 repositories
-
Relation Extraction among Multiple Entities Using a Dual Pointer Network with a Multi-Head Attention Mechanism 1 Nov 2019 · 0 repositories
-
Relation Module for Non-Answerable Predictions on Reading Comprehension 1 Nov 2019 · 0 repositories
-
Sarah's Participation in WAT 2019 1 Nov 2019 · 0 repositories
-
Selecting, Planning, and Rewriting: A Modular Approach for Data-to-Document Generation and Translation 1 Nov 2019 · 0 repositories
-
Self-Adaptive Scaling for Learnable Residual Structure 1 Nov 2019 · 0 repositories
-
Sentence-Level Propaganda Detection in News Articles with Transfer Learning and BERT-BiLSTM-Capsule Model 1 Nov 2019 · 0 repositories
-
SUDA-Alibaba at MRP 2019: Graph-Based Models with BERT 1 Nov 2019 · 0 repositories
-
SUM-QE: a BERT-based Summary Quality Estimation Model 1 Nov 2019 · 0 repositories
-
Supervised neural machine translation based on data augmentation and improved training & inference process 1 Nov 2019 · 0 repositories
-
SYSTRAN @ WAT 2019: Russian-Japanese News Commentary task 1 Nov 2019 · 0 repositories
-
SYSTRAN @ WNGT 2019: DGT Task 1 Nov 2019 · 0 repositories
-
Team DOMLIN: Exploiting Evidence Enhancement for the FEVER Shared Task 1 Nov 2019 · 0 repositories
-
The Concordia NLG Surface Realizer at SRST 2019 1 Nov 2019 · 0 repositories
-
Transfer Learning in Biomedical Named Entity Recognition: An Evaluation of BERT in the PharmaCoNER task 1 Nov 2019 · 0 repositories
-
Transformer and seq2seq model for Paraphrase Generation 1 Nov 2019 · 0 repositories
-
Transformer-based Model for Single Documents Neural Summarization 1 Nov 2019 · 0 repositories
-
Transformer Dissection: An Unified Understanding for Transformer's Attention via the Lens of Kernel 1 Nov 2019 · 0 repositories
-
``Transforming'' Delete, Retrieve, Generate Approach for Controlled Text Style Transfer 1 Nov 2019 · 0 repositories
-
TUPA at MRP 2019: A Multi-Task Baseline System 1 Nov 2019 · 0 repositories