Methods › General › Regularization › Attention Dropout › Papers, page 88
Attention Dropout
Papers archive 2025-07-28
archive papers tagged: 10,892 · with a code link: 4,634 · where Syntology ran a sample: 1,270 (1,043 with a run with no instrument failure, 227 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (1,270 of 10,892 tagged: 1,043 with a run with no instrument failure, 227 where every run was a failure of Syntology's instrument)
Page 88 of 109: papers 8,701 to 8,800 of 10,892, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
SKILLBERT: “SKILLING” THE BERT TO CLASSIFY SKILLS! 8 Mar 2021 · 0 repositories
-
Syntax-BERT: Improving Pre-trained Transformers with Syntax Trees 7 Mar 2021 · 1 repository · arXiv:2103.04350
-
Fine-tuning Pretrained Multilingual BERT Model for Indonesian Aspect-based Sentiment Analysis 5 Mar 2021 · 0 repositories · arXiv:2103.03732
-
MalBERT: Using Transformers for Cybersecurity and Malicious Software Detection 5 Mar 2021 · 0 repositories · arXiv:2103.03806
-
Non-invasive Self-attention for Side Information Fusion in Sequential Recommendation 5 Mar 2021 · 0 repositories · arXiv:2103.03578
-
Hardware Acceleration of Fully Quantized BERT for Efficient Natural Language Processing 4 Mar 2021 · 0 repositories · arXiv:2103.02800
-
Few-shot Learning for Slot Tagging with Attentive Relational Network 3 Mar 2021 · 0 repositories · arXiv:2103.02333
-
Natural Language Understanding for Argumentative Dialogue Systems in the Opinion Building Domain 3 Mar 2021 · 0 repositories · arXiv:2103.02691
-
Disentangling Syntax and Semantics in the Brain with Deep Networks 2 Mar 2021 · 0 repositories · arXiv:2103.01620
-
Hate Towards the Political Opponent: A Twitter Corpus Study of the 2020 US Elections on the Basis of Offensive Speech and Stance Detection 2 Mar 2021 · 0 repositories · arXiv:2103.01664
-
BERT-based knowledge extraction method of unstructured domain text 1 Mar 2021 · 0 repositories · arXiv:2103.00728
-
BERT based patent novelty search by training claims to their own description 1 Mar 2021 · 0 repositories · arXiv:2103.01126
-
Combat COVID-19 Infodemic Using Explainable Natural Language Processing Models 1 Mar 2021 · 0 repositories · arXiv:2103.00747
-
Long Document Summarization in a Low Resource Setting using Pretrained Language Models 1 Mar 2021 · 0 repositories · arXiv:2103.00751
-
NLP-CUET@DravidianLangTech-EACL2021: Offensive Language Detection from Multilingual Code-Mixed Text using Transformers 28 Feb 2021 · 1 repository · arXiv:2103.00455
-
NLP-CUET@LT-EDI-EACL2021: Multilingual Code-Mixed Hope Speech Detection using Cross-lingual Representation Learner 28 Feb 2021 · 1 repository · arXiv:2103.00464
-
COVID-19 Tweets Analysis through Transformer Language Models 27 Feb 2021 · 1 repository · arXiv:2103.00199
-
Transformer in Transformer 27 Feb 2021 · 12 repositories · arXiv:2103.00112Syntology official (archive's flag): 4 ran · 16 ran (of which 12 constructed an object rather than computing a result; 15 with no instrument failure: 1 honoured, 1 violated, 13 with no contract checked; 1 where Syntology's instrument failed) · 8 unverified (of 24 harvested samples) · 5 pointer-only (licence)
-
Transformers with Competitive Ensembles of Independent Mechanisms 27 Feb 2021 · 0 repositories · arXiv:2103.00336
-
Multi-task transfer learning for finding actionable information from crisis-related messages on social media 26 Feb 2021 · 0 repositories · arXiv:2102.13395
-
BERT-based Acronym Disambiguation with Multiple Training Strategies 25 Feb 2021 · 0 repositories · arXiv:2103.00488
-
Emotion-Aware, Emotion-Agnostic, or Automatic: Corpus Creation Strategies to Obtain Cognitive Event Appraisal Annotations 25 Feb 2021 · 0 repositories · arXiv:2102.12858
-
PharmKE: Knowledge Extraction Platform for Pharmaceutical Texts using Transfer Learning 25 Feb 2021 · 0 repositories · arXiv:2102.13139
-
Sentiment Analysis of Persian-English Code-mixed Texts 25 Feb 2021 · 1 repository · arXiv:2102.12700
-
From Universal Language Model to Downstream Task: Improving RoBERTa-Based Vietnamese Hate Speech Detection 24 Feb 2021 · 0 repositories · arXiv:2102.12162
-
Hopeful_Men@LT-EDI-EACL2021: Hope Speech Detection Using Indic Transliteration and Transformers 24 Feb 2021 · 0 repositories · arXiv:2102.12082
-
LRG at SemEval-2021 Task 4: Improving Reading Comprehension with Abstract Words using Augmentation, Linguistic Features and Voting 24 Feb 2021 · 1 repository · arXiv:2102.12255
-
NLRG at SemEval-2021 Task 5: Toxic Spans Detection Leveraging BERT-based Token Classification and Span Prediction Techniques 24 Feb 2021 · 1 repository · arXiv:2102.12254
-
PADA: Example-based Prompt Learning for on-the-fly Adaptation to Unseen Domains 24 Feb 2021 · 1 repository · arXiv:2102.12206
-
Task-Specific Pre-Training and Cross Lingual Transfer for Code-Switched Data 24 Feb 2021 · 0 repositories · arXiv:2102.12407
-
Minimally-Supervised Structure-Rich Text Categorization via Learning on Text-Rich Networks 23 Feb 2021 · 0 repositories · arXiv:2102.11479
-
Robust and Transferable Anomaly Detection in Log Data using Pre-Trained Language Models 23 Feb 2021 · 0 repositories · arXiv:2102.11570
-
VisualCheXbert: Addressing the Discrepancy Between Radiology Report Labels and Image Labels 23 Feb 2021 · 1 repository · arXiv:2102.11467Syntology official: harvested, nothing ran · 0 ran · 2 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
Conditional Positional Encodings for Vision Transformers 22 Feb 2021 · 2 repositories · arXiv:2102.10882
-
Evaluating Contextualized Language Models for Hungarian 22 Feb 2021 · 1 repository · arXiv:2102.10848
-
Generating Human Readable Transcript for Automatic Speech Recognition with Pre-trained Language Model 22 Feb 2021 · 0 repositories · arXiv:2102.11114
-
MixUp Training Leads to Reduced Overfitting and Improved Calibration for the Transformer Architecture 22 Feb 2021 · 0 repositories · arXiv:2102.11402
-
Parallelizing Legendre Memory Unit Training 22 Feb 2021 · 2 repositories · arXiv:2102.11417
-
RUBERT: A Bilingual Roman Urdu BERT Using Cross Lingual Transfer Learning 22 Feb 2021 · 0 repositories · arXiv:2102.11278
-
Using Prior Knowledge to Guide BERT's Attention in Semantic Textual Matching Tasks 22 Feb 2021 · 1 repository · arXiv:2102.10934
-
Pre-Training BERT on Arabic Tweets: Practical Considerations 21 Feb 2021 · 0 repositories · arXiv:2102.10684
-
Web-based Application for Detecting Indonesian Clickbait Headlines using IndoBERT 21 Feb 2021 · 0 repositories · arXiv:2102.10601
-
Calibrate Before Use: Improving Few-Shot Performance of Language Models 19 Feb 2021 · 5 repositories · arXiv:2102.09690Syntology official: harvested, nothing ran · 0 ran · 4 unverified (of 4 harvested samples)
-
Learning Dynamic BERT via Trainable Gate Variables and a Bi-modal Regularizer 19 Feb 2021 · 0 repositories · arXiv:2102.09727
-
Towards Emotion Recognition in Hindi-English Code-Mixed Data: A Transformer Based Approach 19 Feb 2021 · 1 repository · arXiv:2102.09943
-
Using Transformer based Ensemble Learning to classify Scientific Articles 19 Feb 2021 · 2 repositories · arXiv:2102.09991
-
Analysis Of Contextual and Non-Contextual Word Embedding Models For Hindi NER With Web Application For Data Collection 18 Feb 2021 · 1 repository
-
Quiz-Style Question Generation for News Stories 18 Feb 2021 · 2 repositories · arXiv:2102.09094
-
Training Large-Scale News Recommenders with Pretrained Language Models in the Loop 18 Feb 2021 · 1 repository · arXiv:2102.09268
-
UnibucKernel: Geolocating Swiss German Jodels Using Ensemble Learning 18 Feb 2021 · 0 repositories · arXiv:2102.09379
-
Leveraging Query Resolution and Reading Comprehension for Conversational Passage Retrieval 17 Feb 2021 · 0 repositories · arXiv:2102.08795
-
SciDr at SDU-2020: IDEAS -- Identifying and Disambiguating Everyday Acronyms for Scientific Domain 17 Feb 2021 · 2 repositories · arXiv:2102.08818
-
TCN: Table Convolutional Network for Web Table Interpretation 17 Feb 2021 · 1 repository · arXiv:2102.09460
-
THEaiTRE 1.0: Interactive generation of theatre play scripts 17 Feb 2021 · 0 repositories · arXiv:2102.08892
-
COCO-LM: Correcting and Contrasting Text Sequences for Language Model Pretraining 16 Feb 2021 · 2 repositories · arXiv:2102.08473Syntology community repositories only · 5 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 4 where Syntology's instrument failed) · 1 unverified (of 6 harvested samples)
-
Exploring Transformers in Natural Language Generation: GPT, BERT, and XLNet 16 Feb 2021 · 1 repository · arXiv:2102.08036
-
Have Attention Heads in BERT Learned Constituency Grammar? 16 Feb 2021 · 0 repositories · arXiv:2102.07926
-
Non-Autoregressive Text Generation with Pre-trained Language Models 16 Feb 2021 · 1 repository · arXiv:2102.08220
-
TeraPipe: Token-Level Pipeline Parallelism for Training Large-Scale Language Models 16 Feb 2021 · 1 repository · arXiv:2102.07988Syntology official (archive's flag): 7 ran · 7 ran (of which 3 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 4 where Syntology's instrument failed) · 1 unverified (of 8 harvested samples) · 8 pointer-only (licence)
-
DOBF: A Deobfuscation Pre-Training Objective for Programming Languages 15 Feb 2021 · 2 repositories · arXiv:2102.07492
-
Fast End-to-End Speech Recognition via Non-Autoregressive Models and Cross-Modal Knowledge Transferring from BERT 15 Feb 2021 · 0 repositories · arXiv:2102.07594
-
Improved Customer Transaction Classification using Semi-Supervised Knowledge Distillation 15 Feb 2021 · 0 repositories · arXiv:2102.07635
-
Prompt Programming for Large Language Models: Beyond the Few-Shot Paradigm 15 Feb 2021 · 0 repositories · arXiv:2102.07350
-
The corruptive force of AI-generated advice 15 Feb 2021 · 0 repositories · arXiv:2102.07536
-
Within-Document Event Coreference with BERT-Based Contextualized Representations 15 Feb 2021 · 0 repositories · arXiv:2102.09600
-
indicnlp@kgp at DravidianLangTech-EACL2021: Offensive Language Identification in Dravidian Languages 14 Feb 2021 · 1 repository · arXiv:2102.07150
-
indicnlp@ kgp at DravidianLangTech-EACL2021: Offensive Language Identification in Dravidian Languages 14 Feb 2021 · 1 repository
-
Characterizing English Variation across Social Media Communities with BERT 12 Feb 2021 · 1 repository · arXiv:2102.06820
-
Dancing along Battery: Enabling Transformer with Run-time Reconfigurability on Mobile Devices 12 Feb 2021 · 0 repositories · arXiv:2102.06336
-
Dynamic Precision Analog Computing for Neural Networks 12 Feb 2021 · 1 repository · arXiv:2102.06365
-
Exploring Classic and Neural Lexical Translation Models for Information Retrieval: Interpretability, Effectiveness, and Efficiency Benefits 12 Feb 2021 · 2 repositories · arXiv:2102.06815
-
Multiversal views on language models 12 Feb 2021 · 0 repositories · arXiv:2102.06391
-
Optimizing Inference Performance of Transformers on CPUs 12 Feb 2021 · 0 repositories · arXiv:2102.06621
-
AuGPT: Auxiliary Tasks and Data Augmentation for End-To-End Dialogue with Pre-Trained Language Models 9 Feb 2021 · 1 repository · arXiv:2102.05126
-
NewsBERT: Distilling Pre-trained Language Model for Intelligent News Application 9 Feb 2021 · 0 repositories · arXiv:2102.04887
-
Transfer Learning Approach for Arabic Offensive Language Detection System -- BERT-Based Model 9 Feb 2021 · 0 repositories · arXiv:2102.05708
-
A Hybrid Task-Oriented Dialog System with Domain and Task Adaptive Pretraining 8 Feb 2021 · 0 repositories · arXiv:2102.04506
-
Generating Fake Cyber Threat Intelligence Using Transformer-Based Models 8 Feb 2021 · 0 repositories · arXiv:2102.04351
-
Bias Out-of-the-Box: An Empirical Analysis of Intersectional Occupational Biases in Popular Generative Language Models 8 Feb 2021 · 1 repository · arXiv:2102.04130Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 5 where Syntology's instrument failed) · 0 unverified (of 5 harvested samples)
-
Spoiler Alert: Using Natural Language Processing to Detect Spoilers in Book Reviews 7 Feb 2021 · 1 repository · arXiv:2102.03882
-
Jointly Improving Language Understanding and Generation with Quality-Weighted Weak Supervision of Automatic Labeling 6 Feb 2021 · 0 repositories · arXiv:2102.03551
-
Neural Data-to-Text Generation with LM-based Text Augmentation 6 Feb 2021 · 0 repositories · arXiv:2102.03556
-
PipeTransformer: Automated Elastic Pipelining for Distributed Training of Transformers 5 Feb 2021 · 1 repository · arXiv:2102.03161Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
RpBERT: A Text-image Relation Propagation-based BERT Model for Multimodal NER 5 Feb 2021 · 1 repository · arXiv:2102.02967
-
Understanding Emails and Drafting Responses -- An Approach Using GPT-3 5 Feb 2021 · 0 repositories · arXiv:2102.03062
-
1-bit Adam: Communication Efficient Large-Scale Training with Adam's Convergence Speed 4 Feb 2021 · 2 repositories · arXiv:2102.02888
-
Hierarchical Multi-head Attentive Network for Evidence-aware Fake News Detection 4 Feb 2021 · 1 repository · arXiv:2102.02680
-
Understanding the Capabilities, Limitations, and Societal Impact of Large Language Models 4 Feb 2021 · 0 repositories · arXiv:2102.02503
-
Bootstrapping Multilingual AMR with Contextual Word Alignments 3 Feb 2021 · 0 repositories · arXiv:2102.02189
-
HeBERT & HebEMO: a Hebrew BERT Model and a Tool for Polarity Analysis and Emotion Recognition 3 Feb 2021 · 0 repositories · arXiv:2102.01909
-
Introduction to Neural Transfer Learning with Transformers for Social Science Text Analysis 3 Feb 2021 · 0 repositories · arXiv:2102.02111
-
AutoFreeze: Automatically Freezing Model Blocks to Accelerate Fine-tuning 2 Feb 2021 · 1 repository · arXiv:2102.01386Syntology official: no sample here; runs from other or unrecorded repositories · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Clickbait Headline Detection in Indonesian News Sites using Multilingual Bidirectional Encoder Representations from Transformers (M-BERT) 2 Feb 2021 · 0 repositories · arXiv:2102.01497
-
Improving Distantly-Supervised Relation Extraction through BERT-based Label & Instance Embeddings 1 Feb 2021 · 1 repository · arXiv:2102.01156
-
"Is depression related to cannabis?": A knowledge-infused model for Entity and Relation Extraction with Limited Supervision 1 Feb 2021 · 0 repositories · arXiv:2102.01222
-
Polyphone Disambiguation in Mandarin Chinese with Semi-Supervised Learning 1 Feb 2021 · 0 repositories · arXiv:2102.00621
-
Scaling Federated Learning for Fine-tuning of Large Language Models 1 Feb 2021 · 0 repositories · arXiv:2102.00875
-
SJ_AJ@DravidianLangTech-EACL2021: Task-Adaptive Pre-Training of Multilingual BERT models for Offensive Language Identification 1 Feb 2021 · 1 repository · arXiv:2102.01051
-
Text-to-hashtag Generation using Seq2seq Learning 1 Feb 2021 · 1 repository · arXiv:2102.00904
-
[Re] Reproducing Learning to Deceive With Attention-Based Explanations 31 Jan 2021 · 1 repository