Methods › General › Regularization › Attention Dropout › Papers, page 62
Attention Dropout
Papers archive 2025-07-28
archive papers tagged: 10,892 · with a code link: 4,634 · where Syntology ran a sample: 1,270 (1,043 with a run with no instrument failure, 227 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (1,270 of 10,892 tagged: 1,043 with a run with no instrument failure, 227 where every run was a failure of Syntology's instrument)
Page 62 of 109: papers 6,101 to 6,200 of 10,892, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Style transfer and classification in hebrew news items 6 Dec 2022 · 0 repositories · arXiv:2212.03019
-
Audio-Driven Co-Speech Gesture Video Generation 5 Dec 2022 · 0 repositories · arXiv:2212.02350
-
Automatic Generation of Factual News Headlines in Finnish 5 Dec 2022 · 0 repositories · arXiv:2212.02170
-
Video Games as a Corpus: Sentiment Analysis using Fallout New Vegas Dialog 5 Dec 2022 · 0 repositories · arXiv:2212.02168
-
Languages You Know Influence Those You Learn: Impact of Language Characteristics on Multi-Lingual Text-to-Text Transfer 4 Dec 2022 · 0 repositories · arXiv:2212.01757
-
Exploring the Limits of Differentially Private Deep Learning with Group-wise Clipping 3 Dec 2022 · 0 repositories · arXiv:2212.01539
-
Global memory transformer for processing long documents 3 Dec 2022 · 0 repositories · arXiv:2212.01650
-
ColD Fusion: Collaborative Descent for Distributed Multitask Finetuning 2 Dec 2022 · 0 repositories · arXiv:2212.01378
-
Event knowledge in large language models: the gap between the impossible and the unlikely 2 Dec 2022 · 1 repository · arXiv:2212.01488
-
SumREN: Summarizing Reported Speech about Events in News 2 Dec 2022 · 1 repository · arXiv:2212.01146
-
a survey on GPT-3 1 Dec 2022 · 0 repositories · arXiv:2212.00857
-
Adapted Multimodal BERT with Layer-wise Fusion for Sentiment Analysis 1 Dec 2022 · 0 repositories · arXiv:2212.00678
-
Data-Efficient Finetuning Using Cross-Task Nearest Neighbors 1 Dec 2022 · 1 repository · arXiv:2212.00196Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 2 harvested samples)
-
Distilling Reasoning Capabilities into Smaller Language Models 1 Dec 2022 · 1 repository · arXiv:2212.00193
-
BudgetLongformer: Can we Cheaply Pretrain a SotA Legal Language Model From Scratch? 30 Nov 2022 · 0 repositories · arXiv:2211.17135
-
ExtremeBERT: A Toolkit for Accelerating Pretraining of Customized BERT 30 Nov 2022 · 1 repository · arXiv:2211.17201Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
HEAT: Hardware-Efficient Automatic Tensor Decomposition for Transformer Compression 30 Nov 2022 · 0 repositories · arXiv:2211.16749
-
Quadapter: Adapter for GPT-2 Quantization 30 Nov 2022 · 0 repositories · arXiv:2211.16912
-
Composition based oxidation state prediction of materials using deep learning 29 Nov 2022 · 1 repository · arXiv:2211.15895
-
Diverse Multi-Answer Retrieval with Determinantal Point Processes 29 Nov 2022 · 0 repositories · arXiv:2211.16029
-
NoisyQuant: Noisy Bias-Enhanced Post-Training Activation Quantization for Vision Transformers 29 Nov 2022 · 1 repository · arXiv:2211.16056Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
Outfit Generation and Recommendation -- An Experimental Study 29 Nov 2022 · 0 repositories · arXiv:2211.16353
-
Prompted Opinion Summarization with GPT-3.5 29 Nov 2022 · 1 repository · arXiv:2211.15914
-
Automatically Extracting Information in Medical Dialogue: Expert System And Attention for Labelling 28 Nov 2022 · 0 repositories · arXiv:2211.15544
-
DiffusionBERT: Improving Generative Masked Language Models with Diffusion Models 28 Nov 2022 · 1 repository · arXiv:2211.15029Syntology official (archive's flag): 9 ran · 9 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 0 violated, 9 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 11 harvested samples) · 2 pointer-only (licence)
-
GPT-Neo for commonsense reasoning -- a theoretical and practical lens 28 Nov 2022 · 1 repository · arXiv:2211.15593
-
Handling and extracting key entities from customer conversations using Speech recognition and Named Entity recognition 28 Nov 2022 · 0 repositories · arXiv:2211.17107
-
Is it Required? Ranking the Skills Required for a Job-Title 28 Nov 2022 · 0 repositories · arXiv:2212.08553
-
Revisiting Distance Metric Learning for Few-Shot Natural Language Classification 28 Nov 2022 · 0 repositories · arXiv:2211.15202
-
Scientific and Creative Analogies in Pretrained Language Models 28 Nov 2022 · 2 repositories · arXiv:2211.15268
-
ESIE-BERT: Enriching Sub-words Information Explicitly with BERT for Joint Intent Classification and SlotFilling 27 Nov 2022 · 0 repositories · arXiv:2211.14829
-
Detect-Localize-Repair: A Unified Framework for Learning to Debug with CodeT5 27 Nov 2022 · 0 repositories · arXiv:2211.14875
-
Understanding BLOOM: An empirical study on diverse NLP tasks 27 Nov 2022 · 0 repositories · arXiv:2211.14865
-
An Analysis of Social Biases Present in BERT Variants Across Multiple Languages 25 Nov 2022 · 1 repository · arXiv:2211.14402Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
Finetuning BERT on Partially Annotated NER Corpora 25 Nov 2022 · 1 repository · arXiv:2211.14360
-
GPT-3-driven pedagogical agents for training children's curious question-asking skills 25 Nov 2022 · 0 repositories · arXiv:2211.14228
-
InDEX: Indonesian Idiom and Expression Dataset for Cloze Test 24 Nov 2022 · 0 repositories · arXiv:2211.13376
-
Using Selective Masking as a Bridge between Pre-training and Fine-tuning 24 Nov 2022 · 0 repositories · arXiv:2211.13815
-
Holistic Visual-Textual Sentiment Analysis with Prior Models 23 Nov 2022 · 1 repository · arXiv:2211.12981
-
SEAT: Stable and Explainable Attention 23 Nov 2022 · 0 repositories · arXiv:2211.13290
-
SPCXR: Self-supervised Pretraining using Chest X-rays Towards a Domain Specific Foundation Model 23 Nov 2022 · 0 repositories · arXiv:2211.12944
-
Word-Level Representation From Bytes For Language Modeling 23 Nov 2022 · 0 repositories · arXiv:2211.12677
-
Coreference Resolution through a seq2seq Transition-Based System 22 Nov 2022 · 1 repository · arXiv:2211.12142
-
HyperTuning: Toward Adapting Large Language Models without Back-propagation 22 Nov 2022 · 0 repositories · arXiv:2211.12485
-
OLGA : An Ontology and LSTM-based approach for generating Arithmetic Word Problems (AWPs) of transfer type 22 Nov 2022 · 0 repositories · arXiv:2211.12164
-
PromptTTS: Controllable Text-to-Speech with Text Descriptions 22 Nov 2022 · 1 repository · arXiv:2211.12171
-
AF Adapter: Continual Pretraining for Building Chinese Biomedical Language Model 21 Nov 2022 · 1 repository · arXiv:2211.11363
-
Enhancing Self-Consistency and Performance of Pre-Trained Language Models through Natural Language Inference 21 Nov 2022 · 0 repositories · arXiv:2211.11875
-
Exploring the Efficacy of Pre-trained Checkpoints in Text-to-Music Generation Task 21 Nov 2022 · 2 repositories · arXiv:2211.11216
-
L3Cube-HindBERT and DevBERT: Pre-Trained BERT Transformer models for Devanagari based Hindi and Marathi Languages 21 Nov 2022 · 0 repositories · arXiv:2211.11418
-
L3Cube-MahaSBERT and HindSBERT: Sentence BERT Models and Benchmarking BERT Sentence Representations for Hindi and Marathi 21 Nov 2022 · 1 repository · arXiv:2211.11187
-
Language in a Bottle: Language Model Guided Concept Bottlenecks for Interpretable Image Classification 21 Nov 2022 · 2 repositories · arXiv:2211.11158
-
PointCLIP V2: Prompting CLIP and GPT for Powerful 3D Open-world Learning 21 Nov 2022 · 2 repositories · arXiv:2211.11682Syntology official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 1 violated, 3 with no contract checked; 4 where Syntology's instrument failed) · 4 unverified (of 12 harvested samples) · 4 pointer-only (licence)
-
TCBERT: A Technical Report for Chinese Topic Classification BERT 21 Nov 2022 · 0 repositories · arXiv:2211.11304
-
Conceptor-Aided Debiasing of Large Language Models 20 Nov 2022 · 0 repositories · arXiv:2211.11087
-
Detecting Conspiracy Theory Against COVID-19 Vaccines 20 Nov 2022 · 0 repositories · arXiv:2211.13003
-
Feature Weaken: Vicinal Data Augmentation for Classification 20 Nov 2022 · 0 repositories · arXiv:2211.10944
-
Understanding and Improving Knowledge Distillation for Quantization-Aware Training of Large Transformer Encoders 20 Nov 2022 · 1 repository · arXiv:2211.11014
-
UnifiedABSA: A Unified ABSA Framework Based on Multi-task Instruction Tuning 20 Nov 2022 · 0 repositories · arXiv:2211.10986
-
A survey on knowledge-enhanced multimodal learning 19 Nov 2022 · 0 repositories · arXiv:2211.12328
-
Entity-Assisted Language Models for Identifying Check-worthy Sentences 19 Nov 2022 · 0 repositories · arXiv:2211.10678
-
Leveraging Users' Social Network Embeddings for Fake News Detection on Twitter 19 Nov 2022 · 0 repositories · arXiv:2211.10672
-
Metadata Might Make Language Models Better 18 Nov 2022 · 0 repositories · arXiv:2211.10086
-
Where did you tweet from? Inferring the origin locations of tweets based on contextual information 18 Nov 2022 · 0 repositories · arXiv:2211.16506
-
EfficientTrain: Exploring Generalized Curriculum Learning for Training Visual Backbones 17 Nov 2022 · 1 repository · arXiv:2211.09703
-
GLAMI-1M: A Multilingual Image-Text Fashion Dataset 17 Nov 2022 · 1 repository · arXiv:2211.14451
-
Ignore Previous Prompt: Attack Techniques For Language Models 17 Nov 2022 · 1 repository · arXiv:2211.09527Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified (of 6 harvested samples)
-
LongFNT: Long-form Speech Recognition with Factorized Neural Transducer 17 Nov 2022 · 0 repositories · arXiv:2211.09412
-
ProtSi: Prototypical Siamese Network with Data Augmentation for Few-Shot Subjective Answer Evaluation 17 Nov 2022 · 1 repository · arXiv:2211.09855
-
Random-LTD: Random and Layerwise Token Dropping Brings Efficient Training for Large-scale Transformers 17 Nov 2022 · 1 repository · arXiv:2211.11586
-
UniSumm and SummZoo: Unified Model and Diverse Benchmark for Few-Shot Summarization 17 Nov 2022 · 1 repository · arXiv:2211.09783
-
Fast and Accurate FSA System Using ELBERT: An Efficient and Lightweight BERT 16 Nov 2022 · 0 repositories · arXiv:2211.08842
-
Galactica: A Large Language Model for Science 16 Nov 2022 · 1 repository · arXiv:2211.09085Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
TSMind: Alibaba and Soochow University's Submission to the WMT22 Translation Suggestion Task 16 Nov 2022 · 0 repositories · arXiv:2211.08987
-
Unified Question Answering in Slovene 16 Nov 2022 · 1 repository · arXiv:2211.09159
-
ALIGN-MLM: Word Embedding Alignment is Crucial for Multilingual Pre-training 15 Nov 2022 · 1 repository · arXiv:2211.08547
-
An FNet based Auto Encoder for Long Sequence News Story Generation 15 Nov 2022 · 1 repository · arXiv:2211.08295
-
Breakpoint Transformers for Modeling and Tracking Intermediate Beliefs 15 Nov 2022 · 1 repository · arXiv:2211.07950Syntology official: no sample here; runs from other or unrecorded repositories · 3 ran (of which 2 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 1 where Syntology's instrument failed) · 3 unverified (of 6 harvested samples) · 1 pointer-only (licence)
-
Empowering Language Models with Knowledge Graph Reasoning for Question Answering 15 Nov 2022 · 0 repositories · arXiv:2211.08380
-
GLUE-X: Evaluating Natural Language Understanding Models from an Out-of-distribution Generalization Perspective 15 Nov 2022 · 1 repository · arXiv:2211.08073
-
PromptCap: Prompt-Guided Task-Aware Image Captioning 15 Nov 2022 · 1 repository · arXiv:2211.09699
-
RobBERT-2022: Updating a Dutch Language Model to Account for Evolving Language Use 15 Nov 2022 · 0 repositories · arXiv:2211.08192
-
Are Hard Examples also Harder to Explain? A Study with Human and Model-Generated Explanations 14 Nov 2022 · 1 repository · arXiv:2211.07517Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
CST5: Data Augmentation for Code-Switched Semantic Parsing 14 Nov 2022 · 1 repository · arXiv:2211.07514
-
Technological taxonomies for hypernym and hyponym retrieval in patent texts 14 Nov 2022 · 1 repository · arXiv:2212.06039
-
UGIF: UI Grounded Instruction Following 14 Nov 2022 · 0 repositories · arXiv:2211.07615
-
GreenPLM: Cross-Lingual Transfer of Monolingual Pre-Trained Language Models at Almost No Cost 13 Nov 2022 · 1 repository · arXiv:2211.06993
-
Textual Data Augmentation for Patient Outcomes Prediction 13 Nov 2022 · 0 repositories · arXiv:2211.06778
-
Large Language Models Meet Harry Potter: A Bilingual Dataset for Aligning Dialogue Agents with Characters 13 Nov 2022 · 1 repository · arXiv:2211.06869
-
Xu at SemEval-2022 Task 4: Pre-BERT Neural Network Methods vs Post-BERT RoBERTa Approach for Patronizing and Condescending Language Detection 13 Nov 2022 · 1 repository · arXiv:2211.06874
-
Dark patterns in e-commerce: a dataset and its baseline evaluations 12 Nov 2022 · 1 repository · arXiv:2211.06543
-
DocuT5: Seq2seq SQL Generation with Table Documentation 11 Nov 2022 · 0 repositories · arXiv:2211.06193
-
Using Persuasive Writing Strategies to Explain and Detect Health Misinformation 11 Nov 2022 · 1 repository · arXiv:2211.05985
-
Assistive Completion of Agrammatic Aphasic Sentences: A Transfer Learning Approach using Neurolinguistics-based Synthetic Dataset 10 Nov 2022 · 0 repositories · arXiv:2211.05557
-
BERT-Based Combination of Convolutional and Recurrent Neural Network for Indonesian Sentiment Analysis 10 Nov 2022 · 0 repositories · arXiv:2211.05273
-
BERT in Plutarch's Shadows 10 Nov 2022 · 0 repositories · arXiv:2211.05673
-
Biomedical Multi-hop Question Answering Using Knowledge Graph Embeddings and Language Models 10 Nov 2022 · 0 repositories · arXiv:2211.05351
-
PAD-Net: An Efficient Framework for Dynamic Networks 10 Nov 2022 · 1 repository · arXiv:2211.05528
-
On Optimizing the Communication of Model Parallelism 10 Nov 2022 · 0 repositories · arXiv:2211.05322
-
Syntax-Guided Domain Adaptation for Aspect-based Sentiment Analysis 10 Nov 2022 · 0 repositories · arXiv:2211.05457