Methods › General › Regularization › Attention Dropout › Papers, page 55
Attention Dropout
Papers archive 2025-07-28
archive papers tagged: 10,892 · with a code link: 4,634 · where Syntology ran a sample: 1,270 (1,043 with a run with no instrument failure, 227 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (1,270 of 10,892 tagged: 1,043 with a run with no instrument failure, 227 where every run was a failure of Syntology's instrument)
Page 55 of 109: papers 5,401 to 5,500 of 10,892, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Bridging History with AI A Comparative Evaluation of GPT 3.5, GPT4, and GoogleBARD in Predictive Accuracy and Fact Checking 13 May 2023 · 0 repositories · arXiv:2305.07868
-
GPT-Sentinel: Distinguishing Human and ChatGPT Generated Content 13 May 2023 · 2 repositories · arXiv:2305.07969
-
The Machine Psychology of Cooperation: Can GPT models operationalise prompts for altruism, cooperation, competitiveness and selfishness in economic games? 13 May 2023 · 2 repositories · arXiv:2305.07970Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples)
-
PESTS: Persian_English Cross Lingual Corpus for Semantic Textual Similarity 13 May 2023 · 0 repositories · arXiv:2305.07893
-
Learning to Reason over Scene Graphs: A Case Study of Finetuning GPT-2 into a Robot Language Model for Grounded Task Planning 12 May 2023 · 0 repositories · arXiv:2305.07716
-
NL2TL: Transforming Natural Languages to Temporal Logics using Large Language Models 12 May 2023 · 3 repositories · arXiv:2305.07766Syntology official (archive's flag): 3 ran · 5 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 2 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 4 unverified (of 9 harvested samples) · 8 pointer-only (licence)
-
TinyStories: How Small Can Language Models Be and Still Speak Coherent English? 12 May 2023 · 8 repositories · arXiv:2305.07759Syntology 10 ran (of which 0 constructed an object rather than computing a result; 8 with no instrument failure: 0 honoured, 0 violated, 8 with no contract checked; 2 where Syntology's instrument failed) · 8 unverified (of 18 harvested samples)
-
When Giant Language Brains Just Aren't Enough! Domain Pizzazz with Knowledge Sparkle Dust 12 May 2023 · 0 repositories · arXiv:2305.07230
-
A General-Purpose Multilingual Document Encoder 11 May 2023 · 1 repository · arXiv:2305.07016
-
Generative Pre-trained Transformer: A Comprehensive Review on Enabling Technologies, Potential Applications, Emerging Challenges, and Future Directions 11 May 2023 · 0 repositories · arXiv:2305.10435
-
Spear Phishing With Large Language Models 11 May 2023 · 0 repositories · arXiv:2305.06972
-
Overinformative Question Answering by Humans and Machines 11 May 2023 · 0 repositories · arXiv:2305.07151
-
Recommendation as Instruction Following: A Large Language Model Empowered Recommendation Approach 11 May 2023 · 0 repositories · arXiv:2305.07001
-
Transformers for CT Reconstruction From Monoplanar and Biplanar Radiographs 11 May 2023 · 0 repositories · arXiv:2305.06965
-
A Method to Automate the Discharge Summary Hospital Course for Neurology Patients 10 May 2023 · 0 repositories · arXiv:2305.06416
-
Bits of Grass: Does GPT already know how to write like Whitman? 10 May 2023 · 0 repositories · arXiv:2305.11064
-
Davinci the Dualist: the mind-body divide in large language models and in human learners 10 May 2023 · 0 repositories · arXiv:2305.07667
-
Enriching language models with graph-based context information to better understand textual data 10 May 2023 · 1 repository · arXiv:2305.11070
-
Generating medically-accurate summaries of patient-provider dialogue: A multi-stage approach using large language models 10 May 2023 · 0 repositories · arXiv:2305.05982
-
Benchmarking large language models for biomedical natural language processing applications and recommendations 10 May 2023 · 1 repository · arXiv:2305.16326
-
A Black-Box Attack on Code Models via Representation Nearest Neighbor Search 10 May 2023 · 0 repositories · arXiv:2305.05896
-
Summarizing, Simplifying, and Synthesizing Medical Evidence Using GPT-3 (with Varying Success) 10 May 2023 · 1 repository · arXiv:2305.06299
-
A Review of Vision-Language Models and their Performance on the Hateful Memes Challenge 9 May 2023 · 1 repository · arXiv:2305.06159
-
Alleviating Over-smoothing for Unsupervised Sentence Representation 9 May 2023 · 1 repository · arXiv:2305.06154Syntology official: harvested, nothing ran · 0 ran · 4 unverified (of 4 harvested samples)
-
An Exploration of Encoder-Decoder Approaches to Multi-Label Classification for Legal and Biomedical Text 9 May 2023 · 1 repository · arXiv:2305.05627
-
Attack Named Entity Recognition by Entity Boundary Interference 9 May 2023 · 0 repositories · arXiv:2305.05253
-
CodeIE: Large Code Generation Models are Better Few-Shot Information Extractors 9 May 2023 · 1 repository · arXiv:2305.05711Syntology official: harvested, nothing ran · 0 ran · 1 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Detection of depression on social networks using transformers and ensembles 9 May 2023 · 1 repository · arXiv:2305.05325
-
GPT in Game Theory Experiments 9 May 2023 · 0 repositories · arXiv:2305.05516
-
GPT-NAS: Evolutionary Neural Architecture Search with the Generative Pre-Trained Model 9 May 2023 · 0 repositories · arXiv:2305.05351
-
Effects of sub-word segmentation on performance of transformer language models 9 May 2023 · 0 repositories · arXiv:2305.05480
-
StrAE: Autoencoding for Pre-Trained Embeddings using Explicit Structure 9 May 2023 · 0 repositories · arXiv:2305.05588
-
Towards an Automatic Optimisation Model Generator Assisted with Generative Pre-trained Transformer 9 May 2023 · 0 repositories · arXiv:2305.05811
-
Coherent Wave Dynamics and Language Generation of a Generative Pre-trained Transformer 8 May 2023 · 0 repositories · arXiv:2305.05061
-
Do Large Language Models Show Decision Heuristics Similar to Humans? A Case Study Using GPT-3.5 8 May 2023 · 0 repositories · arXiv:2305.04400
-
Explanation-based Finetuning Makes Models More Robust to Spurious Cues 8 May 2023 · 1 repository · arXiv:2305.04990
-
GersteinLab at MEDIQA-Chat 2023: Clinical Note Summarization from Doctor-Patient Conversations through Fine-tuning and In-context Learning 8 May 2023 · 0 repositories · arXiv:2305.05001
-
Multi-Task End-to-End Training Improves Conversational Recommendation 8 May 2023 · 0 repositories · arXiv:2305.06218
-
NeuroComparatives: Neuro-Symbolic Distillation of Comparative Knowledge 8 May 2023 · 1 repository · arXiv:2305.04978
-
PreCog: Exploring the Relation between Memorization and Performance in Pre-trained Language Models 8 May 2023 · 0 repositories · arXiv:2305.04673
-
Revisiting Relation Extraction in the era of Large Language Models 8 May 2023 · 0 repositories · arXiv:2305.05003
-
Unlocking Practical Applications in Legal Domain: Evaluation of GPT for Zero-Shot Semantic Annotation of Legal Texts 8 May 2023 · 0 repositories · arXiv:2305.04417
-
Vulnerability Detection Using Two-Stage Deep Learning Models 8 May 2023 · 0 repositories · arXiv:2305.09673
-
Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting 7 May 2023 · 2 repositories · arXiv:2305.04388Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
Professional Certification Benchmark Dataset: The First 500 Jobs For Large Language Models 7 May 2023 · 0 repositories · arXiv:2305.05377
-
Stanford MLab at SemEval-2023 Task 10: Exploring GloVe- and Transformer-Based Methods for the Explainable Detection of Online Sexism 7 May 2023 · 0 repositories · arXiv:2305.04356
-
Artificial Neuropsychology: Are Large Language Models Developing Executive Functions? 6 May 2023 · 0 repositories · arXiv:2305.04134
-
On the Usage of Continual Learning for Out-of-Distribution Generalization in Pre-trained Language Models of Code 6 May 2023 · 0 repositories · arXiv:2305.04106
-
Plan-and-Solve Prompting: Improving Zero-Shot Chain-of-Thought Reasoning by Large Language Models 6 May 2023 · 3 repositories · arXiv:2305.04091Syntology official: harvested, nothing ran · 0 ran · 2 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
Pre-training Language Model as a Multi-perspective Course Learner 6 May 2023 · 0 repositories · arXiv:2305.03981
-
Refining the Responses of LLMs by Themselves 6 May 2023 · 1 repository · arXiv:2305.04039
-
Rhetorical Role Labeling of Legal Documents using Transformers and Graph Neural Networks 6 May 2023 · 0 repositories · arXiv:2305.04100
-
Adapting Transformer Language Models for Predictive Typing in Brain-Computer Interfaces 5 May 2023 · 0 repositories · arXiv:2305.03819
-
Block the Label and Noise: An N-Gram Masked Speller for Chinese Spell Checking 5 May 2023 · 0 repositories · arXiv:2305.03314
-
CLaC at SemEval-2023 Task 2: Comparing Span-Prediction and Sequence-Labeling approaches for NER 5 May 2023 · 0 repositories · arXiv:2305.03845
-
Harnessing the Power of BERT in the Turkish Clinical Domain: Pretraining Approaches for Limited Data Scenarios 5 May 2023 · 0 repositories · arXiv:2305.03788
-
Otter: A Multi-Modal Model with In-Context Instruction Tuning 5 May 2023 · 1 repository · arXiv:2305.03726
-
Predicting COVID-19 and pneumonia complications from admission texts 5 May 2023 · 0 repositories · arXiv:2305.03661
-
Simulating H.P. Lovecraft horror literature with the ChatGPT large language model 5 May 2023 · 0 repositories · arXiv:2305.03429
-
Using ChatGPT for Entity Matching 5 May 2023 · 1 repository · arXiv:2305.03423
-
Verify-and-Edit: A Knowledge-Enhanced Chain-of-Thought Framework 5 May 2023 · 1 repository · arXiv:2305.03268
-
An automatically discovered chain-of-thought prompt generalizes to novel models and datasets 4 May 2023 · 0 repositories · arXiv:2305.02897
-
AutoML-GPT: Automatic Machine Learning with GPT 4 May 2023 · 0 repositories · arXiv:2305.02499
-
Enhancing Pashto Text Classification using Language Processing Techniques for Single And Multi-Label Analysis 4 May 2023 · 0 repositories · arXiv:2305.03201
-
Gpt-4: A Review on Advancements and Opportunities in Natural Language Processing 4 May 2023 · 0 repositories · arXiv:2305.03195
-
Improving Code Example Recommendations on Informal Documentation Using BERT and Query-Aware LSH: A Comparative Study 4 May 2023 · 1 repository · arXiv:2305.03017
-
Can LLMs Capture Human Preferences? 4 May 2023 · 0 repositories · arXiv:2305.02531
-
Late-Binding Scholarship in the Age of AI: Navigating Legal and Normative Challenges of a New Form of Knowledge Production 4 May 2023 · 0 repositories · arXiv:2305.11058
-
Leveraging BERT Language Model for Arabic Long Document Classification 4 May 2023 · 0 repositories · arXiv:2305.03519
-
PersonaLLM: Investigating the Ability of Large Language Models to Express Personality Traits 4 May 2023 · 1 repository · arXiv:2305.02547Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
A Novel Plagiarism Detection Approach Combining BERT-based Word Embedding, Attention-based LSTMs and an Improved Differential Evolution Algorithm 3 May 2023 · 0 repositories · arXiv:2305.02374
-
Distilling Step-by-Step! Outperforming Larger Language Models with Less Training Data and Smaller Model Sizes 3 May 2023 · 1 repository · arXiv:2305.02301Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples)
-
Entity Tracking in Language Models 3 May 2023 · 1 repository · arXiv:2305.02363
-
evaluating bert and parsbert for analyzing persian advertisement data 3 May 2023 · 0 repositories · arXiv:2305.02426
-
Evaluating BERT-based Scientific Relation Classifiers for Scholarly Knowledge Graph Construction on Digital Library Collections 3 May 2023 · 0 repositories · arXiv:2305.02291
-
Exploring Linguistic Properties of Monolingual BERTs with Typological Classification among Languages 3 May 2023 · 0 repositories · arXiv:2305.02215
-
GPT-RE: In-context Learning for Relation Extraction using Large Language Models 3 May 2023 · 1 repository · arXiv:2305.02105
-
Improving Cancer Hallmark Classification with BERT-based Deep Learning Approach 2 May 2023 · 0 repositories · arXiv:2305.03501
-
Why So Gullible? Enhancing the Robustness of Retrieval-Augmented Models against Counterfactual Noise 2 May 2023 · 1 repository · arXiv:2305.01579
-
FreeLM: Fine-Tuning-Free Language Model 2 May 2023 · 0 repositories · arXiv:2305.01616
-
How to Unleash the Power of Large Language Models for Few-shot Relation Extraction? 2 May 2023 · 2 repositories · arXiv:2305.01555
-
A Paradigm Shift: The Future of Machine Translation Lies with Large Language Models 2 May 2023 · 0 repositories · arXiv:2305.01181
-
Unlimiformer: Long-Range Transformers with Unlimited Length Input 2 May 2023 · 2 repositories · arXiv:2305.01625Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 3 where Syntology's instrument failed) · 1 unverified (of 5 harvested samples)
-
Vision Meets Definitions: Unsupervised Visual Word Sense Disambiguation Incorporating Gloss Information 2 May 2023 · 1 repository · arXiv:2305.01788
-
Automated Paper Screening for Clinical Reviews Using Large Language Models 1 May 2023 · 0 repositories · arXiv:2305.00844
-
Logion: Machine Learning for Greek Philology 1 May 2023 · 0 repositories · arXiv:2305.01099
-
Neural Machine Translation Models with Attention-Based Dropout Layer 1 May 2023 · 1 repository
-
Retrieving Comparative Arguments using Ensemble Methods and Neural Information Retrieval 1 May 2023 · 0 repositories · arXiv:2305.01513
-
SafeWebUH at SemEval-2023 Task 11: Learning Annotator Disagreement in Derogatory Text: Comparison of Direct Training vs Aggregation 1 May 2023 · 1 repository · arXiv:2305.01050
-
Beyond Classification: Financial Reasoning in State-of-the-Art Language Models 30 Apr 2023 · 1 repository · arXiv:2305.01505
-
Using Large Language Models to Generate JUnit Tests: An Empirical Study 30 Apr 2023 · 1 repository · arXiv:2305.00418
-
How does GPT-2 compute greater-than?: Interpreting mathematical abilities in a pre-trained language model 30 Apr 2023 · 3 repositories · arXiv:2305.00586
-
Are the Best Multilingual Document Embeddings simply Based on Sentence Embeddings? 28 Apr 2023 · 1 repository · arXiv:2304.14796
-
Causal Reasoning and Large Language Models: Opening a New Frontier for Causality 28 Apr 2023 · 1 repository · arXiv:2305.00050Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
FlowTransformer: A Transformer Framework for Flow-based Network Intrusion Detection Systems 28 Apr 2023 · 1 repository · arXiv:2304.14746
-
Towards Automated Circuit Discovery for Mechanistic Interpretability 28 Apr 2023 · 4 repositories · arXiv:2304.14997Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 5 harvested samples)
-
Towards Better Domain Adaptation for Self-supervised Models: A Case Study of Child ASR 28 Apr 2023 · 1 repository · arXiv:2305.00115
-
Assessing Text Mining and Technical Analyses on Forecasting Financial Time Series 27 Apr 2023 · 0 repositories · arXiv:2304.14544
-
Framing the News:From Human Perception to Large Language Model Inferences 27 Apr 2023 · 0 repositories · arXiv:2304.14456
-
ICE-Score: Instructing Large Language Models to Evaluate Code 27 Apr 2023 · 2 repositories · arXiv:2304.14317Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 2 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 4 harvested samples)