Methods › General › Regularization › Weight Decay › Papers, page 29
Weight Decay
Papers archive 2025-07-28
archive papers tagged: 10,713 · with a code link: 4,533 · where Syntology ran a sample: 1,291 (1,064 with a run with no instrument failure, 227 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (1,291 of 10,713 tagged: 1,064 with a run with no instrument failure, 227 where every run was a failure of Syntology's instrument)
Page 29 of 108: papers 2,801 to 2,900 of 10,713, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Challenges in Deploying Long-Context Transformers: A Theoretical Peak Performance Analysis 14 May 2024 · 0 repositories · arXiv:2405.08944
-
GPT-3.5 for Grammatical Error Correction 14 May 2024 · 0 repositories · arXiv:2405.08469
-
Refinement of an Epilepsy Dictionary through Human Annotation of Health-related posts on Instagram 14 May 2024 · 0 repositories · arXiv:2405.08784
-
Towards Principled Evaluations of Sparse Autoencoders for Interpretability and Control 14 May 2024 · 0 repositories · arXiv:2405.08366
-
When Large Language Models Meet Optical Networks: Paving the Way for Automation 14 May 2024 · 0 repositories · arXiv:2405.17441
-
Can Language Models Explain Their Own Classification Behavior? 13 May 2024 · 1 repository · arXiv:2405.07436
-
Coding historical causes of death data with Large Language Models 13 May 2024 · 1 repository · arXiv:2405.07560
-
Control Token with Dense Passage Retrieval 13 May 2024 · 0 repositories · arXiv:2405.13008
-
Evaluation of Retrieval-Augmented Generation: A Survey 13 May 2024 · 1 repository · arXiv:2405.07437
-
FreeVA: Offline MLLM as Training-Free Video Assistant 13 May 2024 · 1 repository · arXiv:2405.07798Syntology official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 4 where Syntology's instrument failed) · 1 unverified (of 9 harvested samples) · 1 pointer-only (licence)
-
From Questions to Insightful Answers: Building an Informed Chatbot for University Resources 13 May 2024 · 0 repositories · arXiv:2405.08120
-
MacBehaviour: An R package for behavioural experimentation on large language models 13 May 2024 · 1 repository · arXiv:2405.07495
-
Many-Shot Regurgitation (MSR) Prompting 13 May 2024 · 0 repositories · arXiv:2405.08134
-
Open-vocabulary Auditory Neural Decoding Using fMRI-prompted LLM 13 May 2024 · 0 repositories · arXiv:2405.07840
-
DuetRAG: Collaborative Retrieval-Augmented Generation 12 May 2024 · 0 repositories · arXiv:2405.13002
-
ExplainableDetector: Exploring Transformer-based Language Modeling Approach for SMS Spam Detection with Explainability Analysis 12 May 2024 · 0 repositories · arXiv:2405.08026
-
L(u)PIN: LLM-based Political Ideology Nowcasting 12 May 2024 · 0 repositories · arXiv:2405.07320
-
Learning Reward for Robot Skills Using Large Language Models via Self-Alignment 12 May 2024 · 0 repositories · arXiv:2405.07162
-
Limited Ability of LLMs to Simulate Human Psychological Behaviours: a Psychometric Analysis 12 May 2024 · 1 repository · arXiv:2405.07248Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
Quite Good, but Not Enough: Nationality Bias in Large Language Models -- A Case Study of ChatGPT 11 May 2024 · 1 repository · arXiv:2405.06996
-
RETTA: Retrieval-Enhanced Test-Time Adaptation for Zero-Shot Video Captioning 11 May 2024 · 0 repositories · arXiv:2405.07046
-
TacoERE: Cluster-aware Compression for Event Relation Extraction 11 May 2024 · 0 repositories · arXiv:2405.06890
-
An Initial Introduction to Cooperative Multi-Agent Reinforcement Learning 10 May 2024 · 0 repositories · arXiv:2405.06161
-
A Survey on RAG Meeting LLMs: Towards Retrieval-Augmented Large Language Models 10 May 2024 · 0 repositories · arXiv:2405.06211
-
An Assessment of Model-On-Model Deception 10 May 2024 · 0 repositories · arXiv:2405.12999
-
CANAL -- Cyber Activity News Alerting Language Model: Empirical Approach vs. Expensive LLM 10 May 2024 · 0 repositories · arXiv:2405.06772
-
Characterizing the Accuracy -- Efficiency Trade-off of Low-rank Decomposition in Language Models 10 May 2024 · 0 repositories · arXiv:2405.06626
-
ChatGPTest: opportunities and cautionary tales of utilizing AI for questionnaire pretesting 10 May 2024 · 0 repositories · arXiv:2405.06329
-
A Mixture of Experts Approach to 3D Human Motion Prediction 9 May 2024 · 1 repository · arXiv:2405.06088
-
Can large language models understand uncommon meanings of common words? 9 May 2024 · 0 repositories · arXiv:2405.05741
-
Digital Diagnostics: The Potential Of Large Language Models In Recognizing Symptoms Of Common Illnesses 9 May 2024 · 0 repositories · arXiv:2405.06712
-
Ditto: Quantization-aware Secure Inference of Transformers upon MPC 9 May 2024 · 1 repository · arXiv:2405.05525
-
Iris: An AI-Driven Virtual Tutor For Computer Science Education 9 May 2024 · 0 repositories · arXiv:2405.08008
-
People cannot distinguish GPT-4 from a human in a Turing test 9 May 2024 · 0 repositories · arXiv:2405.08007
-
Reddit-Impacts: A Named Entity Recognition Dataset for Analyzing Clinical and Social Effects of Substance Use Derived from Social Media 9 May 2024 · 0 repositories · arXiv:2405.06145
-
Unveiling the Competitive Dynamics: A Comparative Evaluation of American and Chinese LLMs 9 May 2024 · 0 repositories · arXiv:2405.06713
-
AirGapAgent: Protecting Privacy-Conscious Conversational Agents 8 May 2024 · 0 repositories · arXiv:2405.05175
-
CARE-SD: Classifier-based analysis for recognizing and eliminating stigmatizing and doubt marker labels in electronic health records: model development and validation 8 May 2024 · 1 repository · arXiv:2405.05204
-
Evaluating Students' Open-ended Written Responses with LLMs: Using the RAG Framework for GPT-3.5, GPT-4, Claude-3, and Mistral-Large 8 May 2024 · 0 repositories · arXiv:2405.05444
-
LLMs Can Patch Up Missing Relevance Judgments in Evaluation 8 May 2024 · 0 repositories · arXiv:2405.04727
-
Utilizing Large Language Models to Generate Synthetic Data to Increase the Performance of BERT-Based Neural Networks 8 May 2024 · 0 repositories · arXiv:2405.06695
-
An LLM-Tool Compiler for Fused Parallel Function Calling 7 May 2024 · 0 repositories · arXiv:2405.17438
-
Enriched BERT Embeddings for Scholarly Publication Classification 7 May 2024 · 1 repository · arXiv:2405.04136
-
ERATTA: Extreme RAG for Table To Answers with Large Language Models 7 May 2024 · 0 repositories · arXiv:2405.03963
-
Evaluating Text Summaries Generated by Large Language Models Using OpenAI's GPT 7 May 2024 · 0 repositories · arXiv:2405.04053
-
GPT-Enabled Cybersecurity Training: A Tailored Approach for Effective Awareness 7 May 2024 · 0 repositories · arXiv:2405.04138
-
How does GPT-2 Predict Acronyms? Extracting and Understanding a Circuit via Mechanistic Interpretability 7 May 2024 · 1 repository · arXiv:2405.04156Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 1 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
Long Context Alignment with Short Instructions and Synthesized Positions 7 May 2024 · 0 repositories · arXiv:2405.03939
-
Remote Diffusion 7 May 2024 · 0 repositories · arXiv:2405.04717
-
Revisiting Character-level Adversarial Attacks for Language Models 7 May 2024 · 1 repository · arXiv:2405.04346Syntology official (archive's flag): 23 ran · 23 ran (of which 3 constructed an object rather than computing a result; 13 with no instrument failure: 0 honoured, 0 violated, 13 with no contract checked; 10 where Syntology's instrument failed) · 8 unverified (of 31 harvested samples)
-
Robust Implementation of Retrieval-Augmented Generation on Edge-based Computing-in-Memory Architectures 7 May 2024 · 0 repositories · arXiv:2405.04700
-
SUTRA: Scalable Multilingual Language Model Architecture 7 May 2024 · 0 repositories · arXiv:2405.06694
-
The Silicon Ceiling: Auditing GPT's Race and Gender Biases in Hiring 7 May 2024 · 0 repositories · arXiv:2405.04412
-
Utilizing GPT to Enhance Text Summarization: A Strategy to Minimize Hallucinations 7 May 2024 · 0 repositories · arXiv:2405.04039
-
Anchored Answers: Unravelling Positional Bias in GPT-2's Multiple-Choice Questions 6 May 2024 · 1 repository · arXiv:2405.03205
-
Characterizing the Dilemma of Performance and Index Size in Billion-Scale Vector Search and Breaking It with Second-Tier Memory 6 May 2024 · 0 repositories · arXiv:2405.03267
-
Compressing Long Context for Enhancing RAG with AMR-based Concept Distillation 6 May 2024 · 0 repositories · arXiv:2405.03085
-
Detecting Android Malware: From Neural Embeddings to Hands-On Validation with BERTroid 6 May 2024 · 0 repositories · arXiv:2405.03620
-
Detecting Anti-Semitic Hate Speech using Transformer-based Large Language Models 6 May 2024 · 0 repositories · arXiv:2405.03794
-
ERAGent: Enhancing Retrieval-Augmented Language Models with Improved Accuracy, Efficiency, and Personalization 6 May 2024 · 1 repository · arXiv:2405.06683
-
Hire Me or Not? Examining Language Model's Behavior with Occupation Attributes 6 May 2024 · 1 repository · arXiv:2405.06687Syntology official (archive's flag): 10 ran · 10 ran (of which 0 constructed an object rather than computing a result; 10 with no instrument failure: 0 honoured, 0 violated, 10 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 13 harvested samples) · 13 pointer-only (licence)
-
Large Language Models Reveal Information Operation Goals, Tactics, and Narrative Frames 6 May 2024 · 1 repository · arXiv:2405.03688
-
Structure-Preserving Network Compression Via Low-Rank Induced Training Through Linear Layers Composition 6 May 2024 · 1 repository · arXiv:2405.03089
-
Can Large Language Models Make the Grade? An Empirical Study Evaluating LLMs Ability to Mark Short Answer Questions in K-12 Education 5 May 2024 · 0 repositories · arXiv:2405.02985
-
Labeling supervised fine-tuning data with the scaling law 5 May 2024 · 2 repositories · arXiv:2405.02817
-
Leveraging Lecture Content for Improved Feedback: Explorations with GPT-4 and Retrieval Augmented Generation 5 May 2024 · 0 repositories · arXiv:2405.06681
-
Overconfidence is Key: Verbalized Uncertainty Evaluation in Large Language and Vision-Language Models 5 May 2024 · 0 repositories · arXiv:2405.02917
-
Stochastic RAG: End-to-End Retrieval-Augmented Generation through Expected Utility Maximization 5 May 2024 · 0 repositories · arXiv:2405.02816
-
Unraveling the Dominance of Large Language Models Over Transformer Models for Bangla Natural Language Inference: A Comprehensive Study 5 May 2024 · 1 repository · arXiv:2405.02937
-
A Combination of BERT and Transformer for Vietnamese Spelling Correction 4 May 2024 · 0 repositories · arXiv:2405.02573
-
Analyzing Narrative Processing in Large Language Models (LLMs): Using GPT4 to test BERT 3 May 2024 · 0 repositories · arXiv:2405.02024
-
Comparative Analysis of Retrieval Systems in the Real World 3 May 2024 · 0 repositories · arXiv:2405.02048
-
DALLMi: Domain Adaption for LLM-based Multi-label Classifier 3 May 2024 · 1 repository · arXiv:2405.01883
-
Evaluating Large Language Models for Structured Science Summarization in the Open Research Knowledge Graph 3 May 2024 · 0 repositories · arXiv:2405.02105
-
Exploiting ChatGPT for Diagnosing Autism-Associated Language Disorders and Identifying Distinct Features 3 May 2024 · 1 repository · arXiv:2405.01799
-
Exploring Combinatorial Problem Solving with Large Language Models: A Case Study on the Travelling Salesman Problem Using GPT-3.5 Turbo 3 May 2024 · 0 repositories · arXiv:2405.01997
-
Attribution in Scientific Literature: New Benchmark and Methods 3 May 2024 · 0 repositories · arXiv:2405.02228
-
Structural Pruning of Pre-trained Language Models via Neural Architecture Search 3 May 2024 · 1 repository · arXiv:2405.02267
-
A Survey on Large Language Models for Critical Societal Domains: Finance, Healthcare, and Law 2 May 2024 · 1 repository · arXiv:2405.01769
-
Bayesian Optimization with LLM-Based Acquisition Functions for Natural Language Preference Elicitation 2 May 2024 · 0 repositories · arXiv:2405.00981
-
Early Transformers: A study on Efficient Training of Transformer Models through Early-Bird Lottery Tickets 2 May 2024 · 0 repositories · arXiv:2405.02353
-
Investigating Wit, Creativity, and Detectability of Large Language Models in Domain-Specific Writing Style Adaptation of Reddit's Showerthoughts 2 May 2024 · 1 repository · arXiv:2405.01660
-
Progressive Feedforward Collapse of ResNet Training 2 May 2024 · 0 repositories · arXiv:2405.00985
-
The Effectiveness of LLMs as Annotators: A Comparative Overview and Empirical Analysis of Direct Representation 2 May 2024 · 0 repositories · arXiv:2405.01299
-
A Named Entity Recognition and Topic Modeling-based Solution for Locating and Better Assessment of Natural Disasters in Social Media 1 May 2024 · 0 repositories · arXiv:2405.00903
-
CourseAssist: Pedagogically Appropriate AI Tutor for Computer Science Education 1 May 2024 · 0 repositories · arXiv:2407.10246
-
How Can I Improve? Using GPT to Highlight the Desired and Undesired Parts of Open-ended Responses 1 May 2024 · 0 repositories · arXiv:2405.00291
-
Integrating A.I. in Higher Education: Protocol for a Pilot Study with 'SAMCares: An Adaptive Learning Hub' 1 May 2024 · 1 repository · arXiv:2405.00330
-
Opinion Mining Using Pre-Trained Large Language Models: Identifying the Type, Polarity, Intensity, Expression, and Source of Private States 1 May 2024 · 1 repository
-
Better & Faster Large Language Models via Multi-token Prediction 30 Apr 2024 · 1 repository · arXiv:2404.19737
-
Can Large Language Models put 2 and 2 together? Probing for Entailed Arithmetical Relationships 30 Apr 2024 · 0 repositories · arXiv:2404.19432
-
Do Large Language Models Understand Conversational Implicature -- A case study with a chinese sitcom 30 Apr 2024 · 1 repository · arXiv:2404.19509Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 1 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified (of 6 harvested samples) · 6 pointer-only (licence)
-
Graph Neural Network Approach to Semantic Type Detection in Tables 30 Apr 2024 · 1 repository · arXiv:2405.00123
-
Graphical Reasoning: LLM-based Semi-Open Relation Extraction 30 Apr 2024 · 1 repository · arXiv:2405.00216
-
PANGeA: Procedural Artificial Narrative using Generative AI for Turn-Based Video Games 30 Apr 2024 · 0 repositories · arXiv:2404.19721
-
Towards a Search Engine for Machines: Unified Ranking for Multiple Retrieval-Augmented Large Language Models 30 Apr 2024 · 1 repository · arXiv:2405.00175
-
TuBA: Cross-Lingual Transferability of Backdoor Attacks in LLMs with Instruction Tuning 30 Apr 2024 · 0 repositories · arXiv:2404.19597
-
How secure is AI-generated Code: A Large-Scale Comparison of Large Language Models 29 Apr 2024 · 1 repository · arXiv:2404.18353
-
Evaluating and Mitigating Linguistic Discrimination in Large Language Models 29 Apr 2024 · 0 repositories · arXiv:2404.18534
-
FeDeRA:Efficient Fine-tuning of Language Models in Federated Learning Leveraging Weight Decomposition 29 Apr 2024 · 0 repositories · arXiv:2404.18848