Methods › General › Learning Rate Schedules › Linear Warmup With Cosine Annealing › Papers, page 4
Linear Warmup With Cosine Annealing
Papers archive 2025-07-28
archive papers tagged: 3,797 · with a code link: 1,655 · where Syntology ran a sample: 602 (490 with a run with no instrument failure, 112 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (602 of 3,797 tagged: 490 with a run with no instrument failure, 112 where every run was a failure of Syntology's instrument)
Page 4 of 38: papers 301 to 400 of 3,797, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
LIBRA: Measuring Bias of Large Language Model from a Local Context 2 Feb 2025 · 0 repositories · arXiv:2502.01679
-
CoddLLM: Empowering Large Language Models for Data Analytics 1 Feb 2025 · 0 repositories · arXiv:2502.00329
-
Fast Solvers for Discrete Diffusion Models: Theory and Applications of High-Order Algorithms 1 Feb 2025 · 0 repositories · arXiv:2502.00234
-
Can AI Solve the Peer Review Crisis? A Large Scale Cross Model Experiment of LLMs' Performance and Biases in Evaluating over 1000 Economics Papers 31 Jan 2025 · 0 repositories · arXiv:2502.00070
-
KBQA-o1: Agentic Knowledge Base Question Answering with Monte Carlo Tree Search 31 Jan 2025 · 1 repository · arXiv:2501.18922Syntology official (archive's flag): 2 ran · 2 ran (of which 2 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 6 unverified; every one of the 2 samples that ran constructed an object rather than computing a result (of 8 harvested samples)
-
Large Language Models' Accuracy in Emulating Human Experts' Evaluation of Public Sentiments about Heated Tobacco Products on Social Media 31 Jan 2025 · 0 repositories · arXiv:2502.01658
-
AlphaAdam:Asynchronous Masked Optimization with Dynamic Alpha for Selective Updates 30 Jan 2025 · 0 repositories · arXiv:2501.18094
-
Economic Rationality under Specialization: Evidence of Decision Bias in AI Agents 30 Jan 2025 · 0 repositories · arXiv:2501.18190
-
General Embedding vs. Task-Specific Embedding: A Comparative Approach to Enhancing NLP Performance 30 Jan 2025 · 0 repositories
-
Structure Development in List-Sorting Transformers 30 Jan 2025 · 0 repositories · arXiv:2501.18666
-
Unraveling the Capabilities of Language Models in News Summarization 30 Jan 2025 · 1 repository · arXiv:2501.18128
-
WILDCHAT-50M: A Deep Dive Into the Role of Synthetic Data in Post-Training 30 Jan 2025 · 1 repository · arXiv:2501.18511Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 9 harvested samples)
-
Hybrid Graphs for Table-and-Text based Question Answering using LLMs 29 Jan 2025 · 0 repositories · arXiv:2501.17767
-
Graph of Attacks with Pruning: Optimizing Stealthy Jailbreak Prompt Generation for Enhanced LLM Content Moderation 28 Jan 2025 · 1 repository · arXiv:2501.18638
-
Open-Source Retrieval Augmented Generation Framework for Retrieving Accurate Medication Insights from Formularies for African Healthcare Workers 28 Jan 2025 · 0 repositories · arXiv:2502.15722
-
Kernels of Selfhood: GPT-4o shows humanlike patterns of cognitive consistency moderated by free choice 27 Jan 2025 · 0 repositories · arXiv:2502.07088
-
Weight-based Analysis of Detokenization in Language Models: Understanding the First Stage of Inference Without Inference 27 Jan 2025 · 0 repositories · arXiv:2501.15754
-
An AI-Driven Live Systematic Reviews in the Brain-Heart Interconnectome: Minimizing Research Waste and Advancing Evidence Synthesis 25 Jan 2025 · 1 repository · arXiv:2501.17181
-
An Attempt to Unraveling Token Prediction Refinement and Identifying Essential Layers of Large Language Models 25 Jan 2025 · 0 repositories · arXiv:2501.15054
-
Speech Translation Refinement using Large Language Models 25 Jan 2025 · 1 repository · arXiv:2501.15090
-
DarkMind: Latent Chain-of-Thought Backdoor in Customized LLMs 24 Jan 2025 · 0 repositories · arXiv:2501.18617
-
Prompt-Based Cost-Effective Evaluation and Operation of ChatGPT as a Computer Programming Teaching Assistant 24 Jan 2025 · 0 repositories · arXiv:2501.17176
-
Rethinking Table Instruction Tuning 24 Jan 2025 · 1 repository · arXiv:2501.14693
-
LLMs are Vulnerable to Malicious Prompts Disguised as Scientific Language 23 Jan 2025 · 0 repositories · arXiv:2501.14073
-
Exploring GPT's Ability as a Judge in Music Understanding 22 Jan 2025 · 1 repository · arXiv:2501.13261
-
Advancing the Understanding and Evaluation of AR-Generated Scenes: When Vision-Language Models Shine and Stumble 21 Jan 2025 · 1 repository · arXiv:2501.13964
-
Divide-Then-Aggregate: An Efficient Tool Learning Method via Parallel Tool Invocation 21 Jan 2025 · 0 repositories · arXiv:2501.12432
-
FOCUS: First Order Concentrated Updating Scheme 21 Jan 2025 · 0 repositories · arXiv:2501.12243
-
Harnessing Generative Pre-Trained Transformer for Datacenter Packet Trace Generation 21 Jan 2025 · 0 repositories · arXiv:2501.12033
-
Vision-Language Models for Automated Chest X-ray Interpretation: Leveraging ViT and GPT-2 21 Jan 2025 · 0 repositories · arXiv:2501.12356
-
Synthetic Data Can Mislead Evaluations: Membership Inference as Machine Text Detection 20 Jan 2025 · 0 repositories · arXiv:2501.11786
-
Trustformer: A Trusted Federated Transformer 20 Jan 2025 · 0 repositories · arXiv:2501.11706
-
From Arabic Text to Puzzles: LLM-Driven Development of Arabic Educational Crosswords 19 Jan 2025 · 0 repositories · arXiv:2501.11035
-
FSMoE: A Flexible and Scalable Training System for Sparse Mixture-of-Experts Models 18 Jan 2025 · 0 repositories · arXiv:2501.10714
-
Bias in Decision-Making for AI's Ethical Dilemmas: A Comparative Study of ChatGPT and Claude 17 Jan 2025 · 1 repository · arXiv:2501.10484
-
Confidence Estimation for Error Detection in Text-to-SQL Systems 16 Jan 2025 · 1 repository · arXiv:2501.09527
-
Perspective Transition of Large Language Models for Solving Subjective Tasks 16 Jan 2025 · 0 repositories · arXiv:2501.09265
-
Generative AI Takes a Statistics Exam: A Comparison of Performance between ChatGPT3.5, ChatGPT4, and ChatGPT4o-mini 15 Jan 2025 · 0 repositories · arXiv:2501.09171
-
The Impact of Big Five Personality Traits on AI Agent Decision-Making in Public Spaces: A Social Simulation Study 15 Jan 2025 · 0 repositories · arXiv:2503.15497
-
Investigating Energy Efficiency and Performance Trade-offs in LLM Inference Across Tasks and DVFS Settings 14 Jan 2025 · 0 repositories · arXiv:2501.08219
-
FinerWeb-10BT: Refining Web Data with LLM-Based Line-Level Filtering 13 Jan 2025 · 1 repository · arXiv:2501.07314
-
GPT as a Monte Carlo Language Tree: A Probabilistic Perspective 13 Jan 2025 · 0 repositories · arXiv:2501.07641
-
How GPT learns layer by layer 13 Jan 2025 · 1 repository · arXiv:2501.07108
-
ZNO-Eval: Benchmarking reasoning capabilities of large language models in Ukrainian 12 Jan 2025 · 1 repository · arXiv:2501.06715
-
Assessing instructor-AI cooperation for grading essay-type questions in an introductory sociology course 11 Jan 2025 · 1 repository · arXiv:2501.06461
-
OpenAI ChatGPT interprets Radiological Images: GPT-4 as a Medical Doctor for a Fast Check-Up 9 Jan 2025 · 0 repositories · arXiv:2501.06269
-
UAV-VLA: Vision-Language-Action System for Large Scale Aerial Mission Generation 9 Jan 2025 · 1 repository · arXiv:2501.05014
-
Integrating LLMs with ITS: Recent Advances, Potentials, Challenges, and Future Directions 8 Jan 2025 · 0 repositories · arXiv:2501.04437
-
Scaling Large Language Model Training on Frontier with Low-Bandwidth Partitioning 8 Jan 2025 · 0 repositories · arXiv:2501.04266
-
Finding A Voice: Evaluating African American Dialect Generation for Chatbot Technology 7 Jan 2025 · 1 repository · arXiv:2501.03441
-
Decoding fMRI Data into Captions using Prefix Language Modeling 5 Jan 2025 · 1 repository · arXiv:2501.02570
-
A Survey on Large Language Models with some Insights on their Capabilities and Limitations 3 Jan 2025 · 0 repositories · arXiv:2501.04040
-
AgentRefine: Enhancing Agent Generalization through Refinement Tuning 3 Jan 2025 · 0 repositories · arXiv:2501.01702
-
Large Language Models for Mental Health Diagnostic Assessments: Exploring The Potential of Large Language Models for Assisting with Mental Health Diagnostic Assessments -- The Depression and Anxiety Case 2 Jan 2025 · 0 repositories · arXiv:2501.01305
-
Predicting the Performance of Black-box LLMs through Self-Queries 2 Jan 2025 · 1 repository · arXiv:2501.01558Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 2 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
Column Property Annotation using Large Language Models 1 Jan 2025 · 1 repository
-
CRRG-CLIP: Automatic Generation of Chest Radiology Reports and Classification of Chest Radiographs 31 Dec 2024 · 1 repository · arXiv:2501.01989
-
Why Are Positional Encodings Nonessential for Deep Autoregressive Transformers? Revisiting a Petroglyph 31 Dec 2024 · 0 repositories · arXiv:2501.00659
-
Comparative Performance of Advanced NLP Models and LLMs in Multilingual Geo-Entity Detection 29 Dec 2024 · 0 repositories · arXiv:2412.20414
-
ELECTRA and GPT-4o: Cost-Effective Partners for Sentiment Analysis 29 Dec 2024 · 1 repository · arXiv:2501.00062
-
On Adversarial Robustness of Language Models in Transfer Learning 29 Dec 2024 · 0 repositories · arXiv:2501.00066
-
Assessing Text Classification Methods for Cyberbullying Detection on Social Media Platforms 27 Dec 2024 · 0 repositories · arXiv:2412.19928
-
DrivingWorld: Constructing World Model for Autonomous Driving via Video GPT 27 Dec 2024 · 1 repository · arXiv:2412.19505Syntology official (archive's flag): 9 ran · 9 ran (of which 0 constructed an object rather than computing a result; 8 with no instrument failure: 0 honoured, 1 violated, 7 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 9 harvested samples) · 2 pointer-only (licence)
-
Feature Alignment-Based Knowledge Distillation for Efficient Compression of Large Language Models 27 Dec 2024 · 0 repositories · arXiv:2412.19449
-
Generative Pretrained Embedding and Hierarchical Irregular Time Series Representation for Daily Living Activity Recognition 27 Dec 2024 · 1 repository · arXiv:2412.19732
-
Sentiment trading with large language models 26 Dec 2024 · 0 repositories · arXiv:2412.19245
-
SAFLITE: Fuzzing Autonomous Systems via Large Language Models 25 Dec 2024 · 0 repositories · arXiv:2412.18727
-
Whose Morality Do They Speak? Unraveling Cultural Bias in Multilingual Language Models 25 Dec 2024 · 0 repositories · arXiv:2412.18863
-
Do Language Models Understand the Cognitive Tasks Given to Them? Investigations with the N-Back Paradigm 24 Dec 2024 · 0 repositories · arXiv:2412.18120
-
Bridging Auditory Perception and Language Comprehension through MEG-Driven Encoding Models 22 Dec 2024 · 0 repositories · arXiv:2501.03246
-
On Fusing ChatGPT and Ensemble Learning in Discon-tinuous Named Entity Recognition in Health Corpora 22 Dec 2024 · 0 repositories · arXiv:2412.16976
-
PsychAdapter: Adapting LLM Transformers to Reflect Traits, Personality and Mental Health 22 Dec 2024 · 1 repository · arXiv:2412.16882
-
Robustness of Large Language Models Against Adversarial Attacks 22 Dec 2024 · 0 repositories · arXiv:2412.17011
-
TAR3D: Creating High-Quality 3D Assets via Next-Part Prediction 22 Dec 2024 · 0 repositories · arXiv:2412.16919
-
Assessing Social Alignment: Do Personality-Prompted Large Language Models Behave Like Humans? 21 Dec 2024 · 0 repositories · arXiv:2412.16772
-
Evaluating the Performance of Large Language Models in Scientific Claim Detection and Classification 21 Dec 2024 · 0 repositories · arXiv:2412.16486
-
Improving FIM Code Completions via Context & Curriculum Based Learning 21 Dec 2024 · 0 repositories · arXiv:2412.16589
-
Adversarial Robustness through Dynamic Ensemble Learning 20 Dec 2024 · 0 repositories · arXiv:2412.16254
-
Can LLMs Obfuscate Code? A Systematic Analysis of Large Language Models into Assembly Code Obfuscation 20 Dec 2024 · 0 repositories · arXiv:2412.16135
-
Human-Readable Adversarial Prompts: An Investigation into LLM Vulnerabilities Using Situational Context 20 Dec 2024 · 0 repositories · arXiv:2412.16359
-
Linguistic Features Extracted by GPT-4 Improve Alzheimer's Disease Detection based on Spontaneous Speech 20 Dec 2024 · 1 repository · arXiv:2412.15772
-
Graph-Convolutional Networks: Named Entity Recognition and Large Language Model Embedding in Document Clustering 19 Dec 2024 · 0 repositories · arXiv:2412.14867
-
How good is GPT at writing political speeches for the White House? 19 Dec 2024 · 0 repositories · arXiv:2412.14617
-
LLMs as mediators: Can they diagnose conflicts accurately? 19 Dec 2024 · 0 repositories · arXiv:2412.14675
-
Relational Programming with Foundation Models 19 Dec 2024 · 0 repositories · arXiv:2412.14515
-
ResoFilter: Fine-grained Synthetic Data Filtering for Large Language Models through Data-Parameter Resonance Analysis 19 Dec 2024 · 1 repository · arXiv:2412.14809
-
TOMG-Bench: Evaluating LLMs on Text-based Open Molecule Generation 19 Dec 2024 · 1 repository · arXiv:2412.14642
-
Autonomous Microscopy Experiments through Large Language Model Agents 18 Dec 2024 · 1 repository · arXiv:2501.10385
-
Mix-LN: Unleashing the Power of Deeper Layers by Combining Pre-LN and Post-LN 18 Dec 2024 · 1 repository · arXiv:2412.13795
-
Detecting Document-level Paraphrased Machine Generated Content: Mimicking Human Writing Style and Involving Discourse Features 17 Dec 2024 · 0 repositories · arXiv:2412.12679
-
LLMs are Also Effective Embedding Models: An In-depth Overview 17 Dec 2024 · 0 repositories · arXiv:2412.12591
-
Causal Diffusion Transformers for Generative Modeling 16 Dec 2024 · 1 repository · arXiv:2412.12095Syntology official (archive's flag): 7 ran · 8 ran (of which 3 constructed an object rather than computing a result; 6 with no instrument failure: 1 honoured, 0 violated, 5 with no contract checked; 2 where Syntology's instrument failed) · 2 unverified (of 10 harvested samples) · 10 pointer-only (licence)
-
Glimpse: Enabling White-Box Methods to Use Proprietary Models for Zero-Shot LLM-Generated Text Detection 16 Dec 2024 · 1 repository · arXiv:2412.11506Syntology official: not harvested · 0 ran · 1 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Look Ahead Text Understanding and LLM Stitching 16 Dec 2024 · 1 repository · arXiv:2412.17836
-
No More Adam: Learning Rate Scaling at Initialization is All You Need 16 Dec 2024 · 1 repository · arXiv:2412.11768
-
Priority-Aware Model-Distributed Inference at Edge Networks 16 Dec 2024 · 0 repositories · arXiv:2412.12371
-
Do large language vision models understand 3D shapes? 14 Dec 2024 · 1 repository · arXiv:2412.10908
-
Does Multiple Choice Have a Future in the Age of Generative AI? A Posttest-only RCT 13 Dec 2024 · 1 repository · arXiv:2412.10267
-
Reasoner Outperforms: Generative Stance Detection with Rationalization for Social Media 13 Dec 2024 · 0 repositories · arXiv:2412.10266
-
Adversarial Vulnerabilities in Large Language Models for Time Series Forecasting 11 Dec 2024 · 1 repository · arXiv:2412.08099