Methods › General › Regularization › Entropy Regularization › Papers, page 5
Entropy Regularization
Papers archive 2025-07-28
archive papers tagged: 1,128 · with a code link: 451 · where Syntology ran a sample: 156 (129 with a run with no instrument failure, 27 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (156 of 1,128 tagged: 129 with a run with no instrument failure, 27 where every run was a failure of Syntology's instrument)
Page 5 of 12: papers 401 to 500 of 1,128, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
MAPO: Advancing Multilingual Reasoning through Multilingual Alignment-as-Preference Optimization 12 Jan 2024 · 1 repository · arXiv:2401.06838Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 1 where Syntology's instrument failed) · 4 unverified (of 10 harvested samples) · 10 pointer-only (licence)
-
Autonomous Navigation of Tractor-Trailer Vehicles through Roundabout Intersections 10 Jan 2024 · 0 repositories · arXiv:2401.04980
-
Feedback-Guided Autonomous Driving 1 Jan 2024 · 0 repositories
-
Beyond PID Controllers: PPO with Neuralized PID Policy for Proton Beam Intensity Control in Mu2e 28 Dec 2023 · 0 repositories · arXiv:2312.17372
-
Preference as Reward, Maximum Preference Optimization with Importance Sampling 27 Dec 2023 · 0 repositories · arXiv:2312.16430
-
Some things are more CRINGE than others: Iterative Preference Optimization with the Pairwise Cringe Loss 27 Dec 2023 · 1 repository · arXiv:2312.16682Syntology 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 8 harvested samples) · 1 pointer-only (licence)
-
Agent based modelling for continuously varying supply chains 24 Dec 2023 · 0 repositories · arXiv:2312.15502
-
DriveLM: Driving with Graph Visual Question Answering 21 Dec 2023 · 3 repositories · arXiv:2312.14150
-
Realistic Rainy Weather Simulation for LiDARs in CARLA Simulator 20 Dec 2023 · 1 repository · arXiv:2312.12772
-
Colored Noise in PPO: Improved Exploration and Performance through Correlated Action Sampling 18 Dec 2023 · 1 repository · arXiv:2312.11091
-
Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint 18 Dec 2023 · 3 repositories · arXiv:2312.11456
-
DriveMLM: Aligning Multi-Modal Large Language Models with Behavioral Planning States for Autonomous Driving 14 Dec 2023 · 2 repositories · arXiv:2312.09245Syntology community repositories only · 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
Gradient Informed Proximal Policy Optimization 14 Dec 2023 · 1 repository · arXiv:2312.08710Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 5 harvested samples) · 5 pointer-only (licence)
-
Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations 14 Dec 2023 · 3 repositories · arXiv:2312.08935
-
Challenges of YOLO Series for Object Detection in Extremely Heavy Rain: CALRA Simulator based Synthetic Evaluation Dataset 13 Dec 2023 · 0 repositories · arXiv:2312.07976
-
The Effective Horizon Explains Deep RL Performance in Stochastic Environments 13 Dec 2023 · 1 repository · arXiv:2312.08369
-
A dynamical clipping approach with task feedback for Proximal Policy Optimization 12 Dec 2023 · 0 repositories · arXiv:2312.07624
-
SkyScenes: A Synthetic Dataset for Aerial Scene Understanding 11 Dec 2023 · 0 repositories · arXiv:2312.06719
-
Graph-based Prediction and Planning Policy Network (GP3Net) for scalable self-driving in dynamic environments using Deep Reinforcement Learning 10 Dec 2023 · 0 repositories · arXiv:2312.05784
-
Can language agents be alternatives to PPO? A Preliminary Empirical Study On OpenAI Gym 6 Dec 2023 · 1 repository · arXiv:2312.03290
-
A Reliable Representation with Bidirectional Transition Model for Visual Reinforcement Learning Generalization 4 Dec 2023 · 0 repositories · arXiv:2312.01915
-
Data-efficient Deep Reinforcement Learning for Vehicle Trajectory Control 30 Nov 2023 · 0 repositories · arXiv:2311.18393
-
EpiTESTER: Testing Autonomous Vehicles with Epigenetic Algorithm and Attention Mechanism 30 Nov 2023 · 1 repository · arXiv:2312.00207
-
Safe Reinforcement Learning in a Simulated Robotic Arm 28 Nov 2023 · 0 repositories · arXiv:2312.09468
-
Automated Lane Merging via Game Theory and Branch Model Predictive Control 25 Nov 2023 · 1 repository · arXiv:2311.14916
-
On optimal tracking portfolio in incomplete markets: The reinforcement learning approach 24 Nov 2023 · 0 repositories · arXiv:2311.14318
-
Variational Annealing on Graphs for Combinatorial Optimization 23 Nov 2023 · 1 repository · arXiv:2311.14156Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 0 where Syntology's instrument failed) · 9 unverified (of 16 harvested samples) · 16 pointer-only (licence)
-
Attacking Motion Planners Using Adversarial Perception Errors 21 Nov 2023 · 0 repositories · arXiv:2311.12722
-
Nav-Q: Quantum Deep Reinforcement Learning for Collision-Free Navigation of Self-Driving Cars 20 Nov 2023 · 1 repository · arXiv:2311.12875
-
Bridging Data-Driven and Knowledge-Driven Approaches for Safety-Critical Scenario Generation in Automated Vehicle Validation 18 Nov 2023 · 0 repositories · arXiv:2311.10937
-
Automatic Generation of Scenarios for System-level Simulation-based Verification of Autonomous Driving Systems 16 Nov 2023 · 0 repositories · arXiv:2311.09784
-
Safety Aware Autonomous Path Planning Using Model Predictive Reinforcement Learning for Inland Waterways 16 Nov 2023 · 0 repositories · arXiv:2311.09878
-
Clipped-Objective Policy Gradients for Pessimistic Policy Optimization 10 Nov 2023 · 1 repository · arXiv:2311.05846
-
Black-Box Prompt Optimization: Aligning Large Language Models without Model Training 7 Nov 2023 · 1 repository · arXiv:2311.04155Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 7 harvested samples)
-
Epidemic Decision-making System Based Federated Reinforcement Learning 3 Nov 2023 · 0 repositories · arXiv:2311.01749
-
Robust Adversarial Reinforcement Learning via Bounded Rationality Curricula 3 Nov 2023 · 0 repositories · arXiv:2311.01642
-
Dropout Strategy in Reinforcement Learning: Limiting the Surrogate Objective Variance in Policy Optimization Methods 31 Oct 2023 · 0 repositories · arXiv:2310.20380
-
Bird's Eye View Based Pretrained World model for Visual Navigation 28 Oct 2023 · 0 repositories · arXiv:2310.18847
-
Reward Scale Robustness for Proximal Policy Optimization via DreamerV3 Tricks 26 Oct 2023 · 0 repositories · arXiv:2310.17805Syntology 4 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples) · 1 pointer-only (licence)
-
CycleAlign: Iterative Distillation from Black-box LLM to White-box Models for Better Human Alignment 25 Oct 2023 · 0 repositories · arXiv:2310.16271
-
SuperHF: Supervised Iterative Learning from Human Feedback 25 Oct 2023 · 1 repository · arXiv:2310.16763Syntology official: harvested, nothing ran · 0 ran · 1 unverified (of 1 harvested sample)
-
Safe Navigation: Training Autonomous Vehicles using Deep Reinforcement Learning in CARLA 23 Oct 2023 · 1 repository · arXiv:2311.10735Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 2 harvested samples)
-
Automatic Unit Test Data Generation and Actor-Critic Reinforcement Learning for Code Synthesis 20 Oct 2023 · 2 repositories · arXiv:2310.13669
-
LeTFuser: Light-weight End-to-end Transformer-Based Sensor Fusion for Autonomous Driving with Multi-Task Learning 19 Oct 2023 · 1 repository · arXiv:2310.13135
-
ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models 16 Oct 2023 · 3 repositories · arXiv:2310.10505Syntology official (archive's flag): 11 ran · 11 ran (of which 0 constructed an object rather than computing a result; 10 with no instrument failure: 0 honoured, 0 violated, 10 with no contract checked; 1 where Syntology's instrument failed) · 2 unverified (of 13 harvested samples) · 13 pointer-only (licence)
-
Optimizing the Placement of Roadside LiDARs for Autonomous Driving 11 Oct 2023 · 0 repositories · arXiv:2310.07247
-
Distributional Soft Actor-Critic with Three Refinements 9 Oct 2023 · 2 repositories · arXiv:2310.05858Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 5 harvested samples) · 5 pointer-only (licence)
-
Reinforcement Learning in the Era of LLMs: What is Essential? What is needed? An RL Perspective on RLHF, Prompting, and Beyond 9 Oct 2023 · 2 repositories · arXiv:2310.06147Syntology 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
DialCoT Meets PPO: Decomposing and Exploring Reasoning Paths in Smaller Language Models 8 Oct 2023 · 1 repository · arXiv:2310.05074
-
FP3O: Enabling Proximal Policy Optimization in Multi-Agent Cooperation with Parameter-Sharing Versatility 8 Oct 2023 · 0 repositories · arXiv:2310.05053
-
Proximal Policy Optimization-Based Reinforcement Learning Approach for DC-DC Boost Converter Control: A Comparative Evaluation Against Traditional Control Techniques 4 Oct 2023 · 0 repositories · arXiv:2310.02945
-
Reward Model Ensembles Help Mitigate Overoptimization 4 Oct 2023 · 2 repositories · arXiv:2310.02743Syntology official (archive's flag): 2 ran · 5 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 5 where Syntology's instrument failed) · 0 unverified (of 5 harvested samples) · 3 pointer-only (licence)
-
Adapting LLM Agents with Universal Feedback in Communication 1 Oct 2023 · 0 repositories · arXiv:2310.01444
-
Deep Reinforcement Learning for Autonomous Vehicle Intersection Navigation 30 Sep 2023 · 0 repositories · arXiv:2310.08595
-
Pairwise Proximal Policy Optimization: Harnessing Relative Feedback for LLM Alignment 30 Sep 2023 · 0 repositories · arXiv:2310.00212
-
CARLA: Adjusted common average referencing for cortico-cortical evoked potential data 29 Sep 2023 · 1 repository · arXiv:2310.00185
-
Cleanba: A Reproducible and Efficient Distributed Reinforcement Learning Platform 29 Sep 2023 · 1 repository · arXiv:2310.00036Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified (of 8 harvested samples) · 8 pointer-only (licence)
-
Autonomous Driving using Spiking Neural Networks on Dynamic Vision Sensor Data: A Case Study of Traffic Light Change Detection 27 Sep 2023 · 1 repository · arXiv:2311.09225
-
SHACIRA: Scalable HAsh-grid Compression for Implicit Neural Representations 27 Sep 2023 · 0 repositories · arXiv:2309.15848
-
Effective Multi-Agent Deep Reinforcement Learning Control with Relative Entropy Regularization 26 Sep 2023 · 1 repository · arXiv:2309.14727
-
Don't throw away your value model! Generating more preferable text with Value-Guided Monte-Carlo Tree Search decoding 26 Sep 2023 · 0 repositories · arXiv:2309.15028
-
Enhancing data efficiency in reinforcement learning: a novel imagination mechanism based on mesh information propagation 25 Sep 2023 · 2 repositories · arXiv:2309.14243
-
Boosting Offline Reinforcement Learning for Autonomous Driving with Hierarchical Latent Skills 24 Sep 2023 · 0 repositories · arXiv:2309.13614
-
Learning Actions and Control of Focus of Attention with a Log-Polar-like Sensor 22 Sep 2023 · 0 repositories · arXiv:2309.12634
-
Hierarchical Adaptive Value Estimation for Multi-modal Visual Reinforcement Learning 21 Sep 2023 · 1 repository
-
Learning Better with Less: Effective Augmentation for Sample-Efficient Visual Reinforcement Learning 21 Sep 2023 · 1 repository
-
Learning to Drive Anywhere 21 Sep 2023 · 0 repositories · arXiv:2309.12295
-
RRHF: Rank Responses to Align Language Models with Human Feedback 21 Sep 2023 · 1 repository
-
Adaptive Multi-Agent Deep Reinforcement Learning for Timely Healthcare Interventions 20 Sep 2023 · 0 repositories · arXiv:2309.10980
-
Distributional Estimation of Data Uncertainty for Surveillance Face Anti-spoofing 18 Sep 2023 · 0 repositories · arXiv:2309.09485
-
Privileged to Predicted: Towards Sensorimotor Reinforcement Learning for Urban Driving 18 Sep 2023 · 0 repositories · arXiv:2309.09756
-
Stabilizing RLHF through Advantage Model and Selective Rehearsal 18 Sep 2023 · 0 repositories · arXiv:2309.10202
-
Exploring the impact of low-rank adaptation on the performance, efficiency, and regularization of RLHF 16 Sep 2023 · 1 repository · arXiv:2309.09055Syntology official (archive's flag): 12 ran · 12 ran (of which 0 constructed an object rather than computing a result; 12 with no instrument failure: 0 honoured, 0 violated, 12 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 13 harvested samples)
-
What Matters to Enhance Traffic Rule Compliance of Imitation Learning for End-to-End Autonomous Driving 14 Sep 2023 · 0 repositories · arXiv:2309.07808
-
Deep Reinforcement Learning Enabled Joint Deployment and Beamforming in STAR-RIS Assisted Networks 7 Sep 2023 · 0 repositories · arXiv:2309.03520
-
SKoPe3D: A Synthetic Dataset for Vehicle Keypoint Perception in 3D from Traffic Monitoring Cameras 4 Sep 2023 · 0 repositories · arXiv:2309.01324
-
Efficient RLHF: Reducing the Memory Usage of PPO 1 Sep 2023 · 0 repositories · arXiv:2309.00754
-
A reinforcement learning based construction material supply strategy using robotic crane and computer vision for building reconstruction after an earthquake 30 Aug 2023 · 0 repositories · arXiv:2308.16280
-
On Offline Evaluation of 3D Object Detection for Autonomous Driving 24 Aug 2023 · 0 repositories · arXiv:2308.12779
-
Aligning Language Models with Offline Learning from Human Feedback 23 Aug 2023 · 2 repositories · arXiv:2308.12050
-
Pedestrian Environment Model for Automated Driving 17 Aug 2023 · 1 repository · arXiv:2308.09080
-
Robust Autonomous Vehicle Pursuit without Expert Steering Labels 16 Aug 2023 · 0 repositories · arXiv:2308.08380
-
A Deep Recurrent-Reinforcement Learning Method for Intelligent AutoScaling of Serverless Functions 11 Aug 2023 · 1 repository · arXiv:2308.05937
-
Proximal Policy Optimization Actual Combat: Manipulating Output Tokenizer Length 10 Aug 2023 · 0 repositories · arXiv:2308.05585
-
Exploring the Physical World Adversarial Robustness of Vehicle Detection 7 Aug 2023 · 0 repositories · arXiv:2308.03476
-
Cognitive TransFuser: Semantics-guided Transformer-based Sensor Fusion for Improved Waypoint Prediction 4 Aug 2023 · 1 repository · arXiv:2308.02126
-
Interpretable End-to-End Driving Model for Implicit Scene Understanding 2 Aug 2023 · 0 repositories · arXiv:2308.01180
-
Safety Margins for Reinforcement Learning 25 Jul 2023 · 0 repositories · arXiv:2307.13642
-
Parallel Q-Learning: Scaling Off-policy Reinforcement Learning under Massively Parallel Simulation 24 Jul 2023 · 0 repositories · arXiv:2307.12983Syntology 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples)
-
Trust-aware Safe Control for Autonomous Navigation: Estimation of System-to-human Trust for Trust-adaptive Control Barrier Functions 24 Jul 2023 · 3 repositories · arXiv:2307.12815
-
Llama 2: Open Foundation and Fine-Tuned Chat Models 18 Jul 2023 · 19 repositories · arXiv:2307.09288Syntology community repositories only · 33 ran (of which 7 constructed an object rather than computing a result; 22 with no instrument failure: 1 honoured, 1 violated, 20 with no contract checked; 11 where Syntology's instrument failed) · 19 unverified (of 52 harvested samples) · 20 pointer-only (licence)
-
Navigating Uncertainty: The Role of Short-Term Trajectory Prediction in Autonomous Vehicle Safety 11 Jul 2023 · 1 repository · arXiv:2307.05288
-
Secrets of RLHF in Large Language Models Part I: PPO 11 Jul 2023 · 1 repository · arXiv:2307.04964
-
ContainerGym: A Real-World Reinforcement Learning Benchmark for Resource Allocation 6 Jul 2023 · 1 repository · arXiv:2307.02991
-
Navigation of micro-robot swarms for targeted delivery using reinforcement learning 30 Jun 2023 · 0 repositories · arXiv:2306.17598
-
Preference Ranking Optimization for Human Alignment 30 Jun 2023 · 1 repository · arXiv:2306.17492
-
Action and Trajectory Planning for Urban Autonomous Driving with Hierarchical Reinforcement Learning 28 Jun 2023 · 0 repositories · arXiv:2306.15968
-
Creating Valid Adversarial Examples of Malware 23 Jun 2023 · 1 repository · arXiv:2306.13587
-
Lower Complexity Adaptation for Empirical Entropic Optimal Transport 23 Jun 2023 · 1 repository · arXiv:2306.13580
-
Autonomous Driving with Deep Reinforcement Learning in CARLA Simulation 20 Jun 2023 · 0 repositories · arXiv:2306.11217