Methods › General › Regularization › Entropy Regularization › Papers, page 4
Entropy Regularization
Papers archive 2025-07-28
archive papers tagged: 1,128 · with a code link: 451 · where Syntology ran a sample: 156 (129 with a run with no instrument failure, 27 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (156 of 1,128 tagged: 129 with a run with no instrument failure, 27 where every run was a failure of Syntology's instrument)
Page 4 of 12: papers 301 to 400 of 1,128, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Flow of Reasoning:Training LLMs for Divergent Problem Solving with Minimal Examples 9 Jun 2024 · 1 repository · arXiv:2406.05673Syntology official (archive's flag): 11 ran · 11 ran (of which 0 constructed an object rather than computing a result; 11 with no instrument failure: 0 honoured, 0 violated, 11 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified (of 15 harvested samples)
-
Online Policy Distillation with Decision-Attention 8 Jun 2024 · 0 repositories · arXiv:2406.05488
-
Bench2Drive: Towards Multi-Ability Benchmarking of Closed-Loop End-To-End Autonomous Driving 6 Jun 2024 · 4 repositories · arXiv:2406.03877Syntology official (archive's flag): 13 ran · 13 ran (of which 0 constructed an object rather than computing a result; 10 with no instrument failure: 0 honoured, 0 violated, 10 with no contract checked; 3 where Syntology's instrument failed) · 4 unverified (of 17 harvested samples) · 17 pointer-only (licence)
-
Bisimulation Metrics are Optimal Transport Distances, and Can be Computed Efficiently 6 Jun 2024 · 0 repositories · arXiv:2406.04056
-
Optimal Rates of Convergence for Entropy Regularization in Discounted Markov Decision Processes 6 Jun 2024 · 0 repositories · arXiv:2406.04163
-
HackAtari: Atari Learning Environments for Robust and Continual Reinforcement Learning 6 Jun 2024 · 1 repository · arXiv:2406.03997Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
Transductive Off-policy Proximal Policy Optimization 6 Jun 2024 · 0 repositories · arXiv:2406.03894
-
Prompt-based Visual Alignment for Zero-shot Policy Transfer 5 Jun 2024 · 0 repositories · arXiv:2406.03250
-
Aligning Large Language Models via Fine-grained Supervision 4 Jun 2024 · 0 repositories · arXiv:2406.02756
-
LanEvil: Benchmarking the Robustness of Lane Detection to Environmental Illusions 3 Jun 2024 · 0 repositories · arXiv:2406.00934
-
Validity Learning on Failures: Mitigating the Distribution Shift in Autonomous Vehicle Planning 3 Jun 2024 · 0 repositories · arXiv:2406.01544
-
Self-Improving Robust Preference Optimization 3 Jun 2024 · 0 repositories · arXiv:2406.01660
-
Pedestrian intention prediction in Adverse Weather Conditions with Spiking Neural Networks and Dynamic Vision Sensors 1 Jun 2024 · 0 repositories · arXiv:2406.00473
-
A Deep Reinforcement Learning Approach for Trading Optimization in the Forex Market with Multi-Agent Asynchronous Distribution 30 May 2024 · 0 repositories · arXiv:2405.19982
-
Aquatic Navigation: A Challenging Benchmark for Deep Reinforcement Learning 30 May 2024 · 1 repository · arXiv:2405.20534
-
Entropy annealing for policy mirror descent in continuous time and space 30 May 2024 · 0 repositories · arXiv:2405.20250
-
Intrinsic Dynamics-Driven Generalizable Scene Representations for Vision-Oriented Decision-Making Applications 30 May 2024 · 1 repository · arXiv:2405.19736
-
Linear Function Approximation as a Computationally Efficient Method to Solve Classical Reinforcement Learning Challenges 27 May 2024 · 0 repositories · arXiv:2405.20350
-
SCaRL- A Synthetic Multi-Modal Dataset for Autonomous Driving 27 May 2024 · 0 repositories · arXiv:2405.17030
-
Symmetric Reinforcement Learning Loss for Robust Learning on Diverse Tasks and Model Scales 27 May 2024 · 1 repository · arXiv:2405.17618
-
Rewarded Region Replay (R3) for Policy Learning with Discrete Action Space 26 May 2024 · 1 repository · arXiv:2405.16383
-
Diffusion-based Reinforcement Learning via Q-weighted Variational Policy Optimization 25 May 2024 · 1 repository · arXiv:2405.16173
-
Continuously Learning, Adapting, and Improving: A Dual-Process Approach to Autonomous Driving 24 May 2024 · 1 repository · arXiv:2405.15324
-
Extracting Heuristics from Large Language Models for Reward Shaping in Reinforcement Learning 24 May 2024 · 0 repositories · arXiv:2405.15194
-
AGILE: A Novel Reinforcement Learning Framework of LLM Agents 23 May 2024 · 1 repository · arXiv:2405.14751Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples) · 1 pointer-only (licence)
-
ChatScene: Knowledge-Enabled Safety-Critical Scenario Generation for Autonomous Vehicles 22 May 2024 · 1 repository · arXiv:2405.14062Syntology official (archive's flag): 1 ran · 1 ran (of which 1 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified; the one sample that ran constructed an object rather than computing a result (of 1 harvested sample)
-
MaskFuser: Masked Fusion of Joint Multi-Modal Tokenization for End-to-End Autonomous Driving 13 May 2024 · 0 repositories · arXiv:2405.07573
-
Value Augmented Sampling for Language Model Alignment and Personalization 10 May 2024 · 1 repository · arXiv:2405.06639Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
Proximal Policy Optimization with Adaptive Exploration 7 May 2024 · 1 repository · arXiv:2405.04664
-
Guidance Design for Escape Flight Vehicle Using Evolution Strategy Enhanced Deep Reinforcement Learning 4 May 2024 · 0 repositories · arXiv:2405.03711
-
D2PO: Discriminator-Guided DPO with Response Evaluation Models 2 May 2024 · 1 repository · arXiv:2405.01511Syntology official (archive's flag): 11 ran · 11 ran (of which 0 constructed an object rather than computing a result; 11 with no instrument failure: 0 honoured, 0 violated, 11 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 13 harvested samples) · 13 pointer-only (licence)
-
No Representation, No Trust: Connecting Representation, Collapse, and Trust Issues in PPO 1 May 2024 · 1 repository · arXiv:2405.00662Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 1 where Syntology's instrument failed) · 2 unverified (of 8 harvested samples) · 1 pointer-only (licence)
-
Guiding Attention in End-to-End Driving Models 30 Apr 2024 · 1 repository · arXiv:2405.00242
-
DPO Meets PPO: Reinforced Token Optimization for RLHF 29 Apr 2024 · 1 repository · arXiv:2404.18922Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 1 where Syntology's instrument failed) · 2 unverified (of 6 harvested samples) · 1 pointer-only (licence)
-
3D Extended Object Tracking by Fusing Roadside Sparse Radar Point Clouds and Pixel Keypoints 27 Apr 2024 · 2 repositories · arXiv:2404.17903
-
REBEL: Reinforcement Learning via Regressing Relative Rewards 25 Apr 2024 · 3 repositories · arXiv:2404.16767Syntology official (archive's flag): 16 ran · 16 ran (of which 0 constructed an object rather than computing a result; 12 with no instrument failure: 1 honoured, 0 violated, 11 with no contract checked; 4 where Syntology's instrument failed) · 4 unverified (of 20 harvested samples) · 6 pointer-only (licence)
-
ContextualFusion: Context-Based Multi-Sensor Fusion for 3D Object Detection in Adverse Operating Conditions 23 Apr 2024 · 0 repositories · arXiv:2404.14780
-
Is DPO Superior to PPO for LLM Alignment? A Comprehensive Study 16 Apr 2024 · 1 repository · arXiv:2404.10719
-
Joint Physical-Digital Facial Attack Detection Via Simulating Spoofing Clues 12 Apr 2024 · 3 repositories · arXiv:2404.08450
-
SEVD: Synthetic Event-based Vision Dataset for Ego and Fixed Traffic Perception 12 Apr 2024 · 1 repository · arXiv:2404.10540
-
Enhanced Cooperative Perception for Autonomous Vehicles Using Imperfect Communication 10 Apr 2024 · 0 repositories · arXiv:2404.08013
-
Synergy of Large Language Model and Model Driven Engineering for Automated Development of Centralized Vehicular Systems 8 Apr 2024 · 0 repositories · arXiv:2404.05508
-
Prompting Multi-Modal Tokens to Enhance End-to-End Autonomous Driving Imitation Learning with LLMs 7 Apr 2024 · 0 repositories · arXiv:2404.04869
-
A proximal policy optimization based intelligent home solar management 5 Apr 2024 · 0 repositories · arXiv:2404.03888
-
AGL-NET: Aerial-Ground Cross-Modal Global Localization with Varying Scales 4 Apr 2024 · 1 repository · arXiv:2404.03187
-
Addressing Loss of Plasticity and Catastrophic Forgetting in Continual Learning 31 Mar 2024 · 1 repository · arXiv:2404.00781Syntology official (archive's flag): 1 ran · 1 ran (of which 1 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified; the one sample that ran constructed an object rather than computing a result (of 1 harvested sample)
-
Human-compatible driving partners through data-regularized self-play reinforcement learning 28 Mar 2024 · 1 repository · arXiv:2403.19648
-
Scenario-Based Curriculum Generation for Multi-Agent Autonomous Driving 26 Mar 2024 · 1 repository · arXiv:2403.17805
-
DriveCoT: Integrating Chain-of-Thought Reasoning with End-to-End Driving 25 Mar 2024 · 0 repositories · arXiv:2403.16996
-
Policy Optimization finds Nash Equilibrium in Regularized General-Sum LQ Games 25 Mar 2024 · 0 repositories · arXiv:2404.00045
-
Predictable Interval MDPs through Entropy Regularization 25 Mar 2024 · 0 repositories · arXiv:2403.16711
-
Policy Mirror Descent with Lookahead 21 Mar 2024 · 1 repository · arXiv:2403.14156
-
JaxUED: A simple and useable UED library in Jax 19 Mar 2024 · 1 repository · arXiv:2403.13091
-
M2DA: Multi-Modal Fusion Transformer Incorporating Driver Attention for Autonomous Driving 19 Mar 2024 · 0 repositories · arXiv:2403.12552
-
AdaMER-CTC: Connectionist Temporal Classification with Adaptive Maximum Entropy Regularization for Automatic Speech Recognition 18 Mar 2024 · 0 repositories · arXiv:2403.11578
-
Driving Style Alignment for LLM-powered Driver Agent 17 Mar 2024 · 2 repositories · arXiv:2403.11368
-
Are you a robot? Detecting Autonomous Vehicles from Behavior Analysis 14 Mar 2024 · 0 repositories · arXiv:2403.09571
-
Right Place, Right Time! Dynamizing Topological Graphs for Embodied Navigation 14 Mar 2024 · 0 repositories · arXiv:2403.09905
-
Improving Reinforcement Learning from Human Feedback Using Contrastive Rewards 12 Mar 2024 · 0 repositories · arXiv:2403.07708
-
Tractable Joint Prediction and Planning over Discrete Behavior Modes for Urban Driving 12 Mar 2024 · 0 repositories · arXiv:2403.07232
-
Risk-Sensitive RL with Optimized Certainty Equivalents via Reduction to Standard RL 10 Mar 2024 · 0 repositories · arXiv:2403.06323
-
A Sinkhorn-type Algorithm for Constrained Optimal Transport 8 Mar 2024 · 0 repositories · arXiv:2403.05054
-
Teaching Large Language Models to Reason with Reinforcement Learning 7 Mar 2024 · 0 repositories · arXiv:2403.04642
-
COMMIT: Certifying Robustness of Multi-Sensor Fusion Systems against Semantic Attacks 4 Mar 2024 · 0 repositories · arXiv:2403.02329
-
Integrating Efficient Optimal Transport and Functional Maps For Unsupervised Shape Correspondence Learning 4 Mar 2024 · 0 repositories · arXiv:2403.01781
-
Tsallis Entropy Regularization for Linearly Solvable MDP and Linear Quadratic Regulator 4 Mar 2024 · 0 repositories · arXiv:2403.01805
-
Craftax: A Lightning-Fast Benchmark for Open-Ended Reinforcement Learning 26 Feb 2024 · 1 repository · arXiv:2402.16801Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 14 unverified (of 16 harvested samples)
-
Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs 22 Feb 2024 · 0 repositories · arXiv:2402.14740
-
Distributed Radiance Fields for Edge Video Compression and Metaverse Integration in Autonomous Driving 22 Feb 2024 · 0 repositories · arXiv:2402.14642
-
Hybrid Reasoning Based on Large Language Models for Autonomous Car Driving 21 Feb 2024 · 1 repository · arXiv:2402.13602
-
VADv2: End-to-End Vectorized Autonomous Driving via Probabilistic Planning 20 Feb 2024 · 1 repository · arXiv:2402.13243Syntology official (archive's flag): 1 ran · 3 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
Surpassing legacy approaches to PWR core reload optimization with single-objective Reinforcement learning 16 Feb 2024 · 0 repositories · arXiv:2402.11040
-
RS-DPO: A Hybrid Rejection Sampling and Direct Preference Optimization Method for Alignment of Large Language Models 15 Feb 2024 · 0 repositories · arXiv:2402.10038
-
Entropy-regularized Point-based Value Iteration 14 Feb 2024 · 1 repository · arXiv:2402.09388
-
Reducing Texture Bias of Deep Neural Networks via Edge Enhancing Diffusion 14 Feb 2024 · 1 repository · arXiv:2402.09530
-
SemTra: A Semantic Skill Translator for Cross-Domain Zero-Shot Policy Adaptation 12 Feb 2024 · 0 repositories · arXiv:2402.07418
-
Solving Deep Reinforcement Learning Tasks with Evolution Strategies and Linear Policy Networks 10 Feb 2024 · 1 repository · arXiv:2402.06912
-
Entropy-Regularized Token-Level Policy Optimization for Language Agent Reinforcement 9 Feb 2024 · 1 repository · arXiv:2402.06700Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 1 violated, 2 with no contract checked; 3 where Syntology's instrument failed) · 3 unverified (of 9 harvested samples) · 9 pointer-only (licence)
-
Attention-Enhanced Prioritized Proximal Policy Optimization for Adaptive Edge Caching 8 Feb 2024 · 0 repositories · arXiv:2402.14576
-
Convergence for Natural Policy Gradient on Infinite-State Queueing MDPs 7 Feb 2024 · 0 repositories · arXiv:2402.05274
-
Tuning the feedback controller gains is a simple way to improve autonomous driving performance 7 Feb 2024 · 0 repositories · arXiv:2402.05064
-
Averaging n-step Returns Reduces Variance in Reinforcement Learning 6 Feb 2024 · 0 repositories · arXiv:2402.03903Syntology 5 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 1 where Syntology's instrument failed) · 3 unverified (of 8 harvested samples) · 6 pointer-only (licence)
-
Learning to Generate Explainable Stock Predictions using Self-Reflective Large Language Models 6 Feb 2024 · 1 repository · arXiv:2402.03659Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
OASim: an Open and Adaptive Simulator based on Neural Rendering for Autonomous Driving 6 Feb 2024 · 1 repository · arXiv:2402.03830
-
DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models 5 Feb 2024 · 5 repositories · arXiv:2402.03300Syntology official (archive's flag): 8 ran · 14 ran (of which 0 constructed an object rather than computing a result; 10 with no instrument failure: 0 honoured, 0 violated, 10 with no contract checked; 4 where Syntology's instrument failed) · 10 unverified (of 24 harvested samples) · 3 pointer-only (licence)
-
BRAIn: Bayesian Reward-conditioned Amortized Inference for natural language generation from feedback 4 Feb 2024 · 0 repositories · arXiv:2402.02479
-
Hybrid-Prediction Integrated Planning for Autonomous Driving 4 Feb 2024 · 0 repositories · arXiv:2402.02426
-
Parametric-Task MAP-Elites 2 Feb 2024 · 0 repositories · arXiv:2402.01275
-
Equivalence of the Empirical Risk Minimization to Regularization on the Family of f-Divergences 1 Feb 2024 · 0 repositories · arXiv:2402.00501
-
CARFF: Conditional Auto-encoded Radiance Field for 3D Scene Forecasting 31 Jan 2024 · 0 repositories · arXiv:2401.18075
-
Simple Policy Optimization 29 Jan 2024 · 1 repository · arXiv:2401.16025
-
True Knowledge Comes from Practice: Aligning LLMs with Embodied Environments via Reinforcement Learning 25 Jan 2024 · 1 repository · arXiv:2401.14151Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
Linear Alignment: A Closed-form Solution for Aligning Human Preferences without Tuning and Feedback 21 Jan 2024 · 1 repository · arXiv:2401.11458Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
360ORB-SLAM: A Visual SLAM System for Panoramic Images with Depth Completion Network 19 Jan 2024 · 0 repositories · arXiv:2401.10560
-
LangProp: A code optimization framework using Large Language Models applied to driving 18 Jan 2024 · 1 repository · arXiv:2401.10314Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
ReFT: Reasoning with Reinforced Fine-Tuning 17 Jan 2024 · 1 repository · arXiv:2401.08967Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 9 unverified (of 14 harvested samples) · 13 pointer-only (licence)
-
CPPO: Continual Learning for Reinforcement Learning with Human Feedback 16 Jan 2024 · 0 repositories
-
Sum Throughput Maximization in Multi-BD Symbiotic Radio NOMA Network Assisted by Active-STAR-RIS 16 Jan 2024 · 0 repositories · arXiv:2401.08301
-
Beyond Sparse Rewards: Enhancing Reinforcement Learning with Language Model Critique in Text Generation 14 Jan 2024 · 0 repositories · arXiv:2401.07382
-
Aquarium: A Comprehensive Framework for Exploring Predator-Prey Dynamics through Multi-Agent Reinforcement Learning Algorithms 13 Jan 2024 · 1 repository · arXiv:2401.07056