Methods › Reinforcement Learning › Offline Reinforcement Learning Methods › DPO › Papers, page 4
Direct Preference Optimization
DPO
Papers archive 2025-07-28
archive papers tagged: 409 · with a code link: 184 · where Syntology ran a sample: 112 (98 with a run with no instrument failure, 14 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (112 of 409 tagged: 98 with a run with no instrument failure, 14 where every run was a failure of Syntology's instrument)
Page 4 of 5: papers 301 to 400 of 409, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Online DPO: Online Direct Preference Optimization with Fast-Slow Chasing 8 Jun 2024 · 0 repositories · arXiv:2406.05534
-
Aesthetic Post-Training Diffusion Models from Generic Preferences with Step-by-step Preference Optimization 6 Jun 2024 · 1 repository · arXiv:2406.04314Syntology official: harvested, nothing ran · 0 ran · 3 unverified (of 3 harvested samples)
-
Is Free Self-Alignment Possible? 5 Jun 2024 · 0 repositories · arXiv:2406.03642
-
LLMs Beyond English: Scaling the Multilingual Capability of LLMs with Cross-Lingual Feedback 3 Jun 2024 · 0 repositories · arXiv:2406.01771
-
Self-Improving Robust Preference Optimization 3 Jun 2024 · 0 repositories · arXiv:2406.01660
-
The Importance of Online Data: Understanding Preference Fine-tuning via Coverage 3 Jun 2024 · 0 repositories · arXiv:2406.01462
-
Direct Alignment of Language Models via Quality-Aware Self-Refinement 31 May 2024 · 0 repositories · arXiv:2405.21040
-
Exploratory Preference Optimization: Harnessing Implicit Q*-Approximation for Sample-Efficient RLHF 31 May 2024 · 0 repositories · arXiv:2405.21046
-
Learning to Clarify: Multi-turn Conversations with Action-Based Contrastive Self-Training 31 May 2024 · 0 repositories · arXiv:2406.00222
-
Self-Augmented Preference Optimization: Off-Policy Paradigms for Language Model Alignment 31 May 2024 · 1 repository · arXiv:2405.20830
-
Boost Your Human Image Generation Model via Direct Preference Optimization 30 May 2024 · 0 repositories · arXiv:2405.20216
-
Xwin-LM: Strong and Scalable Alignment Practice for LLMs 30 May 2024 · 1 repository · arXiv:2405.20335
-
Preference Learning Algorithms Do Not Learn Preference Rankings 29 May 2024 · 0 repositories · arXiv:2405.19534
-
Robust Preference Optimization through Reward Model Distillation 29 May 2024 · 0 repositories · arXiv:2405.19316
-
Unified Preference Optimization: Language Model Alignment Beyond the Preference Frontier 28 May 2024 · 0 repositories · arXiv:2405.17956
-
Online Merging Optimizers for Boosting Rewards and Mitigating Tax in Alignment 28 May 2024 · 1 repository · arXiv:2405.17931
-
Multi-Reference Preference Optimization for Large Language Models 26 May 2024 · 0 repositories · arXiv:2405.16388
-
Provably Mitigating Overoptimization in RLHF: Your SFT Loss is Implicitly an Adversarial Regularizer 26 May 2024 · 0 repositories · arXiv:2405.16436
-
Triple Preference Optimization: Achieving Better Alignment with Less Data in a Single Step Optimization 26 May 2024 · 1 repository · arXiv:2405.16681Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 6 harvested samples)
-
Direct Preference Optimization With Unobserved Preference Heterogeneity 23 May 2024 · 0 repositories · arXiv:2405.15065
-
MallowsPO: Fine-Tune Your LLM with Preference Dispersions 23 May 2024 · 0 repositories · arXiv:2405.14953
-
SimPO: Simple Preference Optimization with a Reference-Free Reward 23 May 2024 · 2 repositories · arXiv:2405.14734Syntology official (archive's flag): 5 ran · 6 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 1 violated, 4 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 6 harvested samples) · 1 pointer-only (licence)
-
Annotation-Efficient Preference Optimization for Language Model Alignment 22 May 2024 · 1 repository · arXiv:2405.13541
-
Curriculum Direct Preference Optimization for Diffusion and Consistency Models 22 May 2024 · 1 repository · arXiv:2405.13637Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 1 where Syntology's instrument failed) · 5 unverified (of 8 harvested samples) · 8 pointer-only (licence)
-
Model Editing as a Robust and Denoised variant of DPO: A Case Study on Toxicity 22 May 2024 · 2 repositories · arXiv:2405.13967Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 1 where Syntology's instrument failed) · 5 unverified (of 12 harvested samples)
-
Adversarial DPO: Harnessing Harmful Data for Reducing Toxicity with Minimal Impact on Coherence and Evasiveness in Dialogue Agents 21 May 2024 · 0 repositories · arXiv:2405.12900
-
Aligning Transformers with Continuous Feedback via Energy Rank Alignment 21 May 2024 · 2 repositories · arXiv:2405.12961
-
Mining the Explainability and Generalization: Fact Verification Based on Self-Instruction 21 May 2024 · 0 repositories · arXiv:2405.12579
-
OpenRLHF: An Easy-to-use, Scalable and High-performance RLHF Framework 20 May 2024 · 4 repositories · arXiv:2405.11143
-
Quantifying and Optimizing Global Faithfulness in Persona-driven Role-playing 13 May 2024 · 1 repository · arXiv:2405.07726Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 7 harvested samples)
-
Advanced Natural-based interaction for the ITAlian language: LLaMAntino-3-ANITA 11 May 2024 · 0 repositories · arXiv:2405.07101
-
Value Augmented Sampling for Language Model Alignment and Personalization 10 May 2024 · 1 repository · arXiv:2405.06639Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
MoDiPO: text-to-motion alignment via AI-feedback-driven Direct Preference Optimization 6 May 2024 · 0 repositories · arXiv:2405.03803
-
D2PO: Discriminator-Guided DPO with Response Evaluation Models 2 May 2024 · 1 repository · arXiv:2405.01511Syntology official (archive's flag): 11 ran · 11 ran (of which 0 constructed an object rather than computing a result; 11 with no instrument failure: 0 honoured, 0 violated, 11 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 13 harvested samples) · 13 pointer-only (licence)
-
Self-Play Preference Optimization for Language Model Alignment 1 May 2024 · 1 repository · arXiv:2405.00675Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples) · 1 pointer-only (licence)
-
Iterative Reasoning Preference Optimization 30 Apr 2024 · 0 repositories · arXiv:2404.19733
-
DPO Meets PPO: Reinforced Token Optimization for RLHF 29 Apr 2024 · 1 repository · arXiv:2404.18922Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 1 where Syntology's instrument failed) · 2 unverified (of 6 harvested samples) · 1 pointer-only (licence)
-
REBEL: Reinforcement Learning via Regressing Relative Rewards 25 Apr 2024 · 3 repositories · arXiv:2404.16767Syntology official (archive's flag): 16 ran · 16 ran (of which 0 constructed an object rather than computing a result; 12 with no instrument failure: 1 honoured, 0 violated, 11 with no contract checked; 4 where Syntology's instrument failed) · 4 unverified (of 20 harvested samples) · 6 pointer-only (licence)
-
Weak-to-Strong Extrapolation Expedites Alignment 25 Apr 2024 · 1 repository · arXiv:2404.16792
-
DPO: A Differential and Pointwise Control Approach to Reinforcement Learning 24 Apr 2024 · 0 repositories · arXiv:2404.15617Syntology 1 ran (of which 1 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified; the one sample that ran constructed an object rather than computing a result (of 1 harvested sample)
-
Insights into Alignment: Evaluating DPO and its Variants Across Multiple Tasks 23 Apr 2024 · 1 repository · arXiv:2404.14723
-
Filtered Direct Preference Optimization 22 Apr 2024 · 1 repository · arXiv:2404.13846Syntology official: harvested, nothing ran · 2 ran (of which 2 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified; every one of the 2 samples that ran constructed an object rather than computing a result (of 3 harvested samples)
-
From r to Q^*: Your Language Model is Secretly a Q-Function 18 Apr 2024 · 0 repositories · arXiv:2404.12358
-
OpenBezoar: Small, Cost-Effective and Open Models Trained on Mixes of Instruction Data 18 Apr 2024 · 1 repository · arXiv:2404.12195
-
Token-level Direct Preference Optimization 18 Apr 2024 · 1 repository · arXiv:2404.11999Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 5 harvested samples) · 1 pointer-only (licence)
-
Is DPO Superior to PPO for LLM Alignment? A Comprehensive Study 16 Apr 2024 · 1 repository · arXiv:2404.10719
-
Latent Distance Guided Alignment Training for Large Language Models 9 Apr 2024 · 0 repositories · arXiv:2404.06390
-
Binary Classifier Optimization for Large Language Model Alignment 6 Apr 2024 · 0 repositories · arXiv:2404.04656
-
Towards Analyzing and Understanding the Limitations of DPO: A Theoretical Perspective 6 Apr 2024 · 0 repositories · arXiv:2404.04626
-
Direct Preference Optimization of Video Large Multimodal Models from Language Model Reward 1 Apr 2024 · 1 repository · arXiv:2404.01258Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
Regularized Best-of-N Sampling with Minimum Bayes Risk Objective for Language Model Alignment 1 Apr 2024 · 1 repository · arXiv:2404.01054Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 5 unverified (of 10 harvested samples)
-
Comparing Bad Apples to Good Oranges: Aligning Large Language Models via Joint Preference Optimization 31 Mar 2024 · 1 repository · arXiv:2404.00530Syntology official (archive's flag): 11 ran · 11 ran (of which 0 constructed an object rather than computing a result; 11 with no instrument failure: 0 honoured, 0 violated, 11 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 11 harvested samples)
-
Extensive Self-Contrast Enables Feedback-Free Language Model Alignment 31 Mar 2024 · 2 repositories · arXiv:2404.00604
-
Configurable Safety Tuning of Language Models with Synthetic Preference Data 30 Mar 2024 · 1 repository · arXiv:2404.00495
-
Dialectical Alignment: Resolving the Tension of 3H and Security Threats of LLMs 30 Mar 2024 · 0 repositories · arXiv:2404.00486
-
Disentangling Length from Quality in Direct Preference Optimization 28 Mar 2024 · 1 repository · arXiv:2403.19159Syntology 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
Mixed Preference Optimization: Reinforcement Learning with Data Selection and Better Reference Model 28 Mar 2024 · 0 repositories · arXiv:2403.19443
-
sDPO: Don't Use Your Data All at Once 28 Mar 2024 · 0 repositories · arXiv:2403.19270
-
Detoxifying Large Language Models via Knowledge Editing 21 Mar 2024 · 1 repository · arXiv:2403.14472Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 3 unverified (of 4 harvested samples)
-
Human Alignment of Large Language Models through Online Preference Optimisation 13 Mar 2024 · 0 repositories · arXiv:2403.08635
-
Curry-DPO: Enhancing Alignment using Curriculum Learning & Ranked Preferences 12 Mar 2024 · 0 repositories · arXiv:2403.07230
-
Negating Negatives: Alignment with Human Negative Samples via Distributional Dispreference Optimization 6 Mar 2024 · 1 repository · arXiv:2403.03419
-
Enhancing LLM Safety via Constrained Direct Preference Optimization 4 Mar 2024 · 0 repositories · arXiv:2403.02475
-
Reward Model Learning vs. Direct Policy Optimization: A Comparative Analysis of Learning from Human Preferences 4 Mar 2024 · 0 repositories · arXiv:2403.01857
-
Trial and Error: Exploration-Based Trajectory Optimization for LLM Agents 4 Mar 2024 · 2 repositories · arXiv:2403.02502Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 1 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 7 harvested samples) · 7 pointer-only (licence)
-
Provably Robust DPO: Aligning Language Models with Noisy Feedback 1 Mar 2024 · 0 repositories · arXiv:2403.00409
-
Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs 22 Feb 2024 · 0 repositories · arXiv:2402.14740
-
Smaug: Fixing Failure Modes of Preference Optimisation with DPO-Positive 20 Feb 2024 · 2 repositories · arXiv:2402.13228
-
Direct Large Language Model Alignment Through Self-Rewarding Contrastive Prompt Distillation 19 Feb 2024 · 1 repository · arXiv:2402.11907Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
NOTE: Notable generation Of patient Text summaries through Efficient approach based on direct preference optimization 19 Feb 2024 · 0 repositories · arXiv:2402.11882
-
Direct Preference Optimization with an Offset 16 Feb 2024 · 2 repositories · arXiv:2402.10571Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 6 harvested samples) · 1 pointer-only (licence)
-
Multi-modal Preference Alignment Remedies Degradation of Visual Instruction Tuning on Language Models 16 Feb 2024 · 2 repositories · arXiv:2402.10884Syntology official (archive's flag): 4 ran · 5 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 1 honoured, 2 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 5 harvested samples) · 5 pointer-only (licence)
-
RS-DPO: A Hybrid Rejection Sampling and Direct Preference Optimization Method for Alignment of Large Language Models 15 Feb 2024 · 0 repositories · arXiv:2402.10038
-
ICDPO: Effectively Borrowing Alignment Capability of Others via In-context Direct Preference Optimization 14 Feb 2024 · 1 repository · arXiv:2402.09320
-
Reinforcement Learning from Human Feedback with Active Queries 14 Feb 2024 · 0 repositories · arXiv:2402.09401
-
Active Preference Learning for Large Language Models 12 Feb 2024 · 0 repositories · arXiv:2402.08114
-
Refined Direct Preference Optimization with Synthetic Data for Behavioral Alignment of LLMs 12 Feb 2024 · 1 repository · arXiv:2402.08005
-
Relative Preference Optimization: Enhancing LLM Alignment through Contrasting Responses across Identical and Diverse Prompts 12 Feb 2024 · 1 repository · arXiv:2402.10958Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples)
-
Suppressing Pink Elephants with Direct Principle Feedback 12 Feb 2024 · 0 repositories · arXiv:2402.07896
-
V-STaR: Training Verifiers for Self-Taught Reasoners 9 Feb 2024 · 0 repositories · arXiv:2402.06457
-
Generalized Preference Optimization: A Unified Approach to Offline Alignment 8 Feb 2024 · 0 repositories · arXiv:2402.05749
-
Noise Contrastive Alignment of Language Models with Explicit Rewards 8 Feb 2024 · 3 repositories · arXiv:2402.05369Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
Direct Language Model Alignment from Online AI Feedback 7 Feb 2024 · 0 repositories · arXiv:2402.04792
-
BRAIn: Bayesian Reward-conditioned Amortized Inference for natural language generation from feedback 4 Feb 2024 · 0 repositories · arXiv:2402.02479
-
LiPO: Listwise Preference Optimization through Learning-to-Rank 2 Feb 2024 · 1 repository · arXiv:2402.01878Syntology 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 6 harvested samples)
-
Towards Efficient Exact Optimization of Language Model Alignment 1 Feb 2024 · 2 repositories · arXiv:2402.00856Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 5 harvested samples)
-
Self-Rewarding Language Models 18 Jan 2024 · 3 repositories · arXiv:2401.10020Syntology 6 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 2 violated, 3 with no contract checked; 1 where Syntology's instrument failed) · 5 unverified (of 11 harvested samples)
-
Aligning Large Language Models with Counterfactual DPO 17 Jan 2024 · 0 repositories · arXiv:2401.09566
-
A Mechanistic Understanding of Alignment Algorithms: A Case Study on DPO and Toxicity 3 Jan 2024 · 2 repositories · arXiv:2401.01967Syntology official (archive's flag): 12 ran · 12 ran (of which 0 constructed an object rather than computing a result; 12 with no instrument failure: 0 honoured, 0 violated, 12 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 14 harvested samples)
-
Preference as Reward, Maximum Preference Optimization with Importance Sampling 27 Dec 2023 · 0 repositories · arXiv:2312.16430
-
Some things are more CRINGE than others: Iterative Preference Optimization with the Pairwise Cringe Loss 27 Dec 2023 · 1 repository · arXiv:2312.16682Syntology 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 8 harvested samples) · 1 pointer-only (licence)
-
What Makes Good Data for Alignment? A Comprehensive Study of Automatic Data Selection in Instruction Tuning 25 Dec 2023 · 1 repository · arXiv:2312.15685Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 6 harvested samples) · 1 pointer-only (licence)
-
Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint 18 Dec 2023 · 3 repositories · arXiv:2312.11456
-
Policy Optimization in RLHF: The Impact of Out-of-preference Data 17 Dec 2023 · 1 repository · arXiv:2312.10584Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Silkie: Preference Distillation for Large Visual Language Models 17 Dec 2023 · 0 repositories · arXiv:2312.10665
-
ULMA: Unified Language Model Alignment with Human Demonstration and Point-wise Preference 5 Dec 2023 · 1 repository · arXiv:2312.02554Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
Using Human Feedback to Fine-tune Diffusion Models without Any Reward Model 22 Nov 2023 · 1 repository · arXiv:2311.13231
-
Diffusion Model Alignment Using Direct Preference Optimization 21 Nov 2023 · 2 repositories · arXiv:2311.12908Syntology 8 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 1 violated, 3 with no contract checked; 4 where Syntology's instrument failed) · 2 unverified (of 10 harvested samples) · 10 pointer-only (licence)
-
Direct Preference Optimization for Neural Machine Translation with Minimum Bayes Risk Decoding 14 Nov 2023 · 1 repository · arXiv:2311.08380Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 5 harvested samples) · 1 pointer-only (licence)
-
Black-Box Prompt Optimization: Aligning Large Language Models without Model Training 7 Nov 2023 · 1 repository · arXiv:2311.04155Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 7 harvested samples)