Papers › Prompting for explanations improves Adversarial NLI. Is this true? {Yes} it is {true}...
Prompting for explanations improves Adversarial NLI. Is this true? {Yes} it is {true} because {it weakens superficial cues}
Pride Kavumba, Ana Brassard, Benjamin Heinzerling, Kentaro Inui
Explanation prompts ask language models to not only assign a particular label to a giveninput, such as true, entailment, or contradiction in the case of natural language inference but also to generate a free-text explanation that supports this label. For example: “This is label because explanation.” While this type of prompt was originally introduced with the aim of improving model interpretability, we showhere that explanation prompts also improve robustness to adversarial perturbations in naturallanguage inference benchmarks. Compared to prompting for labels only, explanation prompting consistently yields stronger performance on adversarial benchmarks, outperforming the state of the art on Adversarial Natural Language Inference, Counterfactually-Augmented Natural Language Inference, and SNLI-Hard datasets. We argue that the increase in robustness is due to the fact that prompting for explanations weakens superficial cues. Specifically, single tokens that are highly predictive of the correct answer in the label-only setting become uninformative when the model also has to generate explanations.
Code
No code repository is listed for this paper in the archive or in Syntology's graph.
Code Syntology ran Syntology
Not run by Syntology. Nothing on this page verifies that the listed code works.
Tasks
Results from the paper archive 2025-07-28
| Task | Dataset | Model | Metric | Value | Rank at snapshot | Leaderboard | Report |
|---|---|---|---|---|---|---|---|
| Natural Language Inference | ANLI test | T5-3B (explanation prompting) | A1 | 81.8 | #1 of 25 | Archive leaderboard | report |
| Natural Language Inference | ANLI test | T5-3B (explanation prompting) | A2 | 72.5 | #1 of 25 | Archive leaderboard | report |
| Natural Language Inference | ANLI test | T5-3B (explanation prompting) | A3 | 74.8 | #1 of 25 | Archive leaderboard | report |
| Natural Language Inference | ANLI test | T0-11B (explanation prompting) | A1 | 75.6 | #2 of 25 | Archive leaderboard | report |
| Natural Language Inference | ANLI test | T0-11B (explanation prompting) | A2 | 60.6 | #2 of 25 | Archive leaderboard | report |
| Natural Language Inference | ANLI test | T0-11B (explanation prompting) | A3 | 59.9 | #2 of 25 | Archive leaderboard | report |
Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections