Papers › Universal Evasion Attacks on Summarization Scoring

Universal Evasion Attacks on Summarization Scoring

25 Oct 2022arXiv:2210.14260archive 2025-07-28

Wenchuan Mu, Kwan Hui Lim

The automatic scoring of summaries is important as it guides the development of summarizers. Scoring is also complex, as it involves multiple aspects such as fluency, grammar, and even textual entailment with the source text. However, summary scoring has not been considered a machine learning task to study its accuracy and robustness. In this study, we place automatic scoring in the context of regression machine learning tasks and perform evasion attacks to explore its robustness. Attack systems predict a non-summary string from each input, and these non-summary strings achieve competitive scores with good summarizers on the most popular metrics: ROUGE, METEOR, and BERTScore. Attack systems also "outperform" state-of-the-art summarization methods on ROUGE-1 and ROUGE-L, and score the second-highest on METEOR. Furthermore, a BERTScore backdoor is observed: a simple trigger can score higher than any automatic summarization method. The evasion attacks in this work indicate the low robustness of current scoring systems at the system level. We hope that our highlighting of these proposed attacks will facilitate the development of summary scores.

PaperPDFCode

Code

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Abstractive Text SummarizationDocument SummarizationNatural Language Inference

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Abstractive Text Summarization CNN / Daily Mail Scrambled code + broken (alter) ROUGE-1 48.18 #1 of 53 Archive leaderboard report
Abstractive Text Summarization CNN / Daily Mail Scrambled code + broken (alter) ROUGE-2 19.84 #1 of 53 Archive leaderboard report
Abstractive Text Summarization CNN / Daily Mail Scrambled code + broken (alter) ROUGE-L 45.35 #1 of 53 Archive leaderboard report
Abstractive Text Summarization CNN / Daily Mail Scrambled code + broken ROUGE-1 46.71 #5 of 53 Archive leaderboard report
Abstractive Text Summarization CNN / Daily Mail Scrambled code + broken ROUGE-2 20.39 #5 of 53 Archive leaderboard report
Abstractive Text Summarization CNN / Daily Mail Scrambled code + broken ROUGE-L 43.56 #5 of 53 Archive leaderboard report
Document Summarization CNN / Daily Mail Scrambled code + broken (alter) ROUGE-1 48.18 #1 of 26 Archive leaderboard report
Document Summarization CNN / Daily Mail Scrambled code + broken (alter) ROUGE-2 19.84 #1 of 26 Archive leaderboard report
Document Summarization CNN / Daily Mail Scrambled code + broken (alter) ROUGE-L 45.35 #1 of 26 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections