Papers › Red Teaming for Large Language Models At Scale: Tackling Hallucinations on Mathematics Tasks

Red Teaming for Large Language Models At Scale: Tackling Hallucinations on Mathematics Tasks

30 Dec 2023arXiv:2401.00290archive 2025-07-28

Aleksander Buszydlik, Karol Dobiczek, Michał Teodor Okoń, Konrad Skublicki, Philip Lippmann, Jie Yang

We consider the problem of red teaming LLMs on elementary calculations and algebraic tasks to evaluate how various prompting techniques affect the quality of outputs. We present a framework to procedurally generate numerical questions and puzzles, and compare the results with and without the application of several red teaming techniques. Our findings suggest that even though structured reasoning and providing worked-out examples slow down the deterioration of the quality of answers, the gpt-3.5-turbo and gpt-4 models are not well suited for elementary calculations and reasoning tasks, also when being red teamed.

PaperPDFCode

Code

redteamingforllms/redteamingforllms officialmentioned in paper report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Red Teaming

Results from the paper archive 2025-07-28

No leaderboard rows for this paper in the archive.

Methods

CoT PromptingGPT-4Multi-Head AttentionTransformer

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections