Papers › Step-wise Policy for Rare-tool Knowledge (SPaRK): Offline RL that Drives Diverse Tool...

Step-wise Policy for Rare-tool Knowledge (SPaRK): Offline RL that Drives Diverse Tool Use in LLMs

15 Jul 2025arXiv:2507.11371archive 2025-07-28

Gabriel Bo, Koa Chang, Justin Gu

We present Step-wise Policy for Rare-tool Knowledge (SPaRK), a novel reinforcement learning framework that teaches large language models to explore diverse tool usage patterns beyond conventional high-temperature sampling. Building on recent advances in step-wise reinforcement learning, we introduce a dual-objective reward system that simultaneously optimizes for answer quality and tool diversity, training a Llama-3.1 8B model through offline PPO on synthetically generated trajectories from the MMLU-Pro dataset. Our approach uniquely employs a rarity-first exploitation strategy where a GPT-4o judge scores candidate actions across eight distinct tools plus chain-of-thought reasoning, with the policy favoring less-frequently used but still viable tools to encourage systematic exploration. Empirical results demonstrate that SPaRK achieves competitive performance across 14 MMLU-Pro categories while exhibiting significantly higher entropy in tool selection compared to both baseline and supervised fine-tuning approaches, suggesting that algorithmic exploration through explicit tool diversity can enhance reasoning capabilities without sacrificing accuracy.

PaperPDFCode

Code

gabrielkmbo/explore-rl officialpytorch report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

DiversityMMLUOffline RLReinforcement Learningreinforcement-learning

Datasets

Introduced by this paper, per the archive.

SWiRL MMLU-Pro

Results from the paper archive 2025-07-28

No leaderboard rows for this paper in the archive.

Methods

PPO

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections