Methods › Natural Language Processing › Text Data Augmentation › DART

DART

32 papers tagged archive 2025-07-28

archive 2025-07-28 Description, source and code snippet are the archive's method entry.

🎯 DART-Math

Difficulty-Aware Rejection Tuning for Mathematical Problem-Solving

📝 Paper@arXiv | 🤗 Datasets&Models@HF | 🐱 Code@GitHub

🐦 Thread@X(Twitter) | 🐶 中文博客@知乎 | 📊 Leaderboard@PapersWithCode | 📑 BibTeX

Datasets: DART-Math

DART-Math datasets are the state-of-the-art and data-efficient open-source instruction tuning datasets for mathematical reasoning.

DART-Math-Hard contains \~585k mathematical QA pair samples constructed by applying DARS-Prop2Diff to the query set from MATH and GSK8K training sets, achieves SOTA on many challenging mathematical reasoning benchmarks. It introduces a deliberate bias towards hard queries, opposite to vanilla rejection sampling.

Performance produced by DART-Math-Hard is usually but not necessarily slightly better (\~1% absolutely) than DART-Math-Uniform, which contains \~591k samples constructed by applying DARS-Uniform.

Comparison between Mathematical Instruction Tuning Datasets

Most of previous datasets are constructed with ChatGPT, and many of them are not open-source, especially for ones of the best performance.

Math SFT Dataset # of Samples MATH GSM8K College Synthesis Agent(s) Open-Source
WizardMath 96k 32.3 80.4 23.1 GPT-4 ✗
MetaMathQA 395k 29.8 76.5 19.3 GPT-3.5 ✓
MMIQC 2294k 37.4 75.4 28.5 GPT-4+GPT-3.5+Human ✓
Orca-Math 200k -- -- -- GPT-4 ✓
Xwin-Math-V1.1 1440k 45.5 84.9 27.6 GPT-4 ✗
KPMath-Plus 1576k 46.8 82.1 -– GPT-4 ✗
MathScaleQA 2021k 35.2 74.8 21.8 GPT-3.5+Human ✗
DART-Math-Uniform 591k 43.5 82.6 26.9 DeepSeekMath-7B-RL ✓
DART-Math-Hard 585k 45.5 81.1 29.4 DeepSeekMath-7B-RL ✓

MATH and GSM8K are in-domain, while College(Math) is out-of-domain. Performance here are of models fine-tuned from Mistral-7B, except for Xwin-Math-V1.1 based on Llama2-7B. Bold/Italic means the best/second best score here.

Dataset Construction: DARS - Difficulty-Aware Rejection Sampling

Previous works usually synthesize data from proprietary models to augment existing datasets, followed by instruction tuning to achieve top-tier results. However, our analysis of these datasets reveals severe biases towards easy queries, with frequent failures to generate any correct response for the most challenging queries.

Motivated by the observation above, we propose to Difficulty-Aware Rejection Sampling (DARS), to collect more responses for more difficult queries. Specifically, we introduce two strategies to increase the number of correct responses for difficult queries:

1) Uniform, which involves sampling responses for each query until each query accumulates kᵤ correct responses, where kᵤ is a preset hyperparameter determined by the desired size of the synthetic dataset; 2) Prop2Diff, where we continue sampling responses until the number of correct responses for each query is proportional to its difficulty score. The most challenging queries will receive kₚ responses and kp is a hyperparameter. This method introduces a deliberate bias in the opposite direction to vanilla rejection sampling, towards more difficult queries, inspired by previous works that demonstrate difficult samples can be more effective to enhance model capabilities (Sorscher et al., 2022; Liu et al., 2024b).

See Figure 1 (Right) for examples of DART-Math-Uniform by DARS-Uniform and DART-Math-Hard by DARS-Prop2Diff.

Citation

If you find our data, model or code useful for your work, please kindly cite our paper:

@article{tong2024dartmath,
  title={DART-Math: Difficulty-Aware Rejection Tuning for Mathematical Problem-Solving},
  author={Yuxuan Tong and Xiwen Zhang and Rui Wang and Ruidong Wu and Junxian He},
  year={2024},
  eprint={2407.13690},
  archivePrefix={arXiv},
  primaryClass={cs.CL},
  url={https://arxiv.org/abs/2407.13690},
}
See Code · hkust-nlp/dart-math

Papers archive 2025-07-28

30 shown of 32, newest first. Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code; it is per sample and not a correctness claim.

Tasks archive 2025-07-28

20 shown of 55 tasks the archive attaches to papers tagged with this method, by distinct papers. A task without a page in the catalog is plain text.

TaskPapers
Text Generation6
Data-to-Text Generation4
Image Generation2
Large Language Model2
Object Detection2
Pseudo Label2
Red Teaming2
object-detection2
Active Learning1
Adversarial Robustness1
Anomaly Detection1
Arithmetic Reasoning1
Bayesian Inference1
Computational Efficiency1
Continual Learning1
Decoder1
Denoising1
Domain Adaptation1
Edge-computing1
Image Classification1

Usage over time archive 2025-07-28

Papers per year tagged with DART: 2023 to 2025, peak 24 24 0 2023: 2 papers 2023 2024: 24 papers 2024 2025: 6 papers 2025
Papers per year the archive tags with this method, by the paper's archive date (32 dated). Bars are counts, not a trend claim.

Components: the archive holds no method-to-method composition, so PwC's Components table cannot be rebuilt; the Papers list carries no Results column for the same reason (the archive does not join its leaderboard rows to method tags).

Categories archive 2025-07-28

Text Data Augmentation

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections