Methods › Natural Language Processing › Text Data Augmentation › SynthesizRR

Synthesize by Retrieval and Refinement

SynthesizRR

1 paper tagged archive 2025-07-28

Introduced by Abhishek Divekar et al. in SynthesizRR: Generating Diverse Datasets with Retrieval Augmentation

archive 2025-07-28 Description, source and code snippet are the archive's method entry.

It is often desirable to distill the capabilities of large language models (LLMs) into smaller student models due to compute and memory constraints. One way to do this for classification tasks is via dataset synthesis. Prior approaches to synthesis use few-shot prompting, which relies on the LLM's parametric knowledge to generate usable examples. However, this leads to issues of repetition, bias towards popular entities, and stylistic differences from human text. We propose Synthesize by Retrieval and Refinement (SynthesizRR), which uses retrieval augmentation to introduce variety into the dataset synthesis process: as retrieved passages vary, the LLM is seeded with different content to generate its examples. We find that SynthesizRR greatly improves lexical and semantic diversity, similarity to human-written text, and distillation performance,

PaperSource

Papers archive 2025-07-28

1 shown of 1, newest first. Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code; it is per sample and not a correctness claim.

Tasks archive 2025-07-28

10 tasks the archive attaches to papers tagged with this method, by distinct papers. A task without a page in the catalog is plain text.

TaskPapers
Bias Detection1
Diversity1
Humor Detection1
Information Retrieval1
Language Modelling1
Product Categorization1
Retrieval1
Sentiment Analysis1
Synthetic Data Generation1
Topic Classification1

Usage over time archive 2025-07-28

Papers per year tagged with SynthesizRR: 2024 to 2024, peak 1 1 0 2024: 1 paper 2024
Papers per year the archive tags with this method, by the paper's archive date (1 dated). Bars are counts, not a trend claim.

Components: the archive holds no method-to-method composition, so PwC's Components table cannot be rebuilt; the Papers list carries no Results column for the same reason (the archive does not join its leaderboard rows to method tags).

Categories archive 2025-07-28

Text Data Augmentation

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections