Papers › LM-CPPF: Paraphrasing-Guided Data Augmentation for Contrastive Prompt-Based Few-Shot...
LM-CPPF: Paraphrasing-Guided Data Augmentation for Contrastive Prompt-Based Few-Shot Fine-Tuning
Amirhossein Abaskohi, Sascha Rothe, Yadollah Yaghoobzadeh
In recent years, there has been significant progress in developing pre-trained language models for NLP. However, these models often struggle when fine-tuned on small datasets. To address this issue, researchers have proposed various adaptation approaches. Prompt-based tuning is arguably the most common way, especially for larger models. Previous research shows that adding contrastive learning to prompt-based fine-tuning is effective as it helps the model generate embeddings that are more distinguishable between classes, and it can also be more sample-efficient as the model learns from positive and negative examples simultaneously. One of the most important components of contrastive learning is data augmentation, but unlike computer vision, effective data augmentation for NLP is still challenging. This paper proposes LM-CPPF, Contrastive Paraphrasing-guided Prompt-based Fine-tuning of Language Models, which leverages prompt-based few-shot paraphrasing using generative language models, especially large language models such as GPT-3 and OPT-175B, for data augmentation. Our experiments on multiple text classification benchmarks show that this augmentation method outperforms other methods, such as easy data augmentation, back translation, and multiple templates.
In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.
Code
Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.
Code Syntology ran Syntology
Not run by Syntology. Nothing on this page verifies that the listed code works.
Tasks
Results from the paper archive 2025-07-28
| Task | Dataset | Model | Metric | Value | Rank at snapshot | Leaderboard | Report |
|---|---|---|---|---|---|---|---|
| Linguistic Acceptability | CoLA | LM-CPPF RoBERTa-base | Accuracy | 14.1% | #42 of 43 | Archive leaderboard | report |
| Natural Language Inference | MultiNLI | LM-CPPF RoBERTa-base | Accuracy | 68.4 | #65 of 67 | Archive leaderboard | report |
| Natural Language Inference | QNLI | LM-CPPF RoBERTa-base | Accuracy | 70.2% | #41 of 43 | Archive leaderboard | report |
| Sentiment Analysis | CR | LM-CPPF RoBERTa-base | Accuracy | 93.3 | #2 of 9 | Archive leaderboard | report |
| Sentiment Analysis | SST-2 Binary classification | LM-CPPF RoBERTa-base | Accuracy | 93.2 | #42 of 87 | Archive leaderboard | report |
| Sentiment Analysis | SST-5 Fine-grained classification | LM-CPPF RoBERTa-base | Accuracy | 54.9 | #7 of 31 | Archive leaderboard | report |
Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.
Methods
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections