{"url":"/task/text-augmentation","name":"Text Augmentation","slug":"text-augmentation","description_markdown":"You can read these blog posts to get an overview of the approaches.  \r\n\r\n- [**A Visual Survey of Data Augmentation in NLP**](https://amitness.com/2020/05/data-augmentation-for-nlp/)","categories":[{"name":"Computer Vision","url":"/area/computer-vision"},{"name":"Methodology","url":"/area/methodology"},{"name":"Natural Language Processing","url":"/area/natural-language-processing"}],"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","slug_source":"derived"},"counts":{"papers_tagged":97,"papers_with_code":41,"benchmarks":0,"benchmark_tables_in_archive":0,"benchmark_tables_shown":0,"benchmark_tables_withheld_as_spam":0,"benchmark_definition":"a leaderboard table with at least one row; benchmark_tables_shown also counts the zero-row tables; benchmark_tables_in_archive adds the tables withheld as spam","datasets":0,"subtasks":0,"parent_tasks":1},"benchmarks":[],"datasets":[],"subtasks":[],"parent_tasks":[{"url":"/task/data-augmentation","name":"Data Augmentation"}],"papers":{"order":"repositories listed in the archive (desc), then date (desc); the archive holds no stars","population":"papers tagged with this task that list at least one repository in the archive","shown":30,"of":41,"tagged_in_all":97,"items":[{"url":"/paper/eda-easy-data-augmentation-techniques-for","title":"EDA: Easy Data Augmentation Techniques for Boosting Performance on Text Classification Tasks","date":"2019-01-31","arxiv_id":"1901.11196","repositories_listed":16,"syntology":{"n":13,"n_ran":2,"n_unverified":11,"n_pointer_only":1}},{"url":"/paper/data-augmentation-via-dependency-tree-1","title":"Data Augmentation via Dependency Tree Morphing for Low-Resource Languages","date":"2019-03-22","arxiv_id":"1903.09460","repositories_listed":2,"syntology":null},{"url":"/paper/contextual-augmentation-data-augmentation-by","title":"Contextual Augmentation: Data Augmentation by Words with Paradigmatic Relations","date":"2018-05-16","arxiv_id":"1805.06201","repositories_listed":2,"syntology":{"n":1,"n_ran":0,"n_unverified":1,"n_pointer_only":0}},{"url":"/paper/brightcookies-at-semeval-2025-task-9","title":"BrightCookies at SemEval-2025 Task 9: Exploring Data Augmentation for Food Hazard Classification","date":"2025-04-29","arxiv_id":"2504.20703","repositories_listed":1,"syntology":null},{"url":"/paper/words-or-vision-do-vision-language-models","title":"Words or Vision: Do Vision-Language Models Have Blind Faith in Text?","date":"2025-03-04","arxiv_id":"2503.02199","repositories_listed":1,"syntology":{"n":3,"n_ran":3,"n_unverified":0,"n_pointer_only":0}},{"url":"/paper/laser-efficient-language-guided-segmentation","title":"Laser: Efficient Language-Guided Segmentation in Neural Radiance Fields","date":"2025-01-31","arxiv_id":"2501.19084","repositories_listed":1,"syntology":null},{"url":"/paper/image-text-and-speech-data-augmentation-using","title":"Image, Text, and Speech Data Augmentation using Multimodal LLMs for Deep Learning: A Survey","date":"2025-01-29","arxiv_id":"2501.18648","repositories_listed":1,"syntology":null},{"url":"/paper/building-a-multi-modal-spatiotemporal-expert","title":"Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP","date":"2024-12-13","arxiv_id":"2412.09895","repositories_listed":1,"syntology":null},{"url":"/paper/use-random-selection-for-now-investigation-of","title":"Use Random Selection for Now: Investigation of Few-Shot Selection Strategies in LLM-based Text Augmentation for Classification","date":"2024-10-14","arxiv_id":"2410.10756","repositories_listed":1,"syntology":null},{"url":"/paper/pairaug-what-can-augmented-image-text-pairs","title":"PairAug: What Can Augmented Image-Text Pairs Do for Radiology?","date":"2024-04-07","arxiv_id":"2404.04960","repositories_listed":1,"syntology":{"n":9,"n_ran":4,"n_unverified":5,"n_pointer_only":9}},{"url":"/paper/edda-a-encoder-decoder-data-augmentation","title":"EDDA: A Encoder-Decoder Data Augmentation Framework for Zero-Shot Stance Detection","date":"2024-03-23","arxiv_id":"2403.15715","repositories_listed":1,"syntology":null},{"url":"/paper/a-data-centric-approach-for-unsupervised","title":"Multimodal Unsupervised Domain Generalization by Retrieving Across the Modality Gap","date":"2024-02-06","arxiv_id":"2402.04416","repositories_listed":1,"syntology":null},{"url":"/paper/a-survey-on-data-augmentation-in-large-model","title":"A Survey on Data Augmentation in Large Model Era","date":"2024-01-27","arxiv_id":"2401.15422","repositories_listed":1,"syntology":null},{"url":"/paper/effects-of-diversity-incentives-on-sample","title":"Effects of diversity incentives on sample diversity and downstream model performance in LLM-based text augmentation","date":"2024-01-12","arxiv_id":"2401.06643","repositories_listed":1,"syntology":{"n":3,"n_ran":3,"n_unverified":0,"n_pointer_only":3}},{"url":"/paper/from-big-to-small-without-losing-it-all-text","title":"From Big to Small Without Losing It All: Text Augmentation with ChatGPT for Efficient Sentiment Analysis","date":"2023-12-07","arxiv_id":"2312.04720","repositories_listed":1,"syntology":null},{"url":"/paper/teaching-specific-scientific-knowledge-into","title":"Teaching Specific Scientific Knowledge into Large Language Models through Additional Training","date":"2023-12-06","arxiv_id":"2312.03360","repositories_listed":1,"syntology":null},{"url":"/paper/covid-19-vaccine-misinformation-in-middle","title":"COVID-19 Vaccine Misinformation in Middle Income Countries","date":"2023-11-30","arxiv_id":"2311.18195","repositories_listed":1,"syntology":null},{"url":"/paper/clap-contrastive-learning-with-augmented","title":"CLAP: Isolating Content from Style through Contrastive Learning with Augmented Prompts","date":"2023-11-28","arxiv_id":"2311.16445","repositories_listed":1,"syntology":{"n":3,"n_ran":3,"n_unverified":0,"n_pointer_only":3}},{"url":"/paper/pretraining-language-models-with-text","title":"Pretraining Language Models with Text-Attributed Heterogeneous Graphs","date":"2023-10-19","arxiv_id":"2310.12580","repositories_listed":1,"syntology":null},{"url":"/paper/distributional-data-augmentation-methods-for","title":"Distributional Data Augmentation Methods for Low Resource Language","date":"2023-09-09","arxiv_id":"2309.04862","repositories_listed":1,"syntology":null},{"url":"/paper/story-visualization-by-online-text","title":"Story Visualization by Online Text Augmentation with Context Memory","date":"2023-08-15","arxiv_id":"2308.07575","repositories_listed":1,"syntology":null},{"url":"/paper/sta-self-controlled-text-augmentation-for","title":"STA: Self-controlled Text Augmentation for Improving Text Classifications","date":"2023-02-24","arxiv_id":"2302.12784","repositories_listed":1,"syntology":null},{"url":"/paper/rpn-a-word-vector-level-data-augmentation","title":"RPN: A Word Vector Level Data Augmentation Algorithm in Deep Learning for Language Understanding","date":"2022-12-12","arxiv_id":"2212.05961","repositories_listed":1,"syntology":null},{"url":"/paper/augmentor-or-filter-reconsider-the-role-of","title":"BootAug: Boosting Text Augmentation via Hybrid Instance Filtering Framework","date":"2022-10-06","arxiv_id":"2210.02941","repositories_listed":1,"syntology":null},{"url":"/paper/zemi-learning-zero-shot-semi-parametric","title":"Zemi: Learning Zero-Shot Semi-Parametric Language Models from Multiple Tasks","date":"2022-10-01","arxiv_id":"2210.00185","repositories_listed":1,"syntology":null},{"url":"/paper/adaptation-of-domain-specific-transformer","title":"Adaptation of domain-specific transformer models with text oversampling for sentiment analysis of social media posts on Covid-19 vaccines","date":"2022-09-22","arxiv_id":"2209.10966","repositories_listed":1,"syntology":null},{"url":"/paper/doublemix-simple-interpolation-based-data","title":"DoubleMix: Simple Interpolation-Based Data Augmentation for Text Classification","date":"2022-09-12","arxiv_id":"2209.05297","repositories_listed":1,"syntology":null},{"url":"/paper/selective-text-augmentation-with-word-roles","title":"Selective Text Augmentation with Word Roles for Low-Resource Text Classification","date":"2022-09-04","arxiv_id":"2209.01560","repositories_listed":1,"syntology":null},{"url":"/paper/ban-cap-a-multi-purpose-english-bangla-image","title":"BAN-Cap: A Multi-Purpose English-Bangla Image Descriptions Dataset","date":"2022-05-28","arxiv_id":"2205.14462","repositories_listed":1,"syntology":null},{"url":"/paper/show-me-what-and-tell-me-how-video-synthesis","title":"Show Me What and Tell Me How: Video Synthesis via Multimodal Conditioning","date":"2022-03-04","arxiv_id":"2203.02573","repositories_listed":1,"syntology":null}],"syntology_records":6,"syntology_note":"a paper without a record is not a recorded non-run: it may lack an arXiv id or simply be absent from the graph layer"},"description_links":{"kept":0,"unwrapped_to_text":0,"bare_urls_linked":0,"relative_images_dropped":0,"rule":"internal links are kept only when the target slug exists in the catalog"},"syntology":{"read_at":"2026-09-24T18:15:14+00:00","claim":"Per-sample execution status on synthesized fixtures ('ran N of M samples'); not a correctness claim and not a ranking signal.","status_vocabulary":{"ran_honours":"ran, honoured the contract we drafted","ran_violates":"ran, violated the contract we drafted","ran_draft_wrong":"ran; our contract draft was wrong, not the code","ran_fixture":"ran; our fixture could not drive it","ran":"ran on a synthesized input","unverified":"unverified (harvested, no recorded run)"}},"not_shown":{"libraries":"the archive has no per-task library table","trend_sparklines":"the Trend column of the benchmarks table was a rendered image; it is not in the archive","social_and_latest_sorts":"stars and social signals are not in the archive"}}