Browse State-of-the-Art › Text Augmentation
Text Augmentation
41 papers with code · 0 benchmarks · 0 datasets archive 2025-07-28
You can read these blog posts to get an overview of the approaches.
Description from the archive archive 2025-07-28.
Benchmarks archive 2025-07-28
No benchmark for this task in the archive.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
No dataset record in the archive lists this task.
Subtasks archive 2025-07-28
No subtask under this task in the archive's task tree.
Parent tasks archive 2025-07-28
Most implemented papers archive 2025-07-28
30 shown of 41 papers with code (97 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
31 Jan 2019 16 repositories listed Syntology ran 2 of 13 samples · 11 unverified · 1 pointer-only (licence)We present EDA: easy data augmentation techniques for boosting performance on text classification tasks.
-
22 Mar 2019 2 repositories listedNeural NLP systems achieve high scores in the presence of sizable training dataset.
-
16 May 2018 2 repositories listed Syntology ran 0 of 1 samples · 1 unverifiedWe stochastically replace words with other words that are predicted by a bi-directional language model at the word positions.
-
29 Apr 2025 1 repository listedNone of the three augmentation techniques consistently improved overall performance for classifying hazards and products.
-
4 Mar 2025 1 repository listed Syntology ran 3 of 3 samples · 0 unverifiedVision-Language Models (VLMs) excel in integrating visual and textual information for vision-centric tasks, but their handling of inconsistencies between modalities is underexplored.
-
31 Jan 2025 1 repository listedTo achieve this, we introduce an adapter module and mitigate the noise issue in the dense CLIP feature distillation process through a self-cross-training strategy.
-
29 Jan 2025 1 repository listedAdditionally, we identified potential solutions to these limitations from the literature to enhance the efficacy of data augmentation practices using multimodal LLMs.
-
13 Dec 2024 1 repository listedZero-shot action recognition (ZSAR) requires collaborative multi-modal spatiotemporal understanding.
-
14 Oct 2024 1 repository listedWe evaluate this on in-distribution and out-of-distribution classifier performance.
-
7 Apr 2024 1 repository listed Syntology ran 4 of 9 samples · 5 unverified · 9 pointer-only (licence)Acknowledging this limitation, our objective is to devise a framework capable of concurrently augmenting medical image and text data.
-
23 Mar 2024 1 repository listedTo address these issues, we propose an encoder-decoder data augmentation (EDDA) framework.
-
6 Feb 2024 1 repository listedHowever, most DG methods assume access to abundant source data in the target label space, a requirement that proves overly stringent for numerous real-world applications, where acquiring the same label space as the…
-
27 Jan 2024 1 repository listedLeveraging large models, these data augmentation techniques have outperformed traditional approaches.
-
12 Jan 2024 1 repository listed Syntology ran 3 of 3 samples · 0 unverified · 3 pointer-only (licence)The latest generative large language models (LLMs) have found their application in data augmentation tasks, where small numbers of text samples are LLM-paraphrased and then used to fine-tune downstream models.
-
7 Dec 2023 1 repository listedIn the era of artificial intelligence, data is gold but costly to annotate.
-
6 Dec 2023 1 repository listedThrough additional training, we explore embedding specialized scientific knowledge into the Llama 2 Large Language Model (LLM).
-
30 Nov 2023 1 repository listedThis paper introduces a multilingual dataset of COVID-19 vaccine misinformation, consisting of annotated tweets from three middle-income countries: Brazil, Indonesia, and Nigeria.
-
28 Nov 2023 1 repository listed Syntology ran 3 of 3 samples · 0 unverified · 3 pointer-only (licence)To address this limitation, we adopt a causal generative perspective for multimodal data and propose contrastive learning with data augmentation to disentangle content features from the original representations.
-
19 Oct 2023 1 repository listedIn many real-world scenarios (e.
-
9 Sep 2023 1 repository listedOne of the current state-of-the-art text augmentation techniques is easy data augmentation (EDA), which augments the training data by injecting and replacing synonyms and randomly permuting sentences.
-
15 Aug 2023 1 repository listedStory visualization (SV) is a challenging text-to-image generation task for the difficulty of not only rendering visual details from the text descriptions but also encoding a long-term context across multiple sentences.
-
24 Feb 2023 1 repository listedDespite recent advancements in Machine Learning, many tasks still involve working in low-data regimes which can make solving natural language problems difficult.
-
12 Dec 2022 1 repository listedHowever, existing data augmentation techniques in natural language understanding (NLU) may not fully capture the complexity of natural language variations, and they can be challenging to apply to large datasets.
-
6 Oct 2022 1 repository listedOur experimental results on three classification tasks and nine public datasets show that BootAug addresses the performance drop problem and outperforms state-of-the-art text augmentation methods.
-
1 Oct 2022 1 repository listedNotably, our proposed Zemi_(LARGE) outperforms T0-3B by 16% on all seven evaluation tasks while being 3.
-
22 Sep 2022 1 repository listedCovid-19 has spread across the world and several vaccines have been developed to counter its surge.
-
12 Sep 2022 1 repository listedThis paper proposes a simple yet effective interpolation-based data augmentation approach termed DoubleMix, to improve the robustness of models in text classification.
-
4 Sep 2022 1 repository listedDifferent words may play different roles in text classification, which inspires us to strategically select the proper roles for text augmentation.
-
28 May 2022 1 repository listedAs computers have become efficient at understanding visual information and transforming it into a written representation, research interest in tasks like automatic image captioning has seen a significant leap over the…
-
4 Mar 2022 1 repository listedIn addition, our model can extract visual information as suggested by the text prompt, e.
Syntology lines on 6 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections