Browse State-of-the-Art › Paraphrase Generation
Paraphrase Generation
74 papers with code · 3 benchmarks · 17 datasets archive 2025-07-28
Paraphrase Generation involves transforming a natural language sentence to a new sentence, that has the same semantic meaning but a different syntactic or lexical surface form.
Description from the archive archive 2025-07-28.
Benchmarks archive 2025-07-28
3 leaderboard tables shown for this task, 3 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted.
| Dataset | Best model (first row in archive order) | Paper | Code | Syntology | Compare |
|---|---|---|---|---|---|
| Paralex (2 rows) | HRQ-VAE | Hierarchical Sketch Induction for Paraphrase Generation | code | Syntology ran 3 of 3 samples · 0 unverified | Compare |
| Quora Question Pairs (2 rows) | HRQ-VAE | Hierarchical Sketch Induction for Paraphrase Generation | code | Syntology ran 3 of 3 samples · 0 unverified | Compare |
| MSCOCO (1 row) | HRQ-VAE | Hierarchical Sketch Induction for Paraphrase Generation | code | Syntology ran 3 of 3 samples · 0 unverified | Compare |
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
17 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
1 subtask in the archive's task tree.
Parent tasks archive 2025-07-28
Most implemented papers archive 2025-07-28
30 shown of 74 papers with code (209 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
23 Oct 2023 3 repositories listedCurrent approaches in paraphrase generation and detection heavily rely on a single general similarity score, ignoring the intricate linguistic properties of language.
-
7 Oct 2022 3 repositories listedThe recent success of large language models for text generation poses a severe threat to academic integrity, as plagiarists can generate realistic paraphrases indistinguishable from original work.
-
22 May 2023 2 repositories listedThe emergence of generative large language models (LLMs) raises the question: what will be its impact on crowdsourcing?
-
15 May 2023 2 repositories listed Syntology ran 0 of 15 samples · 15 unverifiedDiffusion models have emerged as a powerful paradigm for generation, obtaining strong performance in various continuous domains.
-
13 Sep 2021 2 repositories listed Syntology ran 2 of 9 samples · 7 unverified · 2 pointer-only (licence)Phrase representations derived from BERT often do not exhibit complex phrasal compositionality, as the model relies instead on lexical similarity to determine semantic relatedness.
-
12 Oct 2020 2 repositories listed Syntology ran 9 of 19 samples · 10 unverifiedModern NLP defines the task of style transfer as modifying the style of a given sentence without appreciably changing its semantics, which implies that the outputs of style transfer systems should be paraphrases of…
-
18 May 2020 2 repositories listedIn these methods, syntactic-guidance is sourced from a separate exemplar sentence.
-
5 May 2020 2 repositories listedParaphrasing natural language sentences is a multifaceted process: it might involve replacing individual words or short phrases, local rearrangement of content, or high-level restructuring like topicalization or…
-
7 Jan 2020 2 repositories listed Syntology ran 0 of 12 samples · 12 unverifiedInspired by variational autoencoders with discrete latent structures, in this work, we propose a latent bag of words (BOW) model for paraphrase generation.
-
3 Jun 2018 2 repositories listed Syntology ran 4 of 17 samples · 13 unverifiedOne way to ensure this is by adding constraints for true paraphrase embeddings to be close and unrelated paraphrase candidate sentence embeddings to be far.
-
28 May 2025 1 repository listedThe PTD model surpasses automated metrics and provides a more reliable framework for evaluating paraphrase quality, advancing paraphrase-type research toward richer, user-aligned language generation and establishing a…
-
11 Feb 2025 1 repository listedThis paper presents ViSP, a high-quality Vietnamese dataset for sentence paraphrasing, consisting of 1.
-
10 Feb 2025 1 repository listed Syntology ran 0 of 1 samples · 1 unverifiedTo understand the complexity of sequence classification tasks, Hahn et al.
-
1 Nov 2024 1 repository listed Syntology ran 1 of 2 samples · 1 unverifiedAs Large Language Models (LLMs) are increasingly deployed in specialized domains with continuously evolving knowledge, the need for timely and precise knowledge injection has become essential.
-
23 Jul 2024 1 repository listedRecent surge in the accessibility of large language models (LLMs) to the general population can lead to untrackable use of such models for medical-related recommendations.
-
18 May 2024 1 repository listedTo address the inference gap, we introduce an optional action token as a placeholder that encourages the model to determine the appropriate action independently when users' intended actions are not provided.
-
13 Apr 2024 1 repository listedParaphrase generation strives to generate high-quality and diverse expressions of a given text, a domain where diffusion models excel.
-
7 Apr 2024 1 repository listedLanguage models, particularly generative models, are susceptible to hallucinations, generating outputs that contradict factual knowledge or the source text.
-
20 Oct 2023 1 repository listedFurthermore, for situations requiring multiple paraphrases for each source sentence, we design a Diverse Templates Search (DTS) algorithm, which can enhance the diversity between paraphrases without sacrificing quality.
-
28 Jul 2023 1 repository listedAfter feeding the input sentence into the encoder of paraphrase modeling, we generate the substitutes based on a novel decoding strategy that concentrates solely on the lexical variations of the complex word.
-
20 Jun 2023 1 repository listedMost existing text generation models follow the sequence-to-sequence paradigm.
-
26 May 2023 1 repository listedExisting fine-tuning methods for this task are costly as all the parameters of the model need to be updated during the training process.
-
26 May 2023 1 repository listedParaphrase generation is a long-standing task in natural language processing (NLP).
-
17 May 2023 1 repository listedWe present a large language model fine-tuned on a diverse collection of task-specific instructions for text editing (a total of 82K instructions).
-
23 Apr 2023 1 repository listedThough individual translated texts are often fluent and preserve meaning, at a large scale, translated texts have statistical tendencies which distinguish them from text originally written in the language…
-
23 Mar 2023 1 repository listedTo increase the robustness of AI-generated text detection to paraphrase attacks, we introduce a simple defense that relies on retrieving semantically-similar generations and must be maintained by a language model API…
-
27 Feb 2023 1 repository listedAugmenting the base neural model with a token-level symbolic datastore is a novel generation paradigm and has achieved promising results in machine translation (MT).
-
17 Jan 2023 1 repository listedIn this paper, we propose a syntactically robust training framework that enables models to be trained on a syntactic-abundant distribution based on diverse paraphrase generation.
-
Language as a Latent Sequence: deep latent variable models for semi-supervised paraphrase generation5 Jan 2023 1 repository listedTo leverage information from text pairs, we additionally introduce a novel supervised model we call dual directional learning (DDL), which is designed to integrate with our proposed VSAR model.
-
5 Sep 2022 1 repository listedWhile Transformers have had significant success in paragraph generation, they treat sentences as linear sequences of tokens and often neglect their hierarchical information.
Syntology lines on 7 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections