Papers › MUSS: Multilingual Unsupervised Sentence Simplification by Mining Paraphrases

MUSS: Multilingual Unsupervised Sentence Simplification by Mining Paraphrases

1 May 2020LREC 2022 6arXiv:2005.00352archive 2025-07-28

Louis Martin, Angela Fan, Éric de la Clergerie, Antoine Bordes, Benoît Sagot

Progress in sentence simplification has been hindered by a lack of labeled parallel simplification data, particularly in languages other than English. We introduce MUSS, a Multilingual Unsupervised Sentence Simplification system that does not require labeled simplification data. MUSS uses a novel approach to sentence simplification that trains strong models using sentence-level paraphrase data instead of proper simplification data. These models leverage unsupervised pretraining and controllable generation mechanisms to flexibly adjust attributes such as length and lexical complexity at inference time. We further present a method to mine such paraphrase data in any language from Common Crawl using semantic sentence embeddings, thus removing the need for labeled data. We evaluate our approach on English, French, and Spanish simplification benchmarks and closely match or outperform the previous best supervised results, despite not using any labeled simplification data. We push the state of the art further by incorporating labeled simplification data.

PaperPDFConference PDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

facebookresearch/muss officialmentioned on GitHub report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Parallel Corpus MiningSentenceText Simplification

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Text Simplification ASSET MUSS (BART+ACCESS Supervised) BLEU 72.98 #2 of 12 Archive leaderboard report
Text Simplification ASSET MUSS (BART+ACCESS Supervised) FKGL 6.05 #2 of 12 Archive leaderboard report
Text Simplification ASSET MUSS (BART+ACCESS Supervised) SARI (EASSE>=0.2.1) 44.15 #2 of 12 Archive leaderboard report
Text Simplification ASSET MUSS (BART+ACCESS Unsupervised) FKGL 8.23 #5 of 12 Archive leaderboard report
Text Simplification ASSET MUSS (BART+ACCESS Unsupervised) SARI (EASSE>=0.2.1) 42.65 #5 of 12 Archive leaderboard report
Text Simplification TurkCorpus MUSS (BART+ACCESS Supervised) BLEU 78.17 #2 of 25 Archive leaderboard report
Text Simplification TurkCorpus MUSS (BART+ACCESS Supervised) FKGL 7.60 #2 of 25 Archive leaderboard report
Text Simplification TurkCorpus MUSS (BART+ACCESS Supervised) SARI (EASSE>=0.2.1) 42.53 #2 of 25 Archive leaderboard report
Text Simplification TurkCorpus MUSS (BART+ACCESS Unsupervised) FKGL 8.79 #6 of 25 Archive leaderboard report
Text Simplification TurkCorpus MUSS (BART+ACCESS Unsupervised) SARI (EASSE>=0.2.1) 40.85 #6 of 25 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections