{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/open-subtitles-paraphrase-corpus-for-six","title":"Open Subtitles Paraphrase Corpus for Six Languages","arxiv_id":"1809.06142","date":"2018-09-17","proceeding":"LREC 2018 5","authors":["Mathias Creutz"],"abstract":"This paper accompanies the release of Opusparcus, a new paraphrase corpus for\nsix European languages: German, English, Finnish, French, Russian, and Swedish.\nThe corpus consists of paraphrases, that is, pairs of sentences in the same\nlanguage that mean approximately the same thing. The paraphrases are extracted\nfrom the OpenSubtitles2016 corpus, which contains subtitles from movies and TV\nshows. The informal and colloquial genre that occurs in subtitles makes such\ndata a very interesting language resource, for instance, from the perspective\nof computer assisted language learning. For each target language, the\nOpusparcus data have been partitioned into three types of data sets: training,\ndevelopment and test sets. The training sets are large, consisting of millions\nof sentence pairs, and have been compiled automatically, with the help of\nprobabilistic ranking functions. The development and test sets consist of\nsentence pairs that have been checked manually; each set contains approximately\n1000 sentence pairs that have been verified to be acceptable paraphrases by two\nannotators.","url_abs":"http://arxiv.org/abs/1809.06142v1","url_pdf":"http://arxiv.org/pdf/1809.06142v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"sentence","task_name":"Sentence"}],"methods":[],"datasets_introduced":[{"slug":"opusparcus","name":"Opusparcus","full_name":""}],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/1809.06142","atlas_url":"https://app.syntology.ai/?focus=1809.06142","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}