Papers › Deep Investigation of Cross-Language Plagiarism Detection Methods

Deep Investigation of Cross-Language Plagiarism Detection Methods

24 May 2017WS 2017 8arXiv:1705.08828archive 2025-07-28

Jeremy Ferrero, Laurent Besacier, Didier Schwab, Frederic Agnes

This paper is a deep investigation of cross-language plagiarism detection methods on a new recently introduced open dataset, which contains parallel and comparable collections of documents with multiple characteristics (different genres, languages and sizes of texts). We investigate cross-language plagiarism detection methods for 6 language pairs on 2 granularities of text units in order to draw robust conclusions on the best methods while deeply analyzing correlations across document styles and languages.

PaperPDFConference PDFCode

Code

FerreroJeremy/Cross-Language-Dataset officialmentioned in paper report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Results from the paper archive 2025-07-28

No leaderboard rows for this paper in the archive.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections