Home › Datasets › language › French
French datasets
archive 2025-07-28
190 datasets carry the language tag "French", ordered by the archive's paper count. Page 2 of 4: 48 shown of 190. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 50 language tags shown of 367, by dataset count; the full filter by modality, task and language is on /datasets
French datasets 49–96 of 190
The Multilingual Reuters Collection dataset comprises over 11,000 articles from six classes in five languages, i.e., English (E), French (F), German (G), Italian (I), and Spanish (S).
13 papers · 0 benchmarks
FLUE (French Language Understanding Evaluation)
FLUE is a French Language Understanding Evaluation benchmark.
12 papers · 0 benchmarks
MM-COVID (Multilingual and Multidimensional COVID-19 Fake News Data Repository)
MM-COVID is a dataset for fake news detection related to COVID-19.
12 papers · 0 benchmarks
MasakhaNEWS is a benchmark dataset for news topic classification covering 16 languages widely spoken in Africa.
12 papers · 0 benchmarks
MultiEURLEX is a multilingual dataset for topic classification of legal documents.
11 papers · 0 benchmarks
Synbols is a dataset generator designed for probing the behavior of learning algorithms.
11 papers · 0 benchmarks
VoxForge is an open speech dataset that was set up to collect transcribed speech for use with Free and Open Source Speech Recognition Engines (on Linux, Windows and Mac).
11 papers · 9 benchmarks
XFORMAL is a multilingual formal style transfer benchmark of multiple formal reformulations of informal text in Brazilian Portuguese, French, and Italian.
10 papers · 0 benchmarks
BIMCV-COVID19+ dataset is a large dataset with chest X-ray images CXR (CR, DX) and computed tomography (CT) imaging of COVID-19 patients along with their radiographic findings, pathologies, polymerase chain reaction (PCR), immunoglobulin G…
9 papers · 0 benchmarks
Global Voices is a multilingual dataset for evaluating cross-lingual summarization methods.
9 papers · 0 benchmarks
EUR-Lex-Sum is a dataset for cross-lingual summarization.
8 papers · 0 benchmarks
Emilia Dataset (An Extensive, Multilingual, and Diverse Speech Dataset for Large-Scale Speech Generation)
Recent advancements in speech generation models have been significantly driven by the use of large-scale training data.
8 papers · 0 benchmarks
RFUND (Revised FUNSD and XFUND)
RFUND is a relabeled version of FUNSD and XFUND datasets, tackling the following issues in their original annotations: 1.
8 papers · 0 benchmarks
RIMES (Reconnaissance & Indexation de données Manuscrites et de fac similÉS / Recognition & Indexing of handwritten documents & faxes)
The RIMES database (Reconnaissance et Indexation de données Manuscrites et de fac similÉS / Recognition and Indexing of handwritten documents and faxes) was created to evaluate automatic systems of recognition and indexing of handwritten…
8 papers · 0 benchmarks
Demetr is a diagnostic dataset with 31K English examples (translated from 10 source languages) for evaluating the sensitivity of MT evaluation metrics to 35 different linguistic perturbations spanning semantic, syntactic, and morphological…
7 papers · 0 benchmarks
BSARD (Belgian Statutory Article Retrieval Dataset)
The Belgian Statutory Article Retrieval Dataset (BSARD) is a French native corpus for studying statutory article retrieval.
6 papers · 1 benchmark
CLAMS (Cross-linguistic Analysis of Models on Syntax)
Targeted syntactic evaluation datasets in 5 languages: English, French, German, Russian, and Hebrew.
6 papers · 0 benchmarks
Concepticon (Concepticon. A Resource for the Linking of Concept Lists)
This resource, our Concepticon, links concept labels from different conceptlists to concept sets.
6 papers · 0 benchmarks
French TimeBank, a corpus for French annotated in ISO-TimeML.
6 papers · 1 benchmark
French Wikipedia is a dataset used for pretraining the CamemBERT French language model.
6 papers · 0 benchmarks
FrenchMedMCQA (FrenchMedMCQA: A French Multiple-Choice Question Answering Dataset for Medical domain)
This paper introduces FrenchMedMCQA, the first publicly available Multiple-Choice Question Answering (MCQA) dataset in French for medical domain.
6 papers · 1 benchmark
GTSinger (GTSinger: A Global Multi-Technique Singing Corpus with Realistic Music Scores for All Singing Tasks)
The scarcity of high-quality and multi-task singing datasets significantly hinders the development of diverse controllable and personalized singing tasks, as existing singing datasets suffer from low quality, limited diversity of languages…
6 papers · 0 benchmarks
We introduce HumanEval-XL, a massively multilingual code generation benchmark specifically crafted to address this deficiency.
6 papers · 0 benchmarks
The MEDIA French corpus is dedicated to semantic extraction from speech in a context of human/machine dialogues.
6 papers · 0 benchmarks
RELX is a benchmark dataset for cross-lingual relation classification in English, French, German, Spanish and Turkish.
6 papers · 0 benchmarks
We present a further analysis of visual modality incompleteness, benchmarking latest MMEA models on our proposed dataset MMEA-UMVM.
6 papers · 3 benchmarks
WikiNEuRal is a high-quality automatically-generated dataset for Multilingual Named Entity Recognition.
6 papers · 0 benchmarks
XQA is a data which consists of a total amount of 90k question-answer pairs in nine languages for cross-lingual open-domain question answering.
6 papers · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
5 papers · 0 benchmarks
MuMiN is a misinformation graph dataset containing rich social media data (tweets, replies, users, images, articles, hashtags), spanning 21 million tweets belonging to 26 thousand Twitter threads, each of which have been semantically…
5 papers · 0 benchmarks
DISRPT2019 (DISRPT2019 shared task on Discourse Unit Segmentation and Connective Detection)
The DISRPT 2019 workshop introduces the first iteration of a cross-formalism shared task on discourse unit segmentation.
4 papers · 0 benchmarks
DiS-ReX is a multilingual dataset for distantly supervised (DS) relation extraction (RE).
4 papers · 0 benchmarks
A new English-French test set for the evaluation of Machine Translation (MT) for informal, written bilingual dialogue.
4 papers · 1 benchmark
GeoCoV19 is a large-scale Twitter dataset containing more than 524 million multilingual tweets.
4 papers · 0 benchmarks
MultiSpider is a large multilingual text-to-SQL dataset which covers seven languages (English, German, French, Spanish, Japanese, Chinese, and Vietnamese).
4 papers · 0 benchmarks
MultiSubs (MultiSubs: A Large-scale Multimodal and Multilingual Dataset)
MultiSubs is a dataset of multilingual subtitles gathered from the OPUS OpenSubtitles dataset, which in turn was sourced from opensubtitles.org.
4 papers · 5 benchmarks
SRL is the task of extracting semantic predicate-argument structures from sentences.
4 papers · 0 benchmarks
XWINO is a multilingual collection of Winograd Schemas in six languages that can be used for evaluation of cross-lingual commonsense reasoning capabilities.
4 papers · 1 benchmark
DAWT (Densely Annotated Wikipedia Texts)
The DAWT dataset consists of Densely Annotated Wikipedia Texts across multiple languages.
3 papers · 0 benchmarks
DISRPT2021 (DISRPT2021 shared task on Discourse Unit Segmentation, Connective Detection and Discourse Relation Classification)
The DISRPT 2021 shared task, co-located with CODI 2021 at EMNLP, introduces the second iteration of a cross-formalism shared task on discourse unit segmentation and connective detection, as well as the first iteration of a cross-formalism…
3 papers · 0 benchmarks
DPB-5L is a Multilingual KG dataset containing 5 KGs in English, French, Japanese, Greek, and Spanish.
3 papers · 1 benchmark
This repository contains gzipped files containing more than 2 million tokens (words) from answers submitted by more than 6,000 students over the course of their first 30 days of using Duolingo.
3 papers · 0 benchmarks
This is a gzipped CSV file containing the 13 million Duolingo student learning traces used in experiments by Settles & Meeder (2016).
3 papers · 0 benchmarks
FreSaDa is a French satire dataset for cross-domain satire detection, which is composed of 11,570 articles from the news domain.
3 papers · 0 benchmarks
The Jamendo Corpus is a voice detection dataset consisting of 93 songs with Creative Commons license from the Jamendo free music sharing website.
3 papers · 0 benchmarks
LLeQA (Long-form Legal Question Answering)
LLeQA is a French native dataset for studying information retrieval and long-form question answering in the legal domain.
3 papers · 0 benchmarks
MRS (Multilingual Reply Suggestion)
MRS, a multilingual reply suggestion dataset with ten languages.
3 papers · 0 benchmarks
Multilingual automatic speech recognition (ASR) in the medical domain serves as a foundational task for various downstream applications such as speech translation, spoken language understanding, and voice-activated assistants.
3 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.