Home › Datasets › language › French

French datasets

archive 2025-07-28

190 datasets carry the language tag "French", ordered by the archive's paper count. Page 3 of 4: 48 shown of 190. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 50 language tags shown of 367, by dataset count; the full filter by modality, task and language is on /datasets

French datasets 97–144 of 190

QALD-9-Plus Dataset Description QALD-9-Plus is the dataset for Knowledge Graph Question Answering (KGQA) based on well-known QALD-9.
3 papers · 1 benchmark
The SmartSpeaker benchmark tests the performance of reacting to music player commands in English as well as in French.
3 papers · 1 benchmark
A vast amount of information in the biomedical domain is available as natural language free text.
3 papers · 0 benchmarks
Travel (Tour & Travels Customer Churn Prediction)
A Tour & Travels Company Wants To Predict Whether A Customer Will Churn Or Not Based On Indicators Given Below.
3 papers · 1 benchmark
TyDiP (A Dataset for Politeness Classification in Nine Typologically Diverse Languages)
A Dataset for Politeness Classification in Nine Typologically Diverse Languages (TyDiP) is a dataset containing three-way politeness annotations for 500 examples in each language, totaling 4.5K examples.
3 papers · 0 benchmarks
CLSE (Corpus of Linguistically Significant Entities)
2 papers · 0 benchmarks
DACCORD is a new dataset dedicated to the task of automatically detecting contradictions between sentences in French.
2 papers · 0 benchmarks
FIJO (French Insurance Job Offer dataset)
This dataset was collected as part of the multidisciplinary project Femmes face aux défis de la transformation numérique : une étude de cas dans le secteur des assurances (Women Facing the Challenges of Digital Transformation: A Case Study…
2 papers · 0 benchmarks
GATITOS (Google's Additional Translations Into Tail-languages: Often Short)
The GATITOS (Google's Additional Translations Into Tail-languages: Often Short) dataset is a high-quality, multi-way parallel dataset of tokens and short phrases, intended for training and improving machine translation models.
2 papers · 0 benchmarks
Timely and effective response to humanitarian crises requires quick and accurate analysis of large amounts of text data, a process that can highly benefit from expert-assisted NLP systems trained on validated and annotated data in the…
2 papers · 0 benchmarks
Jam-ALT (JamALT: A Formatting-Aware Lyrics Transcription Benchmark)
JamALT is a revision of the JamendoLyrics dataset (80 songs in 4 languages), adapted for use as an automatic lyrics transcription (ALT) benchmark.
2 papers · 5 benchmarks
Kaleidoscope (Kaleidoscope: In-language Exams for Massively Multilingual Vision Evaluation)
The evaluation of vision-language models (VLMs) has mainly relied on English-language benchmarks, leaving significant gaps in both multilingual and multicultural coverage.
2 papers · 0 benchmarks
MGSM8KInstruct, the multilingual math reasoning instruction dataset, encompassing ten distinct languages, thus addressing the issue of training data scarcity in multilingual math reasoning.
2 papers · 0 benchmarks
Morph Call is a suite of 46 probing tasks for four Indo-European languages that fall under different morphology: Russian, French, English, and German.
2 papers · 0 benchmarks
MuCo-VQA consist of large-scale (3.7M) multilingual and code-mixed VQA datasets in multiple languages: Hindi (hi), Bengali (bn), Spanish (es), German (de), French (fr) and code-mixed language pairs: en-hi, en-bn, en-fr, en-de and en-es.
2 papers · 0 benchmarks
MultiTACRED is a multilingual version of the large-scale TAC Relation Extraction Dataset.
2 papers · 0 benchmarks
The dataset contains training and evaluation data for 12 languages: - Vietnamese - Romanian - Latvian - Czech - Polish - Slovak - Irish - Hungarian - French - Turkish - Spanish - Croatian For each language, one training, one development…
2 papers · 12 benchmarks
A human-revised dataset for seven languages that allows for the evaluation of multilingual RE systems.
2 papers · 0 benchmarks
SIMARA (SIMARA: a database for key-value information extraction from full-page handwritten documents)
Description We propose a new database for information extraction from historical handwritten documents.
2 papers · 2 benchmarks
Tilde MODEL Corpus (Tilde Multilingual Open Data for European Languages)
Tilde MODEL Corpus is a multilingual corpora for European languages – particularly focused on the smaller languages.
2 papers · 0 benchmarks
WikiCaps is a large-scale multilingual but non-parallel data set for multimodal machine translation and retrieval.
2 papers · 0 benchmarks
WikiTableSet (Wikipedia Table Image Dataset)
WikiTableSet is a large publicly available image-based table recognition dataset in three languages built from Wikipedia.
2 papers · 0 benchmarks
X-WikiRE is a new, large-scale multilingual relation extraction dataset in which relation extraction is framed as a problem of reading comprehension to allow for generalization to unseen relations.
2 papers · 0 benchmarks
ANTILLES (ANTILLES: An Open French Linguistically Enriched Part-of-Speech Corpus)
ANTILLES is a part-of-speech tagging corpus based on UDFrench-GSD which was originally created in 2015 and is based on the universal dependency treebank v2.0.
1 paper · 1 benchmark
AQL-22 (Archive Query Log)
The Archive Query Log (AQL) is a previously unused, comprehensive query log collected at the Internet Archive over the last 25 years.
1 paper · 0 benchmarks
This paper analyses two hitherto unstudied sites sharing state-backed disinformation, Reliable Recent News (rrn.world) and WarOnFakes (waronfakes.com), which publish content in Arabic, Chinese, English, French, German, and Spanish.
1 paper · 0 benchmarks
BAH (Behavioural Ambivalence/Hesitancy)
Recognizing complex emotions linked to ambivalence and hesitancy (A/H) can play a critical role in the personalization and effectiveness of digital behaviour change interventions.
1 paper · 0 benchmarks
Belfort (The Belfort dataset: Handwritten Text Recognition from Crowdsourced Annotations)
The Belfort dataset This dataset includes minutes of Belfort municipal council drawn up between 1790 and 1946.
1 paper · 1 benchmark
CURE (A dataset for Clinical Understanding & Retrieval Evaluation)
CURE is a retrieval dataset with a monolingual and two cross-lingual conditions, with splits spanning ten medical domains.
1 paper · 0 benchmarks
French sentences are sourced from Tatoeba repository and then translated into Congolese Swahili.
1 paper · 0 benchmarks
Deep Indices (multi-spectral leaf/vegetation segmentation)
This dataset inclue multi-spectral acquisition of vegetation for the conception of new DeepIndices.
1 paper · 1 benchmark
These are the test and training data used for experiments presented in BioNLP 2017.
1 paper · 0 benchmarks
Dataset Description EUROPA is a dataset designed for training and evaluating multilingual keyphrase generation models in the legal domain.
1 paper · 0 benchmarks
The EVI dataset is a challenging, multilingual spoken-dialogue dataset with 5,506 dialogues in English, Polish, and French.
1 paper · 3 benchmarks
The ConcoDisco Corpus is an English-French parallel corpus with discourse relations (DRs) and discourse connectives (DCs) annotations.
1 paper · 0 benchmarks
EventEA is an event-centric entity alignment dataset, harvested from EventKG, DBpedia and Wikidata.
1 paper · 0 benchmarks
The FairTranslate Dataset includes 2,418 sentence pairs, each centered around an occupation, designed to assess gender expression and translation in English-French contexts.
1 paper · 0 benchmarks
Fallout New Vegas Dialog is a multilingual sentiment annotated dialog dataset from Fallout New Vegas.
1 paper · 0 benchmarks
The Food Recall Incidents dataset consists of 7,546 short texts (from 5 to 360 characters each), which are the titles of food recall announcements (therefore referred to as title), crawled from 24 public food safety authority websites by…
1 paper · 0 benchmarks
FreCDo (French cross-domain)
FreCDo is a corpus for French dialect identification comprising 413,522 French text samples collected from public news websites in Belgium, Canada, France and Switzerland.
1 paper · 0 benchmarks
Composed of judgments from the French Court of cassation and their corresponding summaries.
1 paper · 0 benchmarks
We collated subcorpora each between 50,000 and 70,000 words, containing samples of national dialects of French across different countries: Algeria, Democratic Republic of Congo, France, Ivory Coast, Morocco and Senegal.
1 paper · 0 benchmarks
GQNLI-FR is a manually translated French version of the GQNLI challenge dataset, originally written in English.
1 paper · 0 benchmarks
GeoEDdA (A Gold Standard Dataset for Geo-semantic Annotation of Diderot & d’Alembert’s Encyclopédie)
Dataset Description - Authors: Ludovic Moncla, Katherine McDonough and Denis Vigier in the framework of the GEODE project.
1 paper · 0 benchmarks
HALvest is a textual dataset comprising 17 billion tokens in 56 languages and 13 domains.
1 paper · 0 benchmarks
HALvest-Geometric is a subset of HALvest: an academic citation network with 238,397 disambiguated authors and 18,662,037 scholarly papers.
1 paper · 0 benchmarks
- Revision: v1.0.0-full-20210527a - DOI: 10.5281/zenodo.4817662 - Authors: J.
1 paper · 0 benchmarks
Dataset Summary INCLUDE is a comprehensive knowledge- and reasoning-centric benchmark across 44 languages that evaluates multilingual LLMs for performance in the actual language environments where they would be deployed.
1 paper · 0 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.