Home › Datasets › language › French

French datasets

archive 2025-07-28

190 datasets carry the language tag "French", ordered by the archive's paper count. Page 4 of 4: 46 shown of 190. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 50 language tags shown of 367, by dataset count; the full filter by modality, task and language is on /datasets

French datasets 145–190 of 190

LEMONADE is a large, expert-annotated dataset for event extraction from news articles in 20 languages: English, Spanish, Arabic, French, Italian, Russian, German, Turkish, Burmese, Indonesian, Ukrainian, Korean, Portuguese, Dutch, Somali,…
1 paper · 0 benchmarks
LSFB Datasets (French Belgian Sign Language Datasets)
Sign Language Datasets for French Belgian Sign Language This dataset is built upon the work of Belgian linguists from the University of Namur.
1 paper · 0 benchmarks
The M-AILABS Speech Dataset is the first large dataset that we are providing free-of-charge, freely usable as training data for speech recognition and speech synthesis.
1 paper · 1 benchmark
M-Phasis (A Feature-Based Corpus of Hate Online)
A corpus of 9k German and French user comments collected from migration-related news articles.
1 paper · 0 benchmarks
M3LS (Multi-Lingual Multi-Modal Summarization Dataset)
Significant developments in techniques such as encoder-decoder models have enabled us to represent information comprising multiple modalities.
1 paper · 0 benchmarks
MAKED (MultiModal MultiLingual Summarization and Keyword Extraction Dataset)
Keyword extraction is an integral task for many downstream problems like clustering, recommendation, search and classification.
1 paper · 0 benchmarks
MSNER (Multilingual Spoken Named Entity Recognition)
This dataset contains named entities annotations for European Parliament recordings in Dutch, French, German and Spanish.
1 paper · 0 benchmarks
MVALUE (Multilingual human VALUE dataset)
Multilingual human VALUE(MVALUE) is a multilingual dataset covering 7 concepts of human values: morality, deontology, utilitarianism, fairness, truthfulness, toxicity and harmfulness, each concept subset of it includes positive and…
1 paper · 0 benchmarks
Mediapi-RGB is a bilingual corpus of French Sign Language (LSF) and written French in the form of subtitled videos, accompanied by complementary data (various representations, segmentation, vocabulary, etc.).
1 paper · 1 benchmark
Mega-COV is a billion-scale dataset from Twitter for studying COVID-19.
1 paper · 0 benchmarks
This data adds textual meta-infomation data to two existing corpora for cross language information retrieval: BoostCLIR, and the Large Scale CLIR Dataset (wiki-clir).
1 paper · 0 benchmarks
Mint (Multilingual Intimacy analysis)
Mint is a new Multilingual intimacy analysis dataset covering 13,384 tweets in 10 languages including English, French, Spanish, Italian, Portuguese, Korean, Dutch, Chinese, Hindi, and Arabic.
1 paper · 0 benchmarks
This dataset contains dialogue lines from the games Knights of the Old Republic 1 & 2 and Neverwinter Nights 1.
1 paper · 0 benchmarks
OLD French Coronavirus Screening Data (OLD (France) Données de laboratoires pour le dépistage : Indicateurs sur les mutations SI-DEP)
The RT-PCR screening tests used and the results of which are reported in SI-DEP made it possible to suspect the presence of the worrisome variant (VOC) Alpha (20I/501Y.V1) and indistinctly from the VOC Beta (20H/501Y.
1 paper · 0 benchmarks
This dataset contains orthographic samples of words in 19 languages (ar, br, de, en, eno, ent, eo, es, fi, fr, fro, it, ko, nl, pt, ru, sh, tr, zh).
1 paper · 0 benchmarks
PolyNews is a multilingual dataset containing news titles in 77 languages and 19 scripts.
1 paper · 0 benchmarks
PolyNews is a multilingual parallel dataset containing news titles 833 language pairs, spanning in 64 languages and 17 scripts.
1 paper · 0 benchmarks
RTE3-FR dataset is the French translation of the Textual Entailment English dataset used in the RTE-3 Challenge (https://nlp.stanford.edu/RTE3-pilot).
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
> The StatCan Dialogue Dataset: Retrieving Data Tables through Conversations with Genuine Intents > > Xing Han Lu, Siva Reddy, Harm de Vries > > EACL 2023 | | | | | | | :--: | :--: | :--: | :--: | :--: | | Code | Huggingface | Request on…
1 paper · 1 benchmark
TQBA++ (Tiny QA Benchmark++)
Ultra-lightweight, multilingual QA eval dataset for rapid testing LLMs.
1 paper · 0 benchmarks
This is the reference headset microphone variant of the VibraVox dataset.
1 paper · 2 benchmarks
This is the in-ear rigid earpiece-embedded microphone variant of the VibraVox dataset.
1 paper · 3 benchmarks
This is the in-ear comply foam-embedded microphone variant of the VibraVox dataset.
1 paper · 3 benchmarks
This is the throat microphone (laryngophone) variant of the VibraVox dataset.
1 paper · 3 benchmarks
WEATHub is a dataset containing 24 languages.
1 paper · 0 benchmarks
WMT 2014 Medical (WMT 2014 Medical Translation Task)
The Medical Translation Task of WMT 2014 addresses the problem of domain-specific and genre-specific machine translation.
1 paper · 0 benchmarks
WMT 2015 News (WMT 2015 News Translation Task)
News translation is a recurring WMT task.
1 paper · 0 benchmarks
WMT 2016 Biomedical (WMT 2016 Biomedical Translation Task)
The Biomedical Translation Shared Task was first introduced at the First Conference of Machine Translation.
1 paper · 0 benchmarks
Vision Language Models (VLMs) often struggle with culture-specific knowledge, particularly in languages other than English and in underrepresented cultural contexts.
1 paper · 0 benchmarks
blbooks (The British Library Books)
This dataset consists of books digitised by the British Library in partnership with Microsoft.
1 paper · 0 benchmarks
Diderot’s Encyclopédie is a reference work from XVIIIth century in Europe that aimed at collecting the knowledge of its era.
1 paper · 0 benchmarks
inaGVAD (InaGVAD : a Challenging French TV and Radio Corpus annotated for Voice Activity Detection and Speaker Gender Segmentation)
InaGVAD is a Voice Activity Detection (VAD) and Speaker Gender Segmentation (SGS) dataset designed for representing the acoustic diversity of French TV and Radio programs.
1 paper · 0 benchmarks
This dataset links all the entries describing named entities of Petit Larousse illustré, a French dictionary published in 1905, to wikidata identifiers.
1 paper · 0 benchmarks
mDRT (Multilingual Diagnostic Rhyme Test)
We present a multilingual test set for conducting speech intelligibility tests in the form of diagnostic rhyme tests.
1 paper · 0 benchmarks
AlexMI (Alex Motor Imagery dataset)
Alex Motor Imagery dataset.
0 papers · 0 benchmarks
Medical report generation (MRG), which aims to automatically generate a textual description of a specific medical image (e.g., a chest X-ray), has recently received increasing research interest.
0 papers · 0 benchmarks
CamNuvem Dataset (CamNuvem: A Robbery Dataset for Video Anomaly Detection)
This dataset focuses only on the robbery category, presenting a new weakly labelled dataset that contains 486 new real–world robbery surveillance videos acquired from public sources.
0 papers · 0 benchmarks
Multi-Spectral Leaf Segmentation (Multi-Spectral Leaf Segmentation For Crop/Weed Identification)
This dataset were acquired with the Airphen (Hyphen, Avignon, France) six-band multi-spectral camera configured using the 450/570/675/710/730/850 nm bands with a 10 nm FWHM.
0 papers · 0 benchmarks
POPCORN (POPCORN: Fictional and Synthetic Intelligence Reports for Named Entity Recognition and Relation Extraction Tasks)
POPCORN is a French dataset consisting of 400 validation texts and 400 training texts, all written and annotated manually.
0 papers · 0 benchmarks
THVD (Talking Head Video Dataset)
About We provide a comprehensive talking-head video dataset with over 50,000 videos, totaling more than 500+ hours of footage and featuring 20,841 unique identities from around the world.
0 papers · 0 benchmarks
The WASABI Song Corpus is a large corpus of songs enriched with metadata extracted from music databases on the Web, and resulting from the processing of song lyrics and from audio analysis.
0 papers · 0 benchmarks
depression interview dataset (depression interview dataset with 1.6 million clinical trail data)
contain the clinical trial dataset
0 papers · 0 benchmarks
--- annotationscreators: - no-annotation language: - fr languagecreators: - found license: - cc-by-4.0 multilinguality: - monolingual prettyname: French Legal Cases Dataset sizecategories: - n>1M sourcedatasets: - la-mousse/INCA-17-01-2025…
0 papers · 0 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.