Home › Datasets › language › French
French datasets
archive 2025-07-28
190 datasets carry the language tag "French", ordered by the archive's paper count. Page 4 of 4: 46 shown of 190. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 50 language tags shown of 367, by dataset count; the full filter by modality, task and language is on /datasets
French datasets 145–190 of 190
LEMONADE is a large, expert-annotated dataset for event extraction from news articles in 20 languages: English, Spanish, Arabic, French, Italian, Russian, German, Turkish, Burmese, Indonesian, Ukrainian, Korean, Portuguese, Dutch, Somali,…
1 paper · 0 benchmarks
Sign Language Datasets for French Belgian Sign Language This dataset is built upon the work of Belgian linguists from the University of Namur.
1 paper · 0 benchmarks
The M-AILABS Speech Dataset is the first large dataset that we are providing free-of-charge, freely usable as training data for speech recognition and speech synthesis.
1 paper · 1 benchmark
M-Phasis (A Feature-Based Corpus of Hate Online)
A corpus of 9k German and French user comments collected from migration-related news articles.
1 paper · 0 benchmarks
M3LS (Multi-Lingual Multi-Modal Summarization Dataset)
Significant developments in techniques such as encoder-decoder models have enabled us to represent information comprising multiple modalities.
1 paper · 0 benchmarks
MAKED (MultiModal MultiLingual Summarization and Keyword Extraction Dataset)
Keyword extraction is an integral task for many downstream problems like clustering, recommendation, search and classification.
1 paper · 0 benchmarks
MSNER (Multilingual Spoken Named Entity Recognition)
This dataset contains named entities annotations for European Parliament recordings in Dutch, French, German and Spanish.
1 paper · 0 benchmarks
MVALUE (Multilingual human VALUE dataset)
Multilingual human VALUE(MVALUE) is a multilingual dataset covering 7 concepts of human values: morality, deontology, utilitarianism, fairness, truthfulness, toxicity and harmfulness, each concept subset of it includes positive and…
1 paper · 0 benchmarks
Mediapi-RGB is a bilingual corpus of French Sign Language (LSF) and written French in the form of subtitled videos, accompanied by complementary data (various representations, segmentation, vocabulary, etc.).
1 paper · 1 benchmark
Mega-COV is a billion-scale dataset from Twitter for studying COVID-19.
1 paper · 0 benchmarks
This data adds textual meta-infomation data to two existing corpora for cross language information retrieval: BoostCLIR, and the Large Scale CLIR Dataset (wiki-clir).
1 paper · 0 benchmarks
Mint (Multilingual Intimacy analysis)
Mint is a new Multilingual intimacy analysis dataset covering 13,384 tweets in 10 languages including English, French, Spanish, Italian, Portuguese, Korean, Dutch, Chinese, Hindi, and Arabic.
1 paper · 0 benchmarks
This dataset contains dialogue lines from the games Knights of the Old Republic 1 & 2 and Neverwinter Nights 1.
1 paper · 0 benchmarks
The RT-PCR screening tests used and the results of which are reported in SI-DEP made it possible to suspect the presence of the worrisome variant (VOC) Alpha (20I/501Y.V1) and indistinctly from the VOC Beta (20H/501Y.
1 paper · 0 benchmarks
This dataset contains orthographic samples of words in 19 languages (ar, br, de, en, eno, ent, eo, es, fi, fr, fro, it, ko, nl, pt, ru, sh, tr, zh).
1 paper · 0 benchmarks
PolyNews is a multilingual dataset containing news titles in 77 languages and 19 scripts.
1 paper · 0 benchmarks
PolyNews is a multilingual parallel dataset containing news titles 833 language pairs, spanning in 64 languages and 17 scripts.
1 paper · 0 benchmarks
RTE3-FR dataset is the French translation of the Textual Entailment English dataset used in the RTE-3 Challenge (https://nlp.stanford.edu/RTE3-pilot).
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
> The StatCan Dialogue Dataset: Retrieving Data Tables through Conversations with Genuine Intents > > Xing Han Lu, Siva Reddy, Harm de Vries > > EACL 2023 | | | | | | | :--: | :--: | :--: | :--: | :--: | | Code | Huggingface | Request on…
1 paper · 1 benchmark
Ultra-lightweight, multilingual QA eval dataset for rapid testing LLMs.
1 paper · 0 benchmarks
This is the forehead accelerometer variant of the VibraVox dataset.
1 paper · 3 benchmarks
This is the reference headset microphone variant of the VibraVox dataset.
1 paper · 2 benchmarks
This is the in-ear rigid earpiece-embedded microphone variant of the VibraVox dataset.
1 paper · 3 benchmarks
This is the in-ear comply foam-embedded microphone variant of the VibraVox dataset.
1 paper · 3 benchmarks
This is the temple vibration pickup variant of the VibraVox dataset.
1 paper · 3 benchmarks
This is the throat microphone (laryngophone) variant of the VibraVox dataset.
1 paper · 3 benchmarks
WEATHub is a dataset containing 24 languages.
1 paper · 0 benchmarks
The Medical Translation Task of WMT 2014 addresses the problem of domain-specific and genre-specific machine translation.
1 paper · 0 benchmarks
News translation is a recurring WMT task.
1 paper · 0 benchmarks
The Biomedical Translation Shared Task was first introduced at the First Conference of Machine Translation.
1 paper · 0 benchmarks
Vision Language Models (VLMs) often struggle with culture-specific knowledge, particularly in languages other than English and in underrepresented cultural contexts.
1 paper · 0 benchmarks
blbooks (The British Library Books)
This dataset consists of books digitised by the British Library in partnership with Microsoft.
1 paper · 0 benchmarks
Diderot’s Encyclopédie is a reference work from XVIIIth century in Europe that aimed at collecting the knowledge of its era.
1 paper · 0 benchmarks
inaGVAD (InaGVAD : a Challenging French TV and Radio Corpus annotated for Voice Activity Detection and Speaker Gender Segmentation)
InaGVAD is a Voice Activity Detection (VAD) and Speaker Gender Segmentation (SGS) dataset designed for representing the acoustic diversity of French TV and Radio programs.
1 paper · 0 benchmarks
This dataset links all the entries describing named entities of Petit Larousse illustré, a French dictionary published in 1905, to wikidata identifiers.
1 paper · 0 benchmarks
mDRT (Multilingual Diagnostic Rhyme Test)
We present a multilingual test set for conducting speech intelligibility tests in the form of diagnostic rhyme tests.
1 paper · 0 benchmarks
AlexMI (Alex Motor Imagery dataset)
Alex Motor Imagery dataset.
0 papers · 0 benchmarks
Medical report generation (MRG), which aims to automatically generate a textual description of a specific medical image (e.g., a chest X-ray), has recently received increasing research interest.
0 papers · 0 benchmarks
This dataset focuses only on the robbery category, presenting a new weakly labelled dataset that contains 486 new real–world robbery surveillance videos acquired from public sources.
0 papers · 0 benchmarks
This dataset were acquired with the Airphen (Hyphen, Avignon, France) six-band multi-spectral camera configured using the 450/570/675/710/730/850 nm bands with a 10 nm FWHM.
0 papers · 0 benchmarks
POPCORN (POPCORN: Fictional and Synthetic Intelligence Reports for Named Entity Recognition and Relation Extraction Tasks)
POPCORN is a French dataset consisting of 400 validation texts and 400 training texts, all written and annotated manually.
0 papers · 0 benchmarks
THVD (Talking Head Video Dataset)
About We provide a comprehensive talking-head video dataset with over 50,000 videos, totaling more than 500+ hours of footage and featuring 20,841 unique identities from around the world.
0 papers · 0 benchmarks
The WASABI Song Corpus is a large corpus of songs enriched with metadata extracted from music databases on the Web, and resulting from the processing of song lyrics and from audio analysis.
0 papers · 0 benchmarks
contain the clinical trial dataset
0 papers · 0 benchmarks
--- annotationscreators: - no-annotation language: - fr languagecreators: - found license: - cc-by-4.0 multilinguality: - monolingual prettyname: French Legal Cases Dataset sizecategories: - n>1M sourcedatasets: - la-mousse/INCA-17-01-2025…
0 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.