Home › Datasets › language › Czech

Czech datasets

archive 2025-07-28

41 datasets carry the language tag "Czech", ordered by the archive's paper count. Page 1 of 1: 41 shown of 41. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 50 language tags shown of 367, by dataset count; the full filter by modality, task and language is on /datasets

Czech datasets 1–41 of 41

The Universal Dependencies (UD) project seeks to develop cross-linguistically consistent treebank annotation of morphology and syntax for multiple languages.
520 papers · 5 benchmarks
Common Voice is an audio dataset that consists of a unique MP3 and corresponding text file.
449 papers · 143 benchmarks
MuST-C currently represents the largest publicly available multilingual corpus (one-to-many) for speech translation.
216 papers · 2 benchmarks
WMT 2016 is a collection of datasets used in shared tasks of the First Conference on Machine Translation.
178 papers · 16 benchmarks
MLQA (MultiLingual Question Answering)
MLQA (MultiLingual Question Answering) is a benchmark dataset for evaluating cross-lingual question answering performance.
167 papers · 1 benchmark
Europarl (European Parliament Proceedings Parallel Corpus)
A corpus of parallel text in 21 European languages from the proceedings of the European Parliament.
128 papers · 1 benchmark
This corpus comprises of monolingual data for 100+ languages and also includes data for romanized languages.
115 papers · 0 benchmarks
WikiANN (PAN-X)
WikiANN, also known as PAN-X, is a multilingual named entity recognition dataset.
72 papers · 4 benchmarks
OSCAR or Open Super-large Crawled ALMAnaCH coRpus is a huge multilingual corpus obtained by language classification and filtering of the Common Crawl corpus using the goclassy architecture.
64 papers · 0 benchmarks
WikiLingua includes ~770k article and summary pairs in 18 languages from WikiHow.
52 papers · 1 benchmark
WMT 2020 is a collection of datasets used in shared tasks of the Fifth Conference on Machine Translation.
33 papers · 0 benchmarks
XM 3600 (Crossmodal 3600)
Research in massively multilingual image captioning has been severely hampered by a lack of high-quality evaluation datasets.
28 papers · 0 benchmarks
The MULTEXT-East resources are a multilingual dataset for language engineering research and development.
25 papers · 0 benchmarks
OASST1 (OpenAssistant Conversations Dataset)
license: apache-2.0 tags: human-feedback sizecategories: 100K Languages with under 1000 messages Vietnamese: 952 Basque: 947 Polish: 886 Hungarian: 811 Arabic: 666 Dutch: 628 Swedish: 512 Turkish: 454 Finnish: 386 Czech: 372 Danish: 358…
25 papers · 0 benchmarks
WMT 2016 News (WMT 2016 News Translation Task)
News translation is a recurring WMT task.
24 papers · 0 benchmarks
AKCES-GEC is a new dataset on grammatical error correction for Czech.
11 papers · 0 benchmarks
MultiEURLEX is a multilingual dataset for topic classification of legal documents.
11 papers · 0 benchmarks
EUR-Lex-Sum is a dataset for cross-lingual summarization.
8 papers · 0 benchmarks
WMT 2018 News (WMT 2018 News Translation Task)
News translation is a recurring WMT task.
8 papers · 0 benchmarks
Demetr is a diagnostic dataset with 31K English examples (translated from 10 source languages) for evaluating the sensitivity of MT evaluation metrics to 35 different linguistic perturbations spanning semantic, syntactic, and morphological…
7 papers · 0 benchmarks
MuMiN is a misinformation graph dataset containing rich social media data (tweets, replies, users, images, articles, hashtags), spanning 21 million tweets belonging to 26 thousand Twitter threads, each of which have been semantically…
5 papers · 0 benchmarks
DaReCzech (Dataset for text relevance ranking in Czech)
DareCzech DaReCzech is a dataset for text relevance ranking in Czech.
4 papers · 1 benchmark
GeoCoV19 is a large-scale Twitter dataset containing more than 524 million multilingual tweets.
4 papers · 0 benchmarks
Czech restaurant information is a dataset for NLG in task-oriented spoken dialogue systems with Czech as the target language.
2 papers · 1 benchmark
Industry Biscuit (Cookie) dataset (Industrial style dataset for the anomaly detection)
The Industrial Biscuits (Cookie) dataset is our internal dataset designed for the anomaly detection task, which captures Tarallini biscuits.
2 papers · 0 benchmarks
The dataset contains training and evaluation data for 12 languages: - Vietnamese - Romanian - Latvian - Czech - Polish - Slovak - Irish - Hungarian - French - Turkish - Spanish - Croatian For each language, one training, one development…
2 papers · 12 benchmarks
COSTRA 1.0 is a dataset of complex sentence transformations.
1 paper · 0 benchmarks
Czech subjectivity dataset of 10k manually annotated subjective and objective sentences from movie reviews and descriptions.
1 paper · 1 benchmark
Dataset Description EUROPA is a dataset designed for training and evaluating multilingual keyphrase generation models in the legal domain.
1 paper · 0 benchmarks
GLAMI-1M (A Multilingual Image-Text Fashion Dataset)
We introduce GLAMI-1M: the largest multilingual image-text classification dataset and benchmark.
1 paper · 1 benchmark
HALvest is a textual dataset comprising 17 billion tokens in 56 languages and 13 domains.
1 paper · 0 benchmarks
PolyNews is a multilingual dataset containing news titles in 77 languages and 19 scripts.
1 paper · 0 benchmarks
PolyNews is a multilingual parallel dataset containing news titles 833 language pairs, spanning in 64 languages and 17 scripts.
1 paper · 0 benchmarks
SumeCzech-NER contains named entity annotations of SumeCzech 1.0, a Czech news-based summarization dataset.
1 paper · 0 benchmarks
Verifee is a dataset of news articles with fine-grained trustworthiness annotations.
1 paper · 0 benchmarks
WMT 2014 Medical (WMT 2014 Medical Translation Task)
The Medical Translation Task of WMT 2014 addresses the problem of domain-specific and genre-specific machine translation.
1 paper · 0 benchmarks
WMT 2015 News (WMT 2015 News Translation Task)
News translation is a recurring WMT task.
1 paper · 0 benchmarks
WMT 2016 IT (WMT 2016 IT Translation Task)
The IT Translation Task is a shared task introduced in the First Conference on Machine Translation.
1 paper · 0 benchmarks
Vision Language Models (VLMs) often struggle with culture-specific knowledge, particularly in languages other than English and in underrepresented cultural contexts.
1 paper · 0 benchmarks
The data originate from the journalistic domain in the Czech language.
0 papers · 0 benchmarks
THVD (Talking Head Video Dataset)
About We provide a comprehensive talking-head video dataset with over 50,000 videos, totaling more than 500+ hours of footage and featuring 20,841 unique identities from around the world.
0 papers · 0 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.