Home › Datasets › language › Italian
Italian datasets
archive 2025-07-28
98 datasets carry the language tag "Italian", ordered by the archive's paper count. Page 2 of 3: 48 shown of 98. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 50 language tags shown of 367, by dataset count; the full filter by modality, task and language is on /datasets
Italian datasets 49–96 of 98
ITALIC: An ITALian Intent Classification Dataset ITALIC is an intent classification dataset for the Italian language, which is the first of its kind.
4 papers · 0 benchmarks
This is a large-scale dataset of tweets associated to thousands of news articles published on Italian disinformation websites in the context of 2019 European elections.
4 papers · 0 benchmarks
LAM(line-level) (The LAM Dataset: A Novel Benchmark for Line-Level Handwritten Text Recognition)
Handwritten Text Recognition (HTR) is an open problem at the intersection of Computer Vision and Natural Language Processing.
4 papers · 1 benchmark
DAWT (Densely Annotated Wikipedia Texts)
The DAWT dataset consists of Densely Annotated Wikipedia Texts across multiple languages.
3 papers · 0 benchmarks
The dataset contains the main components of the news articles published online by the newspaper named Gazzetta di Modena: url of the web page, title, sub-title, text, date of publication, crime category assigned to each news article by the…
3 papers · 0 benchmarks
DivEMT (Post-Editing Effort Across Typologically-diverse Languages)
DivEMT, the first publicly available post-editing study of Neural Machine Translation (NMT) over a typologically diverse set of target languages.
3 papers · 0 benchmarks
This is a gzipped CSV file containing the 13 million Duolingo student learning traces used in experiments by Settles & Meeder (2016).
3 papers · 0 benchmarks
KIND (Kessler Italian Named-entities Dataset)
KIND is an Italian dataset for Named-Entity Recognition.
3 papers · 0 benchmarks
MRS (Multilingual Reply Suggestion)
MRS, a multilingual reply suggestion dataset with ten languages.
3 papers · 0 benchmarks
Multilingual TOP is a dataset for multilingual semantic parsing with human-written sentences as opposed to machine translated ones.
3 papers · 0 benchmarks
Ricordi contains handwritten texts written in Italian.
3 papers · 0 benchmarks
Fanpage dataset, containing news articles taken from Fanpage.
2 papers · 1 benchmark
IlPost dataset, containing news articles taken from IlPost.
2 papers · 1 benchmark
Almawave-SLU is the first Italian dataset for Spoken Language Understanding (SLU).
2 papers · 0 benchmarks
You need to request access to download and use the dataset.
2 papers · 1 benchmark
The Parallel Meaning Bank (PMB), developed at the University of Groningen and building upon the Groningen Meaning Bank, comprises sentences and texts in raw and tokenised format, syntactic analysis, word senses, thematic roles, reference…
2 papers · 0 benchmarks
A human-revised dataset for seven languages that allows for the evaluation of multilingual RE systems.
2 papers · 0 benchmarks
Tilde MODEL Corpus is a multilingual corpora for European languages – particularly focused on the smaller languages.
2 papers · 0 benchmarks
X-WikiRE is a new, large-scale multilingual relation extraction dataset in which relation extraction is framed as a problem of reading comprehension to allow for generalization to unseen relations.
2 papers · 0 benchmarks
The Archive Query Log (AQL) is a previously unused, comprehensive query log collected at the Internet Archive over the last 25 years.
1 paper · 0 benchmarks
This paper analyses two hitherto unstudied sites sharing state-backed disinformation, Reliable Recent News (rrn.world) and WarOnFakes (waronfakes.com), which publish content in Arabic, Chinese, English, French, German, and Spanish.
1 paper · 0 benchmarks
CLIPS (Corpora e Lessici dell'Italiano Parlato e Scritto)
CLIPS, ovvero Corpora e Lessici dell'Italiano Parlato e Scritto, è uno degli otto progetti (Progetto n.
1 paper · 0 benchmarks
The dataset is composed of 95 unique document texts spanning the period 2005-2022.
1 paper · 0 benchmarks
EGO-CH-Gaze (Learning to Detect Attended Objects in Cultural Sites with Gaze Signals and Weak Object Supervision)
To study the problem of weakly supervised attended object detection in cultural sites, we collected and labeled a dataset of egocentric images acquired from subjects visiting a cultural site.
1 paper · 0 benchmarks
Dataset Description EUROPA is a dataset designed for training and evaluating multilingual keyphrase generation models in the legal domain.
1 paper · 0 benchmarks
EmoFilm (Emotional speech from Films)
EmoFilm is a multilingual emotional speech corpus comprising 1115 audio instances produced in English, Italian, and Spanish languages.
1 paper · 0 benchmarks
Fallout New Vegas Dialog is a multilingual sentiment annotated dialog dataset from Fallout New Vegas.
1 paper · 0 benchmarks
The Food Recall Incidents dataset consists of 7,546 short texts (from 5 to 360 characters each), which are the titles of food recall announcements (therefore referred to as title), crawled from 24 public food safety authority websites by…
1 paper · 0 benchmarks
HALvest is a textual dataset comprising 17 billion tokens in 56 languages and 13 domains.
1 paper · 0 benchmarks
Dataset Summary INCLUDE is a comprehensive knowledge- and reasoning-centric benchmark across 44 languages that evaluates multilingual LLMs for performance in the actual language environments where they would be deployed.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
LEMONADE is a large, expert-annotated dataset for event extraction from news articles in 20 languages: English, Spanish, Arabic, French, Italian, Russian, German, Turkish, Burmese, Indonesian, Ukrainian, Korean, Portuguese, Dutch, Somali,…
1 paper · 0 benchmarks
The M-AILABS Speech Dataset is the first large dataset that we are providing free-of-charge, freely usable as training data for speech recognition and speech synthesis.
1 paper · 1 benchmark
The MLSum-it dataset is the translated version (Helsinki-NLP/opus-mt-es-it) of the spanish portion of MLSum, containing news articles taken from BBC/mundo.
1 paper · 1 benchmark
Mega-COV is a billion-scale dataset from Twitter for studying COVID-19.
1 paper · 0 benchmarks
Mint (Multilingual Intimacy analysis)
Mint is a new Multilingual intimacy analysis dataset covering 13,384 tweets in 10 languages including English, French, Spanish, Italian, Portuguese, Korean, Dutch, Chinese, Hindi, and Arabic.
1 paper · 0 benchmarks
This dataset contains dialogue lines from the games Knights of the Old Republic 1 & 2 and Neverwinter Nights 1.
1 paper · 0 benchmarks
The first annotated corpus for multilingual analysis of potentially unfair clauses in online Terms of Service.
1 paper · 0 benchmarks
This dataset contains orthographic samples of words in 19 languages (ar, br, de, en, eno, ent, eo, es, fi, fr, fro, it, ko, nl, pt, ru, sh, tr, zh).
1 paper · 0 benchmarks
PolyNews is a multilingual dataset containing news titles in 77 languages and 19 scripts.
1 paper · 0 benchmarks
PolyNews is a multilingual parallel dataset containing news titles 833 language pairs, spanning in 64 languages and 17 scripts.
1 paper · 0 benchmarks
SQuAD-it is derived from the SQuAD dataset and it is obtained through semi-automatic translation of the SQuAD dataset into Italian.
1 paper · 0 benchmarks
WEATHub is a dataset containing 24 languages.
1 paper · 0 benchmarks
It contains data from two different realities: Food.com, a well-known American recipe site, and Planeat, an Italian site that allows you to plan recipes to save food waste.
1 paper · 0 benchmarks
Vision Language Models (VLMs) often struggle with culture-specific knowledge, particularly in languages other than English and in underrepresented cultural contexts.
1 paper · 0 benchmarks
blbooks (The British Library Books)
This dataset consists of books digitised by the British Library in partnership with Microsoft.
1 paper · 0 benchmarks
gENder-IT is an English-Italian challenge set focusing on the resolution of natural gender phenomena by providing word-level gender tags on the English source side and multiple gender alternative translations, where needed, on the Italian…
1 paper · 0 benchmarks
ChaLearn Pose is a subset of the ChaLearn 2013 Multi-modal gesture dataset from Escalera et al.
0 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.