{"url":"/dataset/italic","name":"ITALIC","full_name":"ITALIC","description_markdown":"**ITALIC: An ITALian Intent Classification Dataset**\r\n\r\nITALIC is an intent classification dataset for the Italian language, which is the first of its kind. \r\nIt includes spoken and written utterances and is annotated with 60 intents. \r\nThe dataset is available on [Zenodo](https://zenodo.org/record/8040649) and connectors ara available for the [HuggingFace Hub](https://huggingface.co/datasets/RiTA-nlp/ITALIC).\r\n\r\n### Data collection\r\n\r\nThe data collection follows the MASSIVE NLU dataset which contains an annotated textual dataset for 60 intents. The data collection process is described in the paper [Massive Natural Language Understanding](https://arxiv.org/abs/2204.08582).\r\n\r\nFollowing the MASSIVE NLU dataset, a pool of 70+ volunteers has been recruited to annotate the dataset. The volunteers were asked to record their voice while reading the utterances (the original text is available on MASSIVE dataset). Together with the audio, the volunteers were asked to provide a self-annotated description of the recording conditions (e.g., background noise, recording device). The audio recordings have also been validated and, in case of errors, re-recorded by the volunteers.\r\n\r\nAll the audio recordings included in the dataset have received a validation from at least two volunteers. All the audio recordings have been validated by native italian speakers (self-annotated).","description_withheld":null,"homepage":"https://github.com/RiTA-nlp/ITALIC/","introduced_date":"2023-06-14","introduced_date_note":null,"introduced_by":{"paper":"/paper/italic-an-italian-intent-classification","title":"ITALIC: An Italian Intent Classification Dataset","first_author":"Alkis Koudounas","url":null},"license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"modalities":[{"name":"Texts","url":"/datasets/modality/texts"},{"name":"Audio","url":"/datasets/modality/audio"}],"tasks":[{"name":"Intent Detection","url":"/task/intent-detection","datasets_with_task":"/datasets/task/intent-detection"},{"name":"Automatic Speech Recognition","url":"/task/automatic-speech-recognition-2","datasets_with_task":"/datasets/task/automatic-speech-recognition-2"},{"name":"Automatic Speech Recognition (ASR)","url":"/task/automatic-speech-recognition","datasets_with_task":"/datasets/task/automatic-speech-recognition"}],"languages":[{"name":"Italian","url":"/datasets/language/italian"}],"variants":["ITALIC"],"data_loaders":[],"num_papers_in_archive":4,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}