{"url":"/dataset/humset","name":"HumSet","full_name":"HumSet","description_markdown":"Timely and effective response to humanitarian crises requires quick and accurate analysis of large amounts of text data, a process that can highly benefit from expert-assisted NLP systems trained on validated and annotated data in the humanitarian response domain. To enable creation of such NLP systems, we introduce and release **HumSet**, a novel and rich multilingual dataset of humanitarian response documents annotated by experts in the humanitarian response community. The dataset provides documents in three languages (English, French, Spanish) and covers a variety of humanitarian crises from 2018 to 2021 across the globe. For each document, **HumSet** provides selected snippets (entries) as well as assigned classes to each entry annotated using common humanitarian information analysis frameworks. **HumSet** also provides novel and challenging entry extraction and multi-label entry classification tasks. In this paper, we take a first step towards approaching these tasks and conduct a set of experiments on  Pre-trained Language Models (PLM) to establish strong baselines for future research in this domain. The dataset is available at [https://blog.thedeep.io/humset/](https://blog.thedeep.io/humset/).","description_withheld":null,"homepage":"https://blog.thedeep.io/humset/","introduced_date":"2022-10-10","introduced_date_note":null,"introduced_by":{"paper":"/paper/humset-dataset-of-multilingual-information","title":"HumSet: Dataset of Multilingual Information Extraction and Classification for Humanitarian Crisis Response","first_author":"Selim Fekih","url":null},"license":null,"modalities":[{"name":"Texts","url":"/datasets/modality/texts"},{"name":"Tabular","url":"/datasets/modality/tabular"}],"tasks":[{"name":"Text Summarization","url":"/task/text-summarization","datasets_with_task":"/datasets/task/text-summarization"},{"name":"Multilabel Text Classification","url":"/task/multilabel-text-classification","datasets_with_task":"/datasets/task/multilabel-text-classification"},{"name":"Humanitarian","url":"/task/humanitarian","datasets_with_task":"/datasets/task/humanitarian"},{"name":"Multilingual NLP","url":"/task/multilingual-nlp","datasets_with_task":"/datasets/task/multilingual-nlp"}],"languages":[{"name":"English","url":"/datasets/language/english"},{"name":"French","url":"/datasets/language/french"},{"name":"Spanish","url":"/datasets/language/spanish"}],"variants":["HumSet"],"data_loaders":[],"num_papers_in_archive":2,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}