{"url":"/dataset/docnli","name":"DocNLI","full_name":null,"description_markdown":"**DocNLI** is a large-scale dataset for document-level NLI. DocNLI is transformed from a broad range of NLP problems and covers multiple genres of text. The premises always stay in the document granularity, whereas the hypotheses vary in length from single sentences to passages with hundreds of words. Additionally, DocNLI has pretty limited artifacts which unfortunately widely exist in some popular sentence-level NLI datasets.","description_withheld":null,"homepage":"https://github.com/salesforce/DocNLI","introduced_date":"2021-06-17","introduced_date_note":null,"introduced_by":{"paper":"/paper/docnli-a-large-scale-dataset-for-document","title":"DocNLI: A Large-scale Dataset for Document-level Natural Language Inference","first_author":"Wenpeng Yin","url":null},"license":{"name":"BSD 3-Clause License","url":"https://github.com/salesforce/DocNLI/blob/main/LICENSE.txt"},"modalities":[{"name":"Texts","url":"/datasets/modality/texts"}],"tasks":[{"name":"Natural Language Inference","url":"/task/natural-language-inference","datasets_with_task":"/datasets/task/natural-language-inference"}],"languages":[{"name":"English","url":"/datasets/language/english"}],"variants":["DocNLI"],"data_loaders":[{"repo":"https://github.com/tensorflow/datasets","url":"https://www.tensorflow.org/datasets/catalog/doc_nli","frameworks":["tf","pytorch","jax"]}],"num_papers_in_archive":8,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-25T09:33:49+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}