{"url":"/dataset/casehold","name":"CaseHOLD","full_name":"Case Holdings On Legal Decisions","description_markdown":"**CaseHOLD** (Case Holdings On Legal Decisions) is a law dataset comprised of over 53,000+ multiple choice questions to identify the relevant holding of a cited case. This dataset presents a fundamental task to lawyers and is both legally meaningful and difficult from an NLP perspective (F1 of 0.4 with a BiLSTM baseline). The citing context from the judicial decision serves as the prompt for the question. The answer choices are holding statements derived from citations following text in a legal decision. There are five answer choices for each citing text. The correct answer is the holding statement that corresponds to the citing text. The four incorrect answers are other holding statements.\r\n\r\nTo read more about the dataset, please see our [paper](https://arxiv.org/abs/2104.08671) or our [blogpost](https://reglab.stanford.edu/data/casehold-benchmark/).","description_withheld":null,"homepage":"https://github.com/reglab/casehold","introduced_date":"2021-04-18","introduced_date_note":null,"introduced_by":{"paper":"/paper/when-does-pretraining-help-assessing-self","title":"When Does Pretraining Help? Assessing Self-Supervised Learning for Law and the CaseHOLD Dataset","first_author":"Lucia Zheng","url":null},"license":{"name":"Unknown","url":null},"modalities":[{"name":"Texts","url":"/datasets/modality/texts"}],"tasks":[{"name":"Question Answering","url":"/task/question-answering","datasets_with_task":"/datasets/task/question-answering"},{"name":"Few-Shot Learning","url":"/task/few-shot-learning","datasets_with_task":"/datasets/task/few-shot-learning"}],"languages":[{"name":"English","url":"/datasets/language/english"}],"variants":["CaseHOLD"],"data_loaders":[{"repo":"https://github.com/reglab/casehold","url":"https://github.com/reglab/casehold","frameworks":["pytorch"]}],"num_papers_in_archive":27,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[{"leaderboard":"/sota/question-answering-on-casehold","task":"Question Answering","dataset_variant":"CaseHOLD","rows":3,"metrics":["Macro F1 (10-fold)"],"first_row_in_archive_order":{"model":"Custom Legal-BERT","paper":"/paper/when-does-pretraining-help-assessing-self","metrics":{"Macro F1 (10-fold)":"69.5"},"code_links":[{"title":"reglab/casehold","url":"https://github.com/reglab/casehold"},{"title":"trusthlt/privacy-legal-nlp-lm","url":"https://github.com/trusthlt/privacy-legal-nlp-lm"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"},{"leaderboard":"/sota/few-shot-learning-on-casehold","task":"Few-Shot Learning","dataset_variant":"CaseHOLD","rows":1,"metrics":["Accuracy"],"first_row_in_archive_order":{"model":"CoT-T5-11B (1024 Shot)","paper":"/paper/the-cot-collection-improving-zero-shot-and","metrics":{"Accuracy":"68.3"},"code_links":[{"title":"kaistai/cot-collection","url":"https://github.com/kaistai/cot-collection"},{"title":"kaist-lklab/cot-collection","url":"https://github.com/kaist-lklab/cot-collection"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"}],"papers_with_a_benchmark_row":[{"paper":"/paper/the-cot-collection-improving-zero-shot-and","title":"The CoT Collection: Improving Zero-shot and Few-shot Learning of Language Models via Chain-of-Thought Fine-Tuning","date":"2023-05-23","rows_on_this_dataset":1,"code_links":2,"syntology":null},{"paper":"/paper/when-does-pretraining-help-assessing-self","title":"When Does Pretraining Help? Assessing Self-Supervised Learning for Law and the CaseHOLD Dataset","date":"2021-04-18","rows_on_this_dataset":3,"code_links":2,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":9,"samples_ran":0,"samples_unverified":9,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}}],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":1,"samples_harvested":9,"samples_ran":0,"samples_unverified":9,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":1,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}