{"url":"/dataset/propara","name":"ProPara","full_name":null,"description_markdown":"The **ProPara** dataset is designed to train and test comprehension of simple paragraphs describing processes (e.g., photosynthesis), designed for the task of predicting, tracking, and answering questions about how entities change during the process.\r\n\r\nProPara aims to promote the research in natural language understanding in the context of procedural text. This requires identifying the actions described in the paragraph and tracking state changes happening to the entities involved. The comprehension task is treated as that of predicting, tracking, and answering questions about how entities change during the procedure. The dataset contains 488 paragraphs and 3,300 sentences. Each paragraph is richly annotated with the existence and locations of all the main entities (the “participants”) at every time step (sentence) throughout the procedure (~81,000 annotations).\r\n\r\nProPara paragraphs are natural (authored by crowdsourcing) rather than synthetic (e.g., in bAbI). Workers were given a prompt (e.g., “What happens during photosynthesis?”) and then asked to author a series of sentences describing the sequence of events in the procedure. From these sentences, participant entities and their existence and locations were identified. The goal of the challenge is to predict the existence and location of each participant, based on sentences in the paragraph.\r\n\r\nSource: [Allen Institute for AI](https://allenai.org/data/propara)","description_withheld":null,"homepage":"https://allenai.org/data/propara","introduced_date":"2018-05-17","introduced_date_note":null,"introduced_by":{"paper":"/paper/tracking-state-changes-in-procedural-text-a","title":"Tracking State Changes in Procedural Text: A Challenge Dataset and Models for Process Paragraph Comprehension","first_author":"Bhavana Dalvi Mishra","url":null},"license":null,"modalities":[{"name":"Texts","url":"/datasets/modality/texts"}],"tasks":[{"name":"Reading Comprehension","url":"/task/reading-comprehension","datasets_with_task":"/datasets/task/reading-comprehension"},{"name":"Procedural Text Understanding","url":"/task/procedural-text-understanding","datasets_with_task":"/datasets/task/procedural-text-understanding"}],"languages":[],"variants":["ProPara"],"data_loaders":[],"num_papers_in_archive":34,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}