{"url":"/dataset/curie","name":"CURIE","full_name":"CURIE: Evaluating LLMs On Multitask Scientific Long Context Understanding and Reasoning","description_markdown":"The data is organized into eight domain-specific subfolders: \"biogr\", \"dft\", \"pdb\", \"geo\", \"mpve\", \"qecc_65\", \"hfd\", and \"hfe\".  Each subfolder contains two further subfolders: \"ground_truth\" and \"inputs\".  Within these, each data instance is stored in a JSON file named record_id.json, where record_id is a unique identifier. The \"biogr\" domain also includes image inputs as record_id.png files alongside the corresponding JSON.\r\n\r\n```bash\r\ndata\r\n    ├── domain\r\n        ├── inputs\r\n        │   └── record_id.json\r\n        └── ground_truth\r\n            └── record_id.json\r\n    └── difficulty_levels.json\r\n\r\n```\r\n\r\nGround truth data varies in structure and content across domains, but all files consistently include a record_id field matching the filename.  Input files have a uniform structure across all domains, containing both a record_id field and a text field representing the input text to LLMs.\r\n\r\nFor the \"biogr\" (geo-referencing) task, for 114 of the 138 examples, we release additional data including the PDF papers that each image was taken from along with other metadata in this Github repo: [https://github.com/google-research/ecology-georeferencing](https://github.com/google-research/ecology-georeferencing)","description_withheld":null,"homepage":"https://github.com/google/curie/tree/main/data","introduced_date":"2025-03-14","introduced_date_note":null,"introduced_by":{"paper":"/paper/curie-evaluating-llms-on-multitask-scientific","title":"CURIE: Evaluating LLMs On Multitask Scientific Long Context Understanding and Reasoning","first_author":"HAO CUI","url":null},"license":null,"modalities":[],"tasks":[],"languages":[],"variants":["CURIE"],"data_loaders":[],"num_papers_in_archive":3,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}