{"url":"/dataset/surgeglobal-evol-instruct","name":"SurgeGlobal/Evol-Instruct","full_name":null,"description_markdown":"## Dataset Generation\r\n\r\n- **Base Model**: h2oai/h2ogpt-gm-oasst1-en-2048-falcon-40b-v2\r\n- **Seed Instructions**: Selected from the databricks/databricks-dolly-15k dataset\r\n- **Generation Approach**: Iterative evolution of instructions using a conversational syntax for in-depth and in-breadth evolving\r\n- **Total Instructions**: 2,304 instruction tuning data samples\r\n\r\n### Dataset Sources\r\n\r\n- **Repository**: [Bitbucket Project](https://bitbucket.org/paladinanalytics/notebooks)\r\n- **Paper**: [Pre-Print](https://arxiv.org/abs/2404.12195)\r\n\r\n### Structure\r\n\r\nThe dataset entries consist of:\r\n- **Instruction**\r\n- **Response**\r\n- **Evolution Strategy** (in-depth or in-breadth)\r\n- **Category** (of the original instruction)\r\n\r\n### Usage\r\n\r\nThe Evol-Instruct Dataset is designed for the automatic evolution of instruction datasets, enhancing the complexity and diversity of instructions to train language models for a wide range of tasks.\r\n\r\n## Citation\r\n\r\nIf you find our work useful, please cite our paper as follows:\r\n```\r\n@misc{surge2024openbezoar,\r\n      title={OpenBezoar: Small, Cost-Effective and Open Models Trained on Mixes of Instruction Data}, \r\n      author={Chandeepa Dissanayake and Lahiru Lowe and Sachith Gunasekara and Yasiru Ratnayake},\r\n      year={2024},\r\n      eprint={2404.12195},\r\n      archivePrefix={arXiv},\r\n      primaryClass={cs.CL}\r\n}\r\n```\r\n\r\n## Dataset Authors\r\n\r\nChandeepa Dissanayake, Lahiru Lowe, Sachith Gunasekara, and Yasiru Ratnayake","description_withheld":null,"homepage":"https://huggingface.co/datasets/SurgeGlobal/Evol-Instruct","introduced_date":"2024-04-18","introduced_date_note":null,"introduced_by":{"paper":"/paper/openbezoar-small-cost-effective-and-open","title":"OpenBezoar: Small, Cost-Effective and Open Models Trained on Mixes of Instruction Data","first_author":"Chandeepa Dissanayake","url":null},"license":{"name":"Apache 2.0","url":null},"modalities":[{"name":"Texts","url":"/datasets/modality/texts"}],"tasks":[{"name":"Instruction Following","url":"/task/instruction-following","datasets_with_task":"/datasets/task/instruction-following"}],"languages":[{"name":"English","url":"/datasets/language/english"}],"variants":["SurgeGlobal/Evol-Instruct"],"data_loaders":[],"num_papers_in_archive":1,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}