{"url":"/dataset/surgeglobal-orca","name":"SurgeGlobal/Orca","full_name":null,"description_markdown":"## Dataset Generation\r\n\r\n-   **Base Model**: h2oai/h2ogpt-gm-oasst1-en-2048-falcon-40b-v2\r\n-   **Seed Instructions**: Derived from the FLAN-v2 Collection.\r\n-   **Generation Approach**: Explanation tuning with detailed responses generated from [h2ogpt-gm-oasst1-en-2048-falcon-40b-v2](https://huggingface.co/h2oai/h2ogpt-gm-oasst1-en-2048-falcon-40b-v2).\r\n-   **Total Instructions**: 5,507 explanation tuning data samples.\r\n\r\n### Dataset Sources\r\n\r\n-   **Repository:**  [Bitbucket Project](https://bitbucket.org/paladinanalytics/notebooks)\r\n-   **Paper :**  [Pre-Print](https://arxiv.org/abs/2404.12195)\r\n\r\n## Structure\r\n\r\nThe dataset entries consist of:\r\n-   **Query**\r\n-   **Response**\r\n-   **System Message**  (when applicable)\r\n\r\n## Usage\r\n\r\nThe Orca Dataset is intended for fine-tuning language models to not only imitate the style but also the reasoning process of LFMs, thereby improving the safety and quality of the models’ responses.\r\n\r\n## Citation\r\n\r\nIf you find our work useful, please cite our paper as follows:\r\n```\r\n@misc{surge2024openbezoar,\r\n      title={OpenBezoar: Small, Cost-Effective and Open Models Trained on Mixes of Instruction Data}, \r\n      author={Chandeepa Dissanayake and Lahiru Lowe and Sachith Gunasekara and Yasiru Ratnayake},\r\n      year={2024},\r\n      eprint={2404.12195},\r\n      archivePrefix={arXiv},\r\n      primaryClass={cs.CL}\r\n}\r\n```\r\n\r\n## Dataset Authors\r\n\r\nChandeepa Dissanayake, Lahiru Lowe, Sachith Gunasekara, and Yasiru Ratnayake","description_withheld":null,"homepage":"https://huggingface.co/datasets/SurgeGlobal/Orca","introduced_date":"2024-04-18","introduced_date_note":null,"introduced_by":{"paper":"/paper/openbezoar-small-cost-effective-and-open","title":"OpenBezoar: Small, Cost-Effective and Open Models Trained on Mixes of Instruction Data","first_author":"Chandeepa Dissanayake","url":null},"license":{"name":"Apache 2.0","url":null},"modalities":[{"name":"Texts","url":"/datasets/modality/texts"}],"tasks":[{"name":"Instruction Following","url":"/task/instruction-following","datasets_with_task":"/datasets/task/instruction-following"}],"languages":[{"name":"English","url":"/datasets/language/english"}],"variants":["SurgeGlobal/Orca"],"data_loaders":[],"num_papers_in_archive":1,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}