{"url":"/dataset/surgeglobal-lamini","name":"SurgeGlobal/LaMini","full_name":null,"description_markdown":"## Overview\r\nThe LaMini Dataset is an instruction dataset generated using [h2ogpt-gm-oasst1-en-2048-falcon-40b-v2](https://huggingface.co/h2oai/h2ogpt-gm-oasst1-en-2048-falcon-40b-v2). It is designed for instruction-tuning pre-trained models to specialize them in a variety of downstream tasks.\r\n\r\n## Dataset Generation\r\n- **Base Model**: h2oai/h2ogpt-gm-oasst1-en-2048-falcon-40b-v2.\r\n- **Seed Instructions**: Sourced from databricks/databricks-dolly-15k dataset.\r\n- **Generation Approach**: Example-guided and topic-guided strategies.\r\n- **Total Instructions**: 1,504 unique instruction examples.\r\n\r\n### Dataset Sources\r\n\r\n- **Repository:** [Bitbucket Project](https://bitbucket.org/paladinanalytics/workspace/projects/OP)\r\n- **Paper :** [Pre-Print](https://arxiv.org/abs/2404.12195)\r\n\r\n## Structure\r\nEach entry in the dataset contains:\r\n- **Instruction**\r\n- **Response**\r\n\r\n## Usage\r\nThe LaMini Dataset can be used to fine-tune language models to improve their ability to follow instructions and generate relevant responses.\r\n\r\n## Access\r\nThe dataset is available on HuggingFace at the following link: [https://huggingface.co/datasets/SurgeGlobal/LaMini](https://huggingface.co/datasets/SurgeGlobal/LaMini)\r\n\r\n## Citation\r\nIf you find our work useful, please cite our paper as follows:\r\n```\r\n@misc{surge2024openbezoar,\r\n      title={OpenBezoar: Small, Cost-Effective and Open Models Trained on Mixes of Instruction Data}, \r\n      author={Chandeepa Dissanayake and Lahiru Lowe and Sachith Gunasekara and Yasiru Ratnayake},\r\n      year={2024},\r\n      eprint={2404.12195},\r\n      archivePrefix={arXiv},\r\n      primaryClass={cs.CL}\r\n}\r\n```\r\n\r\n## Dataset Authors\r\n\r\nChandeepa Dissanayake, Lahiru Lowe, Sachith Gunasekara, and Yasiru Ratnayake","description_withheld":null,"homepage":"https://huggingface.co/datasets/SurgeGlobal/LaMini","introduced_date":"2024-04-18","introduced_date_note":null,"introduced_by":{"paper":"/paper/openbezoar-small-cost-effective-and-open","title":"OpenBezoar: Small, Cost-Effective and Open Models Trained on Mixes of Instruction Data","first_author":"Chandeepa Dissanayake","url":null},"license":{"name":"Apache 2.0","url":null},"modalities":[{"name":"Texts","url":"/datasets/modality/texts"}],"tasks":[{"name":"Instruction Following","url":"/task/instruction-following","datasets_with_task":"/datasets/task/instruction-following"}],"languages":[{"name":"English","url":"/datasets/language/english"}],"variants":["SurgeGlobal/LaMini"],"data_loaders":[],"num_papers_in_archive":1,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}