Datasets › Databricks Dolly 15k

Databricks Dolly 15k (databricks-dolly-15k)

Introduced by Emily M. Bender in 100 Things You Always Wanted to Know about Linguistics But Were Afraid to Ask*1 Jun 2012 archive 2025-07-28

Databricks Dolly 15k is a dataset containing 15,000 high-quality human-generated prompt / response pairs specifically designed for instruction tuning large language models. It is authored by more than 5,000 Databricks employees during March and April of 2023. The training records are natural, expressive and designed to represent a wide range of the behaviors, from brainstorming and content generation to information extraction and summarization.

Source: Free Dolly: Introducing the World's First Truly Open Instruction-Tuned LLM

Benchmarks archive 2025-07-28

No leaderboard in the archive resolves to this dataset.

Papers archive 2025-07-28

No paper in the archive has a leaderboard row on this dataset; the archive counts 4 papers for it but never published that list.

Dataset loaders archive 2025-07-28

No loader listed in the archive.

Tasks archive 2025-07-28

License archive 2025-07-28

Creative Commons Attribution-ShareAlike 3.0 Unported License

Modalities archive 2025-07-28

Languages archive 2025-07-28

Variants archive 2025-07-28

  • Databricks Dolly 15k

1 variant name, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections