Datasets › LIMA

LIMA

Introduced by Chunting Zhou et al. in LIMA: Less Is More for Alignment18 May 2023 archive 2025-07-28

The LIMA dataset is a valuable resource used in natural language processing (NLP) research. Let me provide you with some details:

  1. Origin and Purpose:
  2. The LIMA dataset is derived from the LLaMa language model, which has an impressive 65 billion parameters.
  3. It serves as a fine-tuned version of the LLaMa model, specifically adjusted using approximately 1,000 prompts and responses.

  4. Performance and Applications:

  5. LIMA demonstrates remarkable performance by learning to follow specific response formats from just a handful of examples in the training data.
  6. The dataset covers a wide range of tasks, including complex queries such as planning trip itineraries and speculating about alternate history.
  7. Interestingly, the model tends to generalize well to unseen tasks that were not part of the training data.

  8. License:

  9. The licensing of the LIMA dataset depends on the source data it was derived from:
    • If the source data has a stricter license than CC BY-NC-SA, the LIMA dataset follows the same restrictions.
    • Otherwise, it adheres to the CC BY-NC-SA license.

(1) GAIR/lima · Datasets at Hugging Face. https://huggingface.co/datasets/GAIR/lima. (2) GAIR/lima at main - Hugging Face. https://huggingface.co/datasets/GAIR/lima/tree/main. (3) 日本語LIMAデータセットlima-jaを作成したので公開します. https://zanote.net/ai/lima-ja/. (4) Paper page - LIMA: Less Is More for Alignment - Hugging Face. https://huggingface.co/papers/2305.11206. (5) undefined. https://huggingface.co/datasets/GAIR/lima/.

Benchmarks archive 2025-07-28

No leaderboard in the archive resolves to this dataset.

Papers archive 2025-07-28

No paper in the archive has a leaderboard row on this dataset; the archive counts 107 papers for it but never published that list.

Dataset loaders archive 2025-07-28

No loader listed in the archive.

Tasks archive 2025-07-28

No task tagged in the archive.

License archive 2025-07-28

No licence recorded in the archive. Absence here is not a statement about the dataset's terms.

Modalities archive 2025-07-28

No modality tagged.

Languages archive 2025-07-28

Variants archive 2025-07-28

  • LIMA

1 variant name, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections