{"url":"/dataset/tamil-alpaca-orca","name":"Tamil Alpaca Orca","full_name":null,"description_markdown":"# Dataset Card for \"tamil-alpaca\"\r\n\r\nThis repository includes a Tamil-translated versions of the [Alpaca dataset](https://huggingface.co/datasets/yahma/alpaca-cleaned) and a subset of [OpenOrca](https://huggingface.co/datasets/Open-Orca/OpenOrca) dataset. \r\n\r\nThis dataset is part of the release of Tamil LLaMA family of models – an important step in advancing LLMs for the Tamil language. To dive deep into the development and capabilities of this model, please read the [research paper](https://arxiv.org/abs/2311.05845) and the [introductory blog post (WIP) ]() that outlines our journey and the model's potential impact.\r\n\r\n**GitHub Repository:** [https://github.com/abhinand5/tamil-llama](https://github.com/abhinand5/tamil-llama)\r\n\r\n## Models trained using this dataset\r\n\r\n| Model                    | Type                        | Data              | Base Model           | # Params | Download Links                                                         |\r\n|--------------------------|-----------------------------|-------------------|----------------------|------|------------------------------------------------------------------------|\r\n| Tamil LLaMA 7B Instruct  | Instruction following model | 145k instructions | Tamil LLaMA 7B Base  | 7B   | [HF Hub](https://huggingface.co/abhinand/tamil-llama-7b-instruct-v0.1) |\r\n| Tamil LLaMA 13B Instruct | Instruction following model | 145k instructions | Tamil LLaMA 13B Base | 13B  | [HF Hub](abhinand/tamil-llama-13b-instruct-v0.1)                       |\r\n\r\n## Meet the Developers\r\n\r\nGet to know the creators behind this innovative model and follow their contributions to the field:\r\n\r\n- [Abhinand Balachandran](https://www.linkedin.com/in/abhinand-05/)\r\n\r\n## Citation\r\n\r\nIf you use this model or any of the the Tamil-Llama datasets in your research, please cite:\r\n\r\n```bibtex\r\n@misc{balachandran2023tamilllama,\r\n      title={Tamil-Llama: A New Tamil Language Model Based on Llama 2}, \r\n      author={Abhinand Balachandran},\r\n      year={2023},\r\n      eprint={2311.05845},\r\n      archivePrefix={arXiv},\r\n      primaryClass={cs.CL}\r\n}\r\n```","description_withheld":null,"homepage":"https://huggingface.co/datasets/abhinand/tamil-alpaca-orca","introduced_date":"2023-11-13","introduced_date_note":null,"introduced_by":{"paper":"/paper/tamil-llama-a-new-tamil-language-model-based","title":"Tamil-Llama: A New Tamil Language Model Based on Llama 2","first_author":"Abhinand Balachandran","url":null},"license":{"name":"gpl-3.0","url":null},"modalities":[{"name":"Texts","url":"/datasets/modality/texts"}],"tasks":[{"name":"Text Generation","url":"/task/text-generation","datasets_with_task":"/datasets/task/text-generation"},{"name":"Instruction Following","url":"/task/instruction-following","datasets_with_task":"/datasets/task/instruction-following"}],"languages":[{"name":"Tamil","url":"/datasets/language/tamil"}],"variants":["Tamil Alpaca Orca"],"data_loaders":[],"num_papers_in_archive":1,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}