{"url":"/dataset/farstail","name":"FarsTail","full_name":null,"description_markdown":"Natural Language Inference (NLI), also called Textual Entailment, is an important task in NLP with the goal of determining the inference relationship between a premise p and a hypothesis h. It is a three-class problem, where each pair (p, h) is assigned to one of these classes: \"ENTAILMENT\" if the hypothesis can be inferred from the premise, \"CONTRADICTION\" if the hypothesis contradicts the premise, and \"NEUTRAL\" if none of the above holds. There are large datasets such as SNLI, MNLI, and SciTail for NLI in English, but there are few datasets for poor-data languages like Persian. Persian (Farsi) language is a pluricentric language spoken by around 110 million people in countries like Iran, Afghanistan, and Tajikistan. **FarsTail** is the first relatively large-scale Persian dataset for NLI task. A total of 10,367 samples are generated from a collection of 3,539 multiple-choice questions. The train, validation, and test portions include 7,266, 1,537, and 1,564 instances, respectively.\r\n\r\nSource: [https://github.com/dml-qom/FarsTail](https://github.com/dml-qom/FarsTail)\r\nImage Source: [https://github.com/dml-qom/FarsTail](https://github.com/dml-qom/FarsTail)","description_withheld":null,"homepage":"https://github.com/dml-qom/FarsTail","introduced_date":"2020-09-18","introduced_date_note":null,"introduced_by":{"paper":"/paper/farstail-a-persian-natural-language-inference","title":"FarsTail: A Persian Natural Language Inference Dataset","first_author":"Hossein Amirkhani","url":null},"license":{"name":"Apache License 2.0","url":"https://github.com/dml-qom/FarsTail/blob/master/LICENSE"},"modalities":[{"name":"Texts","url":"/datasets/modality/texts"}],"tasks":[{"name":"Natural Language Inference","url":"/task/natural-language-inference","datasets_with_task":"/datasets/task/natural-language-inference"}],"languages":[{"name":"Persian","url":"/datasets/language/persian"}],"variants":["FarsTail"],"data_loaders":[{"repo":"https://github.com/dml-qom/FarsTail","url":"https://github.com/dml-qom/FarsTail","frameworks":[]}],"num_papers_in_archive":9,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[{"leaderboard":"/sota/natural-language-inference-on-farstail","task":"Natural Language Inference","dataset_variant":"FarsTail","rows":10,"metrics":["% Test Accuracy"],"first_row_in_archive_order":{"model":"mBERT","paper":"/paper/farstail-a-persian-natural-language-inference","metrics":{"% Test Accuracy":"83.38"},"code_links":[{"title":"dml-qom/FarsTail","url":"https://github.com/dml-qom/FarsTail"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"}],"papers_with_a_benchmark_row":[{"paper":"/paper/farstail-a-persian-natural-language-inference","title":"FarsTail: A Persian Natural Language Inference Dataset","date":"2020-09-18","rows_on_this_dataset":10,"code_links":1,"syntology":null}],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}