Home › Datasets › language › Hindi
Hindi datasets
archive 2025-07-28
81 datasets carry the language tag "Hindi", ordered by the archive's paper count. Page 2 of 2: 33 shown of 81. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 50 language tags shown of 367, by dataset count; the full filter by modality, task and language is on /datasets
Hindi datasets 49–81 of 81
This dataset releases a significantly sized standard-abiding Hindi NER dataset containing 109,146 sentences and 2,220,856 tokens, annotated with 3 collapsed tags (PER, LOC, ORG).
1 paper · 1 benchmark
HiNER-original (HiNER: A Large Hindi Named Entity Recognition Dataset)
This dataset releases a significantly sized standard-abiding Hindi NER dataset containing 109,146 sentences and 2,220,856 tokens, annotated with 11 tags.
1 paper · 1 benchmark
For testing refusal behavior in a language-specific setting, we introduce HiXSTest — a set of manually curated prompts in the Hindi language designed to measure exaggerated safety.
1 paper · 0 benchmarks
Hinglish-TOP is a human annotated code-switched semantic parsing dataset containing 10k human annotations for Hindi-English (HINGLISH) code switched utterances, and over 170K CST5 generated code-switched utterances from the TOPv2 dataset.
1 paper · 0 benchmarks
The dataset is taken from the First shared task on Information Extractor for Conversational Systems in Indian Languages (IECSIL) .
1 paper · 1 benchmark
Dataset Summary INCLUDE is a comprehensive knowledge- and reasoning-centric benchmark across 44 languages that evaluates multilingual LLMs for performance in the actual language environments where they would be deployed.
1 paper · 0 benchmarks
IRLCov19 is a multilingual Twitter dataset related to Covid-19 collected in the period between February 2020 to July 2020 specifically for regional languages in India.
1 paper · 0 benchmarks
Full Koo platform Dataset.
1 paper · 0 benchmarks
M3LS (Multi-Lingual Multi-Modal Summarization Dataset)
Significant developments in techniques such as encoder-decoder models have enabled us to represent information comprising multiple modalities.
1 paper · 0 benchmarks
MAKED (MultiModal MultiLingual Summarization and Keyword Extraction Dataset)
Keyword extraction is an integral task for many downstream problems like clustering, recommendation, search and classification.
1 paper · 0 benchmarks
MILU (Multi-task Indic Language Understanding Benchmark)
Overview MILU (Multi-task Indic Language Understanding Benchmark) is a comprehensive evaluation dataset designed to assess the performance of Large Language Models (LLMs) across 11 Indic languages.
1 paper · 0 benchmarks
MalayalamMixSentiment is a Sentiment Analysis Dataset for Code-Mixed Malayalam-English.
1 paper · 0 benchmarks
Mega-COV is a billion-scale dataset from Twitter for studying COVID-19.
1 paper · 0 benchmarks
Mint (Multilingual Intimacy analysis)
Mint is a new Multilingual intimacy analysis dataset covering 13,384 tweets in 10 languages including English, French, Spanish, Italian, Portuguese, Korean, Dutch, Chinese, Hindi, and Arabic.
1 paper · 0 benchmarks
We announce the release of a new multilingual speaker dataset called NITK-IISc Multilingual Multi-accent Speaker Profiling(NISP) dataset.
1 paper · 0 benchmarks
OTSC-Hindi (Occupation Test Set with Simple Context - Hindi)
Test set of sentences in Hindi with simple gender-specific context used to measure gender bias in NMT systems for Hindi-English.
1 paper · 0 benchmarks
Poly-FEVER is a multilingual fact verification benchmark designed to evaluate hallucination detection in large language models (LLMs).
1 paper · 0 benchmarks
PolyNews is a multilingual dataset containing news titles in 77 languages and 19 scripts.
1 paper · 0 benchmarks
PolyNews is a multilingual parallel dataset containing news titles 833 language pairs, spanning in 64 languages and 17 scripts.
1 paper · 0 benchmarks
We release various types of word embeddings for multiple Indian languages.
1 paper · 0 benchmarks
ROAST (Review level Opinion Aspect Sentiment Target Joint Detection for ABSA)
This repository has a review-level multidomain multilingual dataset for Aspect-based Sentiment Analysis(ABSA) for the paper ROAST: Review-level Opinion Aspect Sentiment Target Joint Detection.
1 paper · 0 benchmarks
SAIL 2017 (Sentiment Analysis for Indian Languages)
India is a linguistic area with one of the longest histories of contact, influence, use, teaching and learning of English-in-diaspora in the world (Kachru and Nelson, 2006).
1 paper · 1 benchmark
The ComMA Dataset v0.2 is a multilingual dataset annotated with a hierarchical, fine-grained tagset marking different types of aggression and the "context" in which they occur.
1 paper · 0 benchmarks
WEATHub is a dataset containing 24 languages.
1 paper · 0 benchmarks
Test set of sentences in Hindi with complex coreference involving two entities inspired by WinoBias format of sentences in English.
1 paper · 0 benchmarks
Vision Language Models (VLMs) often struggle with culture-specific knowledge, particularly in languages other than English and in underrepresented cultural contexts.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
We provide a new data set XWikiRef for the task of Cross-lingual Multi-document Summarization.
1 paper · 0 benchmarks
Download free fonts in DaFont style from our extensive collection.
1 paper · 0 benchmarks
We present sentence aligned parallel corpora across 10 Indian Languages - Hindi, Telugu, Tamil, Malayalam, Gujarati, Urdu, Bengali, Oriya, Marathi, Punjabi, and English - many of which are categorized as low resource.
0 papers · 0 benchmarks
MIMIC Meme Dataset (Misogyny Identification in Multimodal Internet Content in Hindi-English Code-Mix Language)
This dataset endeavors to fill the research void by presenting a meticulously curated collection of misogynistic memes in a code-mixed language of Hindi and English.
0 papers · 0 benchmarks
The increase in religiously motivated hate on social media is clear and ongoing.
0 papers · 0 benchmarks
THVD (Talking Head Video Dataset)
About We provide a comprehensive talking-head video dataset with over 50,000 videos, totaling more than 500+ hours of footage and featuring 20,841 unique identities from around the world.
0 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.