{"url":"/dataset/tqba","name":"TQBA++","full_name":"Tiny QA Benchmark++","description_markdown":"Ultra-lightweight, multilingual QA eval dataset for rapid testing LLMs.\r\n\r\nDataset Characteristics:\r\n\r\n- Multilingual: Includes packs for Arabic, Chinese, French, German, Japanese, Korean, Portuguese, Russian, Spanish, and Turkish.\r\n- Compact: Contains a curated English gold-standard set of 52 QA pairs (under 20kB), enabling immediate and resource-friendly evaluations.\r\n- Synthetic Generation: Features a LiteLLM-powered synthetic data generator see [tinyqabenchmarkpp](https://pypi.org/project/tinyqabenchmarkpp/), allowing quick creation of custom evaluation sets tailored to specific domains or languages.\r\n- Metadata Support: Provided in Croissant-compatible formats, ready for seamless integration with modern evaluation harnesses and CI tools.\r\n\r\nMotivation and Content Summary:\r\n\r\nThe primary motivation behind TQB++ is to enable rapid iteration and continuous integration (CI) of language models. Existing evaluation benchmarks typically involve significant computational overhead and slow feedback loops. In contrast, TQB++ offers near-instantaneous assessments of model performance and prompt stability across multiple languages. It is particularly sensitive to issues such as prompt-template regressions, tokenizer drift, and fine-tuning side effects.\r\n\r\nPotential Use Cases:\r\n- Continuous Integration (CI): Immediate detection of breaking changes or regressions in LLM pipelines.\r\n- Multilingual Model Validation: Quickly assess model accuracy and performance across multiple languages without large compute costs.\r\n- Prompt Optimization and Testing: Ideal for iterative prompt refinement workflows, enabling fast feedback loops and effective tuning.\r\n- Teaching and Prototyping: Educational use in courses or workshops, showcasing multilingual LLM evaluation in real-time scenarios.","description_withheld":null,"homepage":"https://huggingface.co/datasets/vincentkoc/tiny_qa_benchmark_pp","introduced_date":"2025-05-20","introduced_date_note":null,"introduced_by":{"paper":"/paper/tiny-qa-benchmark-ultra-lightweight-synthetic","title":"Tiny QA Benchmark++: Ultra-Lightweight, Synthetic Multilingual Dataset Generation & Smoke-Tests for Continuous LLM Evaluation","first_author":"Vincent Koc","url":null},"license":{"name":"Apache-2.0","url":"https://github.com/vincentkoc/tiny_qa_benchmark_pp?tab=readme-ov-file#license"},"modalities":[{"name":"Texts","url":"/datasets/modality/texts"}],"tasks":[{"name":"Question Answering","url":"/task/question-answering","datasets_with_task":"/datasets/task/question-answering"},{"name":"Text Summarization","url":"/task/text-summarization","datasets_with_task":"/datasets/task/text-summarization"},{"name":"TinyQA Benchmark++","url":"/task/tinyqa-benchmark","datasets_with_task":"/datasets/task/tinyqa-benchmark"}],"languages":[{"name":"English","url":"/datasets/language/english"},{"name":"French","url":"/datasets/language/french"},{"name":"Spanish","url":"/datasets/language/spanish"},{"name":"German","url":"/datasets/language/german"},{"name":"Chinese","url":"/datasets/language/chinese"},{"name":"Japanese","url":"/datasets/language/japanese"},{"name":"Russian","url":"/datasets/language/russian"},{"name":"Portuguese","url":"/datasets/language/portuguese"},{"name":"Arabic","url":"/datasets/language/arabic"},{"name":"Turkish","url":"/datasets/language/turkish"}],"variants":["TQBA++"],"data_loaders":[],"num_papers_in_archive":1,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}