Papers › Tiny QA Benchmark++: Ultra-Lightweight, Synthetic Multilingual Dataset Generation &...

Tiny QA Benchmark++: Ultra-Lightweight, Synthetic Multilingual Dataset Generation & Smoke-Tests for Continuous LLM Evaluation

17 May 2025arXiv:2505.12058archive 2025-07-28

Vincent Koc

Tiny QA Benchmark++ (TQB++) presents an ultra-lightweight, multilingual smoke-test suite designed to give large-language-model (LLM) pipelines a unit-test style safety net dataset that runs in seconds with minimal cost. Born out of the tight feedback-loop demands building the Comet Opik prompt-optimization SDK, where waiting on heavyweight benchmarks breaks developer flow. TQB++ couples a 52-item English gold set (less than 20 kB) with a tiny synthetic-data generator pypi package built on provider-agnostic LiteLLM. The generator lets practitioners mint their own tiny packs in any language, domain, or difficulty, while ten ready-made packs already cover Arabic, Chinese, French, German, Japanese, Korean, Portuguese, Russian, Spanish, and Turkish. Every dataset ships with Croissant metadata and plug-and-play files for OpenAI-Evals, LangChain, and standard CI tools, so teams can drop deterministic micro-benchmarks directly into pull-request gates, prompt-engineering loops, and production dashboards without touching GPU budgets. A complete TQB++ run adds only a few seconds to pipeline latency yet reliably flags prompt-template errors, tokenizer drift, and fine-tuning side-effects long before full-scale suites like MMLU or BIG-Bench would finish configuring. The entire framework is released to accelerate continuous, resource-efficient quality assurance across the generative-AI ecosystem.

PaperPDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

vincentkoc/tiny_qa_benchmark_pp officialmentioned in papermentioned on GitHub report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Dataset GenerationLarge Language ModelMMLUPrompt EngineeringTinyQA Benchmark++

1 archive task tag without a task page not shown.

Datasets

Introduced by this paper, per the archive.

TQBA++

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
TinyQA Benchmark++ tinyqabenchmark_core-en gemma-3-4b Exact Match 86.5 #1 of 8 Archive leaderboard report
TinyQA Benchmark++ tinyqabenchmark_core-en mistral-24b-instruct Exact Match 84.6 #2 of 8 Archive leaderboard report
TinyQA Benchmark++ tinyqabenchmark_core-en llama-3.2-3b-instruct Exact Match 84.6 #3 of 8 Archive leaderboard report
TinyQA Benchmark++ tinyqabenchmark_core-en ministral-8b Exact Match 80.8 #4 of 8 Archive leaderboard report
TinyQA Benchmark++ tinyqabenchmark_core-en ministral-3b Exact Match 76.9 #5 of 8 Archive leaderboard report
TinyQA Benchmark++ tinyqabenchmark_core-en llama-3.2-1b-instruct Exact Match 53.8 #6 of 8 Archive leaderboard report
TinyQA Benchmark++ tinyqabenchmark_core-en mistral-7b-instruct Exact Match 50.0 #7 of 8 Archive leaderboard report
TinyQA Benchmark++ tinyqabenchmark_core-en gemma-3-12b Exact Macth 90.4 #8 of 8 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

SET

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections