Papers › Tiny QA Benchmark++: Ultra-Lightweight, Synthetic Multilingual Dataset Generation &...
Tiny QA Benchmark++: Ultra-Lightweight, Synthetic Multilingual Dataset Generation & Smoke-Tests for Continuous LLM Evaluation
Vincent Koc
Tiny QA Benchmark++ (TQB++) presents an ultra-lightweight, multilingual smoke-test suite designed to give large-language-model (LLM) pipelines a unit-test style safety net dataset that runs in seconds with minimal cost. Born out of the tight feedback-loop demands building the Comet Opik prompt-optimization SDK, where waiting on heavyweight benchmarks breaks developer flow. TQB++ couples a 52-item English gold set (less than 20 kB) with a tiny synthetic-data generator pypi package built on provider-agnostic LiteLLM. The generator lets practitioners mint their own tiny packs in any language, domain, or difficulty, while ten ready-made packs already cover Arabic, Chinese, French, German, Japanese, Korean, Portuguese, Russian, Spanish, and Turkish. Every dataset ships with Croissant metadata and plug-and-play files for OpenAI-Evals, LangChain, and standard CI tools, so teams can drop deterministic micro-benchmarks directly into pull-request gates, prompt-engineering loops, and production dashboards without touching GPU budgets. A complete TQB++ run adds only a few seconds to pipeline latency yet reliably flags prompt-template errors, tokenizer drift, and fine-tuning side-effects long before full-scale suites like MMLU or BIG-Bench would finish configuring. The entire framework is released to accelerate continuous, resource-efficient quality assurance across the generative-AI ecosystem.
In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.
Code
Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.
Code Syntology ran Syntology
Not run by Syntology. Nothing on this page verifies that the listed code works.
Tasks
1 archive task tag without a task page not shown.
Datasets
Introduced by this paper, per the archive.
Results from the paper archive 2025-07-28
| Task | Dataset | Model | Metric | Value | Rank at snapshot | Leaderboard | Report |
|---|---|---|---|---|---|---|---|
| TinyQA Benchmark++ | tinyqabenchmark_core-en | gemma-3-4b | Exact Match | 86.5 | #1 of 8 | Archive leaderboard | report |
| TinyQA Benchmark++ | tinyqabenchmark_core-en | mistral-24b-instruct | Exact Match | 84.6 | #2 of 8 | Archive leaderboard | report |
| TinyQA Benchmark++ | tinyqabenchmark_core-en | llama-3.2-3b-instruct | Exact Match | 84.6 | #3 of 8 | Archive leaderboard | report |
| TinyQA Benchmark++ | tinyqabenchmark_core-en | ministral-8b | Exact Match | 80.8 | #4 of 8 | Archive leaderboard | report |
| TinyQA Benchmark++ | tinyqabenchmark_core-en | ministral-3b | Exact Match | 76.9 | #5 of 8 | Archive leaderboard | report |
| TinyQA Benchmark++ | tinyqabenchmark_core-en | llama-3.2-1b-instruct | Exact Match | 53.8 | #6 of 8 | Archive leaderboard | report |
| TinyQA Benchmark++ | tinyqabenchmark_core-en | mistral-7b-instruct | Exact Match | 50.0 | #7 of 8 | Archive leaderboard | report |
| TinyQA Benchmark++ | tinyqabenchmark_core-en | gemma-3-12b | Exact Macth | 90.4 | #8 of 8 | Archive leaderboard | report |
Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.
Methods
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections