{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/tiny-qa-benchmark-ultra-lightweight-synthetic","title":"Tiny QA Benchmark++: Ultra-Lightweight, Synthetic Multilingual Dataset Generation & Smoke-Tests for Continuous LLM Evaluation","arxiv_id":"2505.12058","date":"2025-05-17","proceeding":null,"authors":["Vincent Koc"],"abstract":"Tiny QA Benchmark++ (TQB++) presents an ultra-lightweight, multilingual smoke-test suite designed to give large-language-model (LLM) pipelines a unit-test style safety net dataset that runs in seconds with minimal cost. Born out of the tight feedback-loop demands building the Comet Opik prompt-optimization SDK, where waiting on heavyweight benchmarks breaks developer flow. TQB++ couples a 52-item English gold set (less than 20 kB) with a tiny synthetic-data generator pypi package built on provider-agnostic LiteLLM. The generator lets practitioners mint their own tiny packs in any language, domain, or difficulty, while ten ready-made packs already cover Arabic, Chinese, French, German, Japanese, Korean, Portuguese, Russian, Spanish, and Turkish. Every dataset ships with Croissant metadata and plug-and-play files for OpenAI-Evals, LangChain, and standard CI tools, so teams can drop deterministic micro-benchmarks directly into pull-request gates, prompt-engineering loops, and production dashboards without touching GPU budgets. A complete TQB++ run adds only a few seconds to pipeline latency yet reliably flags prompt-template errors, tokenizer drift, and fine-tuning side-effects long before full-scale suites like MMLU or BIG-Bench would finish configuring. The entire framework is released to accelerate continuous, resource-efficient quality assurance across the generative-AI ecosystem.","url_abs":"https://arxiv.org/abs/2505.12058v1","url_pdf":"https://arxiv.org/pdf/2505.12058v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"tiny-qa-benchmark-ultra-lightweight-synthetic","repo_url":"https://github.com/vincentkoc/tiny_qa_benchmark_pp","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"none","reach":null}],"tasks":[{"task_slug":"dataset-generation","task_name":"Dataset Generation"},{"task_slug":null,"task_name":"GPU"},{"task_slug":"large-language-model","task_name":"Large Language Model"},{"task_slug":"mmlu","task_name":"MMLU"},{"task_slug":"prompt-engineering","task_name":"Prompt Engineering"},{"task_slug":"tinyqa-benchmark","task_name":"TinyQA Benchmark++"}],"methods":[{"method_slug":"set","method_name":"SET"}],"datasets_introduced":[{"slug":"tqba","name":"TQBA++","full_name":"Tiny QA Benchmark++"}],"methods_introduced":[],"results":[{"leaderboard":"/sota/tinyqa-benchmark-on-tinyqabenchmark-core-en","task":"TinyQA Benchmark++","dataset":"tinyqabenchmark_core-en","model":"gemma-3-4b","rank_in_archive_order":1,"of":8,"metrics":{"Exact Match":"86.5"},"uses_additional_data":false},{"leaderboard":"/sota/tinyqa-benchmark-on-tinyqabenchmark-core-en","task":"TinyQA Benchmark++","dataset":"tinyqabenchmark_core-en","model":"mistral-24b-instruct","rank_in_archive_order":2,"of":8,"metrics":{"Exact Match":"84.6"},"uses_additional_data":false},{"leaderboard":"/sota/tinyqa-benchmark-on-tinyqabenchmark-core-en","task":"TinyQA Benchmark++","dataset":"tinyqabenchmark_core-en","model":"llama-3.2-3b-instruct","rank_in_archive_order":3,"of":8,"metrics":{"Exact Match":"84.6"},"uses_additional_data":false},{"leaderboard":"/sota/tinyqa-benchmark-on-tinyqabenchmark-core-en","task":"TinyQA Benchmark++","dataset":"tinyqabenchmark_core-en","model":"ministral-8b","rank_in_archive_order":4,"of":8,"metrics":{"Exact Match":"80.8"},"uses_additional_data":false},{"leaderboard":"/sota/tinyqa-benchmark-on-tinyqabenchmark-core-en","task":"TinyQA Benchmark++","dataset":"tinyqabenchmark_core-en","model":"ministral-3b","rank_in_archive_order":5,"of":8,"metrics":{"Exact Match":"76.9"},"uses_additional_data":false},{"leaderboard":"/sota/tinyqa-benchmark-on-tinyqabenchmark-core-en","task":"TinyQA Benchmark++","dataset":"tinyqabenchmark_core-en","model":"llama-3.2-1b-instruct","rank_in_archive_order":6,"of":8,"metrics":{"Exact Match":"53.8"},"uses_additional_data":false},{"leaderboard":"/sota/tinyqa-benchmark-on-tinyqabenchmark-core-en","task":"TinyQA Benchmark++","dataset":"tinyqabenchmark_core-en","model":"mistral-7b-instruct","rank_in_archive_order":7,"of":8,"metrics":{"Exact Match":"50.0"},"uses_additional_data":false},{"leaderboard":"/sota/tinyqa-benchmark-on-tinyqabenchmark-core-en","task":"TinyQA Benchmark++","dataset":"tinyqabenchmark_core-en","model":"gemma-3-12b","rank_in_archive_order":8,"of":8,"metrics":{"Exact Macth":"90.4"},"uses_additional_data":false}],"syntology":{"syntology_url":null,"atlas_url":"https://app.syntology.ai/?focus=2505.12058","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}