{"url":"/dataset/bigos","name":"BIGOS","full_name":"Benchmark Intended Grouping of Open Speech","description_markdown":"The **Benchmark Intended Grouping of Open Speech (BIGOS)** is a novel corpus specifically designed for **Polish Automatic Speech Recognition (ASR) systems**. This initial version of the benchmark comprises **1,900 audio recordings** from **71 distinct speakers**, sourced from **10 publicly available speech corpora**¹²³.\r\n\r\nHere are some key points about BIGOS:\r\n\r\n1. **Purpose**: BIGOS aims to facilitate systematic benchmarking and tracking of Polish ASR systems over time across a diverse range of publicly available corpora.\r\n2. **Evaluation**: The benchmark evaluates both **proprietary** and **open-source ASR systems** on a diverse set of recordings and their corresponding original transcriptions.\r\n3. **Findings**:\r\n    - The performance of the latest open-source models is **comparable** to that of more established commercial services.\r\n    - Model size significantly influences system accuracy.\r\n    - Scenarios involving highly specialized or spontaneous speech show a decrease in accuracy.\r\n4. **Challenges**: The study discusses the challenges of using public datasets for ASR evaluation and the limitations based on this inaugural benchmark.\r\n5. **Availability**: The BIGOS corpus and associated tools are **publicly available** for replication and customization of the benchmark.\r\n\r\nIn summary, BIGOS provides a valuable resource for advancing Polish ASR research and improving the quality of speech recognition systems in the Polish language¹.\r\nCheck https://www.semanticscholar.org/paper/BIGOS-Benchmark-Intended-Grouping-of-Open-Speech-Junczyk\r\n\r\nSource: Conversation with Bing, 3/18/2024\r\n(1) BIGOS - Benchmark Intended Grouping of Open Speech Corpora for Polish .... https://annals-csis.org/proceedings/2023/pliks/1609.pdf.\r\n(2) Annals of Computer Science and Information Systems, Volume 35. https://annals-csis.org/proceedings/2023/drp/1609.html.\r\n(3) BIGOS - Benchmark Intended Grouping of Open Speech Corpora for Polish .... https://ieeexplore.ieee.org/abstract/document/10306084.","description_withheld":null,"homepage":"https://huggingface.co/datasets/amu-cai/pl-asr-bigos-v2","introduced_date":null,"introduced_date_note":null,"introduced_by":null,"license":null,"modalities":[],"tasks":[],"languages":[],"variants":["BIGOS"],"data_loaders":[],"num_papers_in_archive":0,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-25T09:33:49+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}