{"url":"/dataset/spgispeech","name":"SPGISpeech","full_name":null,"description_markdown":"SPGISpeech (pronounced “speegie-speech”) is a large-scale transcription dataset, freely available for academic research. SPGISpeech is a collection of 5,000 hours of professionally-transcribed financial audio. Contrary to previous transcription datasets, SPGISpeech contains global english accents, strongly varying audio quality as well as both spontaneous and presentation style speech. The transcripts have each been cross-checked by multiple professional editors for high accuracy and are fully formatted including sentence structure and capitalization.\r\n\r\nSPGISpeech consists of 5,000 hours of recorded company earnings calls and associated manual transcription text. The original calls were split based on silences into slices ranging from 5 to 15 seconds to allow easy training of a speech recognition system. The format of each WAV file is single channel, 16kHz, 16 bit audio.\r\n\r\nTranscription text represents the output of several stages of manual post-processing. As such, the text contains polished English orthography following a detailed style guide, including proper casing, punctuation, and denormalized non-standard words such as numbers or acronyms, making SPGISpeech suited for training fully formatted end-to-end models.\r\n\r\nIn general, the transcriptions aim at professional utility rather than linguistic fidelity, and the correspondence between verbatim speech and finalized text is therefore not exact, resulting in the occasional purposeful omission of meeting operator instructions or certain verbal pleasantries.","description_withheld":null,"homepage":"https://datasets.kensho.com/datasets/scribe","introduced_date":"2021-04-05","introduced_date_note":null,"introduced_by":{"paper":"/paper/spgispeech-5000-hours-of-transcribed","title":"SPGISpeech: 5,000 hours of transcribed financial audio for fully formatted end-to-end speech recognition","first_author":"Patrick K. O'Neill","url":null},"license":{"name":"Custom (non-commercial)","url":null},"modalities":[{"name":"Speech","url":"/datasets/modality/speech"}],"tasks":[{"name":"Speech Recognition","url":"/task/speech-recognition","datasets_with_task":"/datasets/task/speech-recognition"},{"name":"Automatic Speech Recognition","url":"/task/automatic-speech-recognition-2","datasets_with_task":"/datasets/task/automatic-speech-recognition-2"}],"languages":[{"name":"English","url":"/datasets/language/english"}],"variants":["SPGISpeech"],"data_loaders":[],"num_papers_in_archive":16,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[{"leaderboard":"/sota/speech-recognition-on-spgispeech","task":"Speech Recognition","dataset_variant":"SPGISpeech","rows":3,"metrics":["Word Error Rate (WER)"],"first_row_in_archive_order":{"model":"Icefall - zipformer transducer","paper":null,"metrics":{"Word Error Rate (WER)":"2.35"},"code_links":[]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"}],"papers_with_a_benchmark_row":[{"paper":"/paper/fast-conformer-with-linearly-scalable","title":"Fast Conformer with Linearly Scalable Attention for Efficient Speech Recognition","date":"2023-05-08","rows_on_this_dataset":1,"code_links":0,"syntology":null},{"paper":"/paper/spgispeech-5000-hours-of-transcribed","title":"SPGISpeech: 5,000 hours of transcribed financial audio for fully formatted end-to-end speech recognition","date":"2021-04-05","rows_on_this_dataset":1,"code_links":1,"syntology":null}],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}