Datasets › Speech Brown
Speech Brown
Dataset Summary
Speech Brown is a comprehensive, synthetic, and diverse paired speech-text dataset in 15 categories, covering a wide range of topics from fiction to religion. This dataset consists of over 55,000 sentence-level samples.
To train the CLASP model, we created this dataset based on the Brown Corpus. The synthetic speech was generated using the NVIDIA Tacotron 2 text-to-speech model.
For more information about our proposed model, please refer to this paper. The dataset generation pipeline, along with code and usage instructions, is available on this GitHub page.
Dataset Statistics
- Total size: Approximately 30 GB.
- Number of samples: 55,173 pairs of speech and text.
- Average tokens per sample: 19.00.
- Maximum tokens in a sample: 48.
- Average characters per sample: 96.72.
- Number of unique tokens: 50,667
- Categories: 15 categories consist of
adventure,belles_lettres,editorial,fiction,government,hobbies,humor,learned,lore,mystery,news,religion,reviews,romance,science_fiction.
Dataset Structure
To ensure ease of use, the dataset is partitioned into 10 parts. Each part can be used independently if it meets the requirements of your task and model.
Metadata Files
- global_metadata: A JSON file containing metadata for all 55,173 samples.
- localized_metadata: A JSON file containing metadata for all samples, categorized into the 10 dataset partitions.
Metadata Fields
- id: The unique identifier for the sample.
- audio_file_path: The file path for the audio in the dataset.
- category: The category of the sample's text.
- text: The corresponding text of the audio file.
Usage Instructions
To use this dataset, download the parts and metadata files as follows:
Option 1: Manual Download
Visit the dataset repository and download all dataset_partX.zip files and the global_metadata.json file.
Option 2: Programmatic Download
Use the huggingface_hub library to download the files programmatically:
from huggingface_hub import hf_hub_download
from zipfile import ZipFile
import os
import json
# Download dataset parts
zip_file_path1 = hf_hub_download(repo_id="llm-lab/SpeechBrown", filename="dataset_part1.zip", repo_type="dataset")
zip_file_path2 = hf_hub_download(repo_id="llm-lab/SpeechBrown", filename="dataset_part2.zip", repo_type="dataset")
# Download other parts...
# Download metadata
metadata_file_path = hf_hub_download(repo_id="llm-lab/SpeechBrown", filename="global_metadata.json", repo_type="dataset")
for i in range(1, 11):
with ZipFile(f'dataset_part{i}.zip', 'r') as zip_ref:
zip_ref.extractall(f'dataset_part{i}')
os.remove(f'dataset_part{i}.zip')
with open('global_metadata.json', 'r') as f:
metadata = json.load(f)
metadata.keys()
Citations
If you find our paper, code, data, or models useful, please cite the paper:
@misc{abootorabi2024claspcontrastivelanguagespeechpretraining,
title={CLASP: Contrastive Language-Speech Pretraining for Multilingual Multimodal Information Retrieval},
author={Mohammad Mahdi Abootorabi and Ehsaneddin Asgari},
year={2024},
eprint={2412.13071},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2412.13071},
}
Contact
If you have questions, please email mahdi.abootorabi2@gmail.com or asgari@berkeley.edu.
Benchmarks archive 2025-07-28
No leaderboard in the archive resolves to this dataset.
Papers archive 2025-07-28
No paper in the archive has a leaderboard row on this dataset; the archive counts 1 paper for it but never published that list.
Dataset loaders archive 2025-07-28
No loader listed in the archive.
Tasks archive 2025-07-28
License archive 2025-07-28
Modalities archive 2025-07-28
Languages archive 2025-07-28
Variants archive 2025-07-28
- Speech Brown
1 variant name, as the archive lists them.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections