{"url":"/dataset/ivrit-ai","name":"ivrit.ai","full_name":"database of Hebrew audio and text content.","description_markdown":"ivrit.ai is a database of Hebrew audio and text content.\r\n\r\n**audio-base** contains the raw, unprocessed sources. About 13,000 hours of speech audio.\r\n\r\n**audio-vad** contains audio snippets generated by applying Silero VAD (https://github.com/snakers4/silero-vad) to the base dataset. v1 data is generated using silero-vad's default parameters. v2 data is generated using min_speech_duration_ms=2000 (milliseconds), and max_speech_duration_s=30 (seconds).\r\n\r\n**audio-transcripts** contain transcriptions for each snippet in the audio-vad dataset.\r\n\r\nYou can find the full list of sources in this dataset under https://www.ivrit.ai/en/credits.","description_withheld":null,"homepage":"https://www.ivrit.ai/en","introduced_date":"2023-07-17","introduced_date_note":null,"introduced_by":{"paper":"/paper/ivrit-ai-a-comprehensive-dataset-of-hebrew","title":"ivrit.ai: A Comprehensive Dataset of Hebrew Speech for AI Research and Development","first_author":"Yanir Marmor","url":null},"license":{"name":"ivrit.ai license","url":"https://www.ivrit.ai/en/the-license/"},"modalities":[],"tasks":[],"languages":[],"variants":[],"data_loaders":[],"num_papers_in_archive":2,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}