{"url":"/dataset/videodb-s-ocr-benchmark-public-collection","name":"VideoDB's OCR Benchmark Public Collection","full_name":"VideoDB's OCR Benchmark Public Collection","description_markdown":"## Dataset Introduction\r\n\r\nThis dataset leverages VideoDB's Public Collection to offer a diverse range of videos featuring text-containing scenes. It spans multiple categories—ranging from finance and legal documents to software UI elements and handwritten notes—ensuring a broad representation of real-world text appearances. Each video is annotated with frame indexes to facilitate consistent and reproducible OCR benchmarks. Currently, the dataset includes over 25 curated videos, yielding thousands of extracted frames that present a variety of text-related challenges.\r\n\r\n### Key Features\r\n\r\n1. **Diverse Text Genres**  \r\n    - **Finance/Business:** Includes news tickers and stock market visuals where text scrolls rapidly.  \r\n    - **Legal/Educational:** Features documents with formal language, diagrams, and formatted text.  \r\n    - **Software/Web Development/UI:** Shows on-screen code editors, browser windows, and other UI elements that test OCR's ability to handle varying font sizes and code snippets.  \r\n    - **Handwriting:** Encompasses both cursive and print handwriting on whiteboards or paper, capturing the challenges of style variability and penmanship.  \r\n    - **Miscellaneous/Other:** Covers signage, billboards, and everyday text in the wild.\r\n\r\n2. **Rich Annotation**  \r\n    - Each video includes **frame indexes** or **scene timestamps** to ensure consistent, reproducible extraction of text segments.  \r\n    - Ground truth (OCR text) is provided for thousands of extracted frames, facilitating **quantitative performance evaluations** (e.g., CER, WER).\r\n\r\n3. **Benchmark-Ready**  \r\n    - The dataset seamlessly integrates with the [ocr-benchmark repository](https://github.com/video-db/ocr-benchmark/) to streamline model evaluation.  \r\n    - Scripts are included for **frame extraction**, **automatic OCR comparison**, and **metric calculation** (CER, WER, accuracy).\r\n\r\n### How to Access\r\n\r\n- **VideoDB Public Collection ID:** `c-c0a2c223-e377-4625-94bf-910501c2a31c`  \r\n  Simply reference this ID within VideoDB to retrieve and review the videos.\r\n- **Ground Truth Files:** Located in the [`ocr_ground_truths`](https://github.com/video-db/ocr-benchmark/tree/main/ocr_ground_truths) directory. The JSON files map each frame to its corresponding textual annotations.\r\n\r\nFor detailed instructions on working with VideoDB Public Collections, please refer to the [official documentation](https://docs.videodb.io/public-collections-102).\r\n\r\n### Licensing and Usage\r\n\r\n- **Usage Restrictions:** The videos are publicly accessible for research and educational use.  \r\n- **Attribution:** Please cite the [Benchmarking Vision-Language Models on Optical Character Recognition in Dynamic Video Environments](https://arxiv.org/abs/2502.06445) paper or repository if you use this dataset in your work.","description_withheld":null,"homepage":"https://docs.videodb.io/public-collections-102","introduced_date":"2025-02-10","introduced_date_note":null,"introduced_by":{"paper":"/paper/benchmarking-vision-language-models-on","title":"Benchmarking Vision-Language Models on Optical Character Recognition in Dynamic Video Environments","first_author":"Sankalp Nagaonkar","url":null},"license":null,"modalities":[{"name":"Images","url":"/datasets/modality/images"},{"name":"Videos","url":"/datasets/modality/videos"},{"name":"Texts","url":"/datasets/modality/texts"}],"tasks":[{"name":"Optical Character Recognition (OCR)","url":"/task/optical-character-recognition","datasets_with_task":"/datasets/task/optical-character-recognition"}],"languages":[{"name":"English","url":"/datasets/language/english"}],"variants":["VideoDB's OCR Benchmark Public Collection"],"data_loaders":[],"num_papers_in_archive":1,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[{"leaderboard":"/sota/optical-character-recognition-ocr-on-videodb","task":"Optical Character Recognition (OCR)","dataset_variant":"VideoDB's OCR Benchmark Public Collection","rows":5,"metrics":["Average Accuracy","Character Error Rate (CER)","Word Error Rate (WER)"],"first_row_in_archive_order":{"model":"GPT-4o","paper":"/paper/benchmarking-vision-language-models-on","metrics":{"Average Accuracy":"76.22","Character Error Rate (CER)":"0.2378","Word Error Rate (WER)":"0.5117"},"code_links":[{"title":"video-db/ocr-benchmark","url":"https://github.com/video-db/ocr-benchmark"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"}],"papers_with_a_benchmark_row":[{"paper":"/paper/benchmarking-vision-language-models-on","title":"Benchmarking Vision-Language Models on Optical Character Recognition in Dynamic Video Environments","date":"2025-02-10","rows_on_this_dataset":5,"code_links":1,"syntology":{"read_at":"2026-09-25T09:33:49+00:00","samples_harvested":5,"samples_ran":4,"samples_unverified":1,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}}],"syntology_totals":{"read_at":"2026-09-25T09:33:49+00:00","papers_with_samples":1,"samples_harvested":5,"samples_ran":4,"samples_unverified":1,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}