{"url":"/dataset/sharegpt4video","name":"ShareGPT4Video","full_name":null,"description_markdown":"The **ShareGPT4Video** dataset is a large-scale resource designed to improve video understanding and generation¹. It features **1.2 million highly descriptive captions**⁴ for video clips, surpassing existing datasets in diversity and information content⁴. The captions cover a wide range of aspects, including world knowledge, object properties, spatial relationships, and aesthetic evaluations⁴.\r\n\r\nThe dataset includes detailed captions of **40K videos** generated by **GPT-4V**¹ and **4.8M videos** generated by **ShareCaptioner-Video**¹. The videos are sourced from YouTube and other user-uploaded video websites, and they cover a variety of scenarios, such as human activities and auto-driving¹.\r\n\r\nThe ShareGPT4Video dataset also provides a basis for the **ShareCaptioner-Video**, an exceptional video captioner capable of efficiently generating high-quality captions for videos of a wide range of resolution, aspect ratio, and duration¹.\r\n\r\nFor example, the dataset includes a detailed caption of a video documenting a meticulous meal preparation by an individual with tattooed forearms¹. The caption describes the individual's actions in detail, from slicing a cucumber to mixing the dressing and adding croutons to the salad¹.\r\n\r\nIn addition to its use in research, the ShareGPT4Video dataset has been used to train the **sharegpt4video-8b** model, an open-source video chatbot². This model was trained on open-source video instruction data and is primarily intended for researchers and hobbyists in computer vision, natural language processing, machine learning, and artificial intelligence².\r\n\r\n(1) arXiv:2406.04325v1 [cs.CV] 6 Jun 2024. https://arxiv.org/pdf/2406.04325.\r\n(2) ShareGPT4V: Improving Large Multi-Modal Models with Better Captions. https://arxiv.org/abs/2311.12793.\r\n(3) Lin-Chen/sharegpt4video-8b · Hugging Face. https://huggingface.co/Lin-Chen/sharegpt4video-8b.\r\n(4) ShareGPT4Video: Improving Video Understanding and Generation with .... https://www.aimodels.fyi/papers/arxiv/sharegpt4video-improving-video-understanding-generation-better-captions.\r\n(5) GitHub - ShareGPT4Omni/ShareGPT4Video: An official implementation of .... https://github.com/ShareGPT4Omni/ShareGPT4Video.\r\n(6) undefined. https://sharegpt4video.github.io/.","description_withheld":null,"homepage":"https://github.com/ShareGPT4Omni/ShareGPT4Video","introduced_date":"2024-06-06","introduced_date_note":null,"introduced_by":{"paper":"/paper/sharegpt4video-improving-video-understanding","title":"ShareGPT4Video: Improving Video Understanding and Generation with Better Captions","first_author":"Lin Chen","url":null},"license":null,"modalities":[],"tasks":[],"languages":[],"variants":["ShareGPT4Video"],"data_loaders":[],"num_papers_in_archive":28,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}