{"url":"/dataset/dataseeds-ai-sample-dataset-dsd","name":"DataSeeds.AI-Sample-Dataset-DSD","full_name":null,"description_markdown":"Dataset Summary\r\nThe DataSeeds.AI Sample Dataset (DSD) is a high-fidelity, human-curated computer vision-ready dataset comprised of 7,772 peer-ranked, fully annotated photographic images, 350,000+ words of descriptive text, and comprehensive metadata. While the DSD is being released under an open source license, a sister dataset of over 10,000 fully annotated and segmented images is available for immediate commercial licensing, and the broader GuruShots ecosystem contains over 100 million images in its catalog.\r\n\r\nEach image includes multi-tier human annotations and semantic segmentation masks. Generously contributed to the community by the GuruShots photography platform, where users engage in themed competitions, the DSD uniquely captures aesthetic preference signals and high-quality technical metadata (EXIF) across an expansive diversity of photographic styles, camera types, and subject matter. The dataset is optimized for fine-tuning and evaluating multimodal vision-language models, especially in scene description and stylistic comprehension tasks.\r\n\r\nTechnical Report - Peer-Ranked Precision: Creating a Foundational Dataset for Fine-Tuning Vision Models from DataSeeds' Annotated Imagery\r\nGithub Repo - Access the complete weights and code which were used to evaluate the DSD -- https://github.com/DataSeeds-ai/DSD-finetune-blip-llava\r\nThis dataset is ready for commercial/non-commercial use.\r\nDataset Structure\r\nSize: 7,772 images (7,010 train, 762 validation)\r\nFormat: Apache Parquet files for metadata, with images in JPG format\r\nTotal Size: ~4.1GB\r\nLanguages: English (annotations)\r\nAnnotation Quality: All annotations were verified through a multi-tier human-in-the-loop process","description_withheld":null,"homepage":"https://huggingface.co/datasets/Dataseeds/DataSeeds.AI-Sample-Dataset-DSD#dataset-summary","introduced_date":null,"introduced_date_note":null,"introduced_by":null,"license":null,"modalities":[],"tasks":[],"languages":[],"variants":["DataSeeds.AI-Sample-Dataset-DSD"],"data_loaders":[],"num_papers_in_archive":1,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}