{"url":"/dataset/phd2","name":"PHD²","full_name":"Personalized Highlight Detection Dataset","description_markdown":"The dataset contains information on what video segments a specific user considers a highlight. Having this kind of data allows for strong personalization models, as specific examples of what a user is interested in help models obtain a fine-grained understanding of that specific user.\r\n\r\nThe data consists of YouTube videos, from which gifs.com users manually extracted their highlights, by creating GIFs from a segment of the full video. Thus, the dataset is similar to PHD-GIFS, with two major differences.\r\n\r\n- Each selection is associated with a user, which is what allows personalization.\r\n- instead of visual matching to find the position in the video from which a GIF was selected, PHD-GIFS uses the timestamps. Thus, the ground truth is free from any alignment errors.\r\n\r\nThe training set contains highlights from 12,972 users. The test set contains highlights from 850 users. \r\n\r\nSource: [Personalized Highlight Detection Dataset](https://github.com/gifs/personalized-highlights-dataset)","description_withheld":null,"homepage":"https://github.com/gifs/personalized-highlights-dataset","introduced_date":"2018-04-18","introduced_date_note":null,"introduced_by":{"paper":"/paper/phd-gifs-personalized-highlight-detection-for","title":"PHD-GIFs: Personalized Highlight Detection for Automatic GIF Creation","first_author":"Ana García del Molino","url":null},"license":null,"modalities":[{"name":"Videos","url":"/datasets/modality/videos"}],"tasks":[],"languages":[],"variants":["PHD²"],"data_loaders":[{"repo":"https://github.com/gifs/personalized-highlights-dataset","url":"https://github.com/gifs/personalized-highlights-dataset","frameworks":[]}],"num_papers_in_archive":4,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}