{"url":"/dataset/se-pef","name":"SE-PEF","full_name":"Stack Exchange - Personalized Expert Finding","description_markdown":"Click to add a brief description of the dataset (Markdown and LaTeX enabled).\r\n\r\nProvide:\r\n\r\n* a high-level explanation of the dataset characteristics\r\n* explain motivations and summary of its coThe problem of personalization in Information Retrieval has been under study for a long time. A well know issue related to this task is the lack of publicly available datasets that can support a comparative evaluation of personalised search systems. To contribute in this respect, this paper introduces SE-PEF (StackExchange - Personalized Expert Finding), a resource useful for designing and evaluating personalized models related to the task of Expert Finding (EF).\r\nThe contributed dataset  includes more than  250k queries and 565k answers from 3,306 experts, which are annotated with a rich set of features modeling the social interactions among the users of a popular cQA platform.\r\nThe results of the preliminary experiments conducted show the appropriateness of SE-PEF to evaluate and to train effective EF models.ntent\r\n* potential use cases of the dataset","description_withheld":null,"homepage":"https://doi.org/10.5281/zenodo.8332748","introduced_date":"2023-09-10","introduced_date_note":null,"introduced_by":{"paper":"/paper/se-pef-a-resource-for-personalized-expert","title":"SE-PEF: a Resource for Personalized Expert Finding","first_author":"Pranav Kasela","url":null},"license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/legalcode"},"modalities":[{"name":"Texts","url":"/datasets/modality/texts"}],"tasks":[{"name":"Entity Retrieval","url":"/task/entity-retrieval","datasets_with_task":"/datasets/task/entity-retrieval"}],"languages":[{"name":"English","url":"/datasets/language/english"}],"variants":["SE-PEF"],"data_loaders":[],"num_papers_in_archive":1,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}