{"url":"/dataset/synthetic-product-desirability-datasets-for","name":"Synthetic Product Desirability Datasets for Sentiment Analysis Testing","full_name":null,"description_markdown":"Overview:\r\nThis collection contains three synthetic datasets produced by gpt-4o-mini for sentiment analysis and PDT (Product Desirability Toolkit) testing. Each dataset contains 1000 hypothetical software product reviews with the aim to produce a diversity of sentiment and text. The datasets were created as part of the research described in:\r\n\r\nJ. D. Hastings, S. Weitl-Harms, J. Doty, Z. L. Myers, and W. Thompson, “Utilizing Large Language Models to Synthesize Product Desirability Datasets,” in Proceedings of the 2024 IEEE International Conference on Big Data (BigData-24), Workshop on Large Language and Foundation Models (WLLFM-24), Dec. 2024. arXiv: 2411.13485 [cs.CL].\r\n\r\nBriefly, each row in the datasets was produced as follows:\r\n1) Word+Review: The LLM selected a word and synthesized a review that would align with a random target sentiment.\r\n2) Review+Word: The LLM produced a review to align with the target sentiment score, and then selected a word appropriate for the review.\r\n3) Supply-Word: A word was supplied to the LLM which was then scored, and a review was produced to align with that score.\r\n\r\nFor sentiment analysis and PDT testing, the two columns of main interest across the datasets are likely 'Selected Word' and 'Hypothetical Review'.\r\n\r\nLicense:\r\nThis data is licensed under the CC Attribution 4.0 international license, and may be taken and used freely with credit given. Cite as:\r\n\r\nHastings, J., Weitl-Harms, S., Doty, J., Myers, Z., & Thompson, W. (2024). Synthetic Product Desirability Datasets for Sentiment Analysis Testing (1.0.0). Zenodo. https://doi.org/10.5281/zenodo.14188456","description_withheld":null,"homepage":"https://doi.org/10.5281/zenodo.14188456","introduced_date":"2024-11-19","introduced_date_note":null,"introduced_by":{"paper":"/paper/utilizing-large-language-models-to-synthesize","title":"Utilizing Large Language Models to Synthesize Product Desirability Datasets","first_author":"John D. Hastings","url":null},"license":{"name":"CC Attribution 4.0 nternationalI","url":null},"modalities":[{"name":"Texts","url":"/datasets/modality/texts"}],"tasks":[{"name":"Sentiment Analysis","url":"/task/sentiment-analysis","datasets_with_task":"/datasets/task/sentiment-analysis"},{"name":"Sentiment Analysis (Product + User)","url":"/task/sentiment-analysis-product-user","datasets_with_task":"/datasets/task/sentiment-analysis-product-user"}],"languages":[{"name":"English","url":"/datasets/language/english"}],"variants":["Synthetic Product Desirability Datasets for Sentiment Analysis Testing"],"data_loaders":[],"num_papers_in_archive":1,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-25T09:33:49+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}