{"url":"/dataset/chatgpt-software-testing","name":"ChatGPT-software-testing","full_name":"ChatGPT Software Testing","description_markdown":"## Dataset Description\r\nOur dataset contains questions from a well-known software testing book **Introduction to Software Testing 2nd Edition** by Ammann and Offutt. \r\nWe use all the text-book questions in Chapters 1 to 5 that have solutions available on the book’s official website. \r\n\r\nOur dataset contains 40 such questions from these five chapters. 31 questions out of the 40 are multipart questions and the rest 9 are independent.\r\nThis tool generates responses from the ChatGPT automatically for these questions. All of these questions are run in both shared and separate context.\r\nMore information about the contexts can be found below.\r\n\r\n### Combined.xlsx\r\nContains all the questions & answers for 3 iterations of both shared and separate contexts. Contains labels for answers and explanations given by ChatGPT.\r\n\r\n### Combined_pair.xlsx\r\nContains the same data as **combined.xlsx** except for the questions that are independent i.e., not part of a multipart question.\r\n\r\n### Combined_analysis.xlsx\r\nContains the result and analysis of the four research questions. Besides, it contains various illustrations for the results.\r\n\r\n### Combined-temp.xlsx\r\nQuestions with missing shared contexts are replaced with the answers for the separate context to easily fetch the data for \r\nRQ2 & RQ3 from a single column.\r\n\r\n### examples folders\r\nContains examples of some interesting response categories.\r\n\r\n### Case Study.pdf\r\nContains the following analysis:\r\n- When responses are likely to be incorrect?\r\n- What are the reasons for being incorrect?\r\n- Can we fix it with prompt engineering?\r\n- Case studies with actual examples.\r\n\r\n\r\n## Separate Context Query\r\nIn separate context queries, we treat each of the 31 multipart questions as an independent question.\r\nEach sub-question is asked in a separate chat thread.\r\nCombining with the nine independent questions, a total of 40 questions are asked for each run. To evaluate the consistency in\r\nChatGPT’s responses, we collect a total of three runs for each question, which results in a total of 120 responses from ChatGPT.\r\n\r\n## Shared Context Query\r\nOur dataset contains six questions that contain total 31 multipart questions or sub-questions and nine questions that do not. \r\nThese six sub-questions are asked in a chat thread that are shared with other sub-questions as long as the sub-questions \r\nrefer to the same code or scenario.","description_withheld":null,"homepage":"https://github.com/sajedjalil/Study-on-ChatGPT/tree/main/dataset","introduced_date":"2023-01-31","introduced_date_note":null,"introduced_by":null,"license":{"name":"MIT","url":"https://github.com/sajedjalil/Study-on-ChatGPT/blob/main/LICENSE"},"modalities":[{"name":"Texts","url":"/datasets/modality/texts"}],"tasks":[{"name":"Chatbot","url":"/task/chatbot","datasets_with_task":"/datasets/task/chatbot"}],"languages":[{"name":"English","url":"/datasets/language/english"}],"variants":["ChatGPT-software-testing"],"data_loaders":[],"num_papers_in_archive":0,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}