Datasets › Multi Lingual Bug Reports
Multi Lingual Bug Reports
Dataset Description
The dataset used in this study comprises bug reports extracted from the Visual Studio Code GitHub repository, specifically focusing on those labeled with the english-please tag. This label indicates that the original submission was written in a language other than English, providing a clear signal for multilingual content. The dataset spans a five-year period (March 2019--June 2024), ensuring a diverse representation of bug types, user environments, and technical contexts.
Characteristics
The dataset contains 1,381 multilingual bug reports, each consisting of: - The original bug report written in a non-English language. - A translated version in English. - Metadata such as issue number, creation date, labels, and status. - Categorization into functional, UI, and performance-related issues based on the content.
Motivation & Summary
This dataset is motivated by the need to improve multilingual bug tracking and translation evaluation. Given the increasing globalization of software development, developers and QA teams frequently encounter bug reports in languages they do not understand. By providing a structured corpus of translated bug reports, this dataset facilitates: - Comparative translation evaluation (e.g., ChatGPT vs AWS Translate vs DeepL). - Linguistic analysis of technical bug reporting across different languages. - Insights into common software issues encountered by diverse users. - Improving multilingual issue tracking through automated labeling and categorization.
Potential Use Cases
This dataset can be beneficial for various research and development applications, including: - Machine Translation Benchmarking: Evaluating the performance of translation models in a technical domain. - Natural Language Processing (NLP) Tasks: Training classifiers to categorize bug reports based on their content. - Software Engineering Research: Understanding trends in bug reporting, issue resolution, and localization challenges. - Automated Bug Triage: Developing AI-driven solutions for assigning and prioritizing bug reports in multilingual repositories.
Benchmarks archive 2025-07-28
All 1 leaderboard whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.
| First row (archive order) | Paper | Code | ||||
|---|---|---|---|---|---|---|
| Machine Translation | Multi Lingual Bug Reports | ChatGPT BERTScore 79 | English Please: Evaluating Machine Translation with... | av9ash/English-Please | 1 | Compare |
Papers archive 2025-07-28
1 shown of 1 paper with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 1. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.
| Date | Samples run Syntology | |||
|---|---|---|---|---|
| English Please: Evaluating Machine Translation with Large Language Models for Multilingual Bug Reports | 1 | 1 | 20 Feb 2025 | not harvested |
Dataset loaders archive 2025-07-28
No loader listed in the archive.
Tasks archive 2025-07-28
License archive 2025-07-28
Modalities archive 2025-07-28
Languages archive 2025-07-28
Variants archive 2025-07-28
- Multi Lingual Bug Reports
1 variant name, as the archive lists them.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections