Datasets › Multi Lingual Bug Reports

Multi Lingual Bug Reports

Introduced by Avinash Patil et al. in English Please: Evaluating Machine Translation with Large Language Models for Multilingual Bug Reports20 Feb 2025 archive 2025-07-28

Dataset Description

The dataset used in this study comprises bug reports extracted from the Visual Studio Code GitHub repository, specifically focusing on those labeled with the english-please tag. This label indicates that the original submission was written in a language other than English, providing a clear signal for multilingual content. The dataset spans a five-year period (March 2019--June 2024), ensuring a diverse representation of bug types, user environments, and technical contexts.

Characteristics

The dataset contains 1,381 multilingual bug reports, each consisting of: - The original bug report written in a non-English language. - A translated version in English. - Metadata such as issue number, creation date, labels, and status. - Categorization into functional, UI, and performance-related issues based on the content.

Motivation & Summary

This dataset is motivated by the need to improve multilingual bug tracking and translation evaluation. Given the increasing globalization of software development, developers and QA teams frequently encounter bug reports in languages they do not understand. By providing a structured corpus of translated bug reports, this dataset facilitates: - Comparative translation evaluation (e.g., ChatGPT vs AWS Translate vs DeepL). - Linguistic analysis of technical bug reporting across different languages. - Insights into common software issues encountered by diverse users. - Improving multilingual issue tracking through automated labeling and categorization.

Potential Use Cases

This dataset can be beneficial for various research and development applications, including: - Machine Translation Benchmarking: Evaluating the performance of translation models in a technical domain. - Natural Language Processing (NLP) Tasks: Training classifiers to categorize bug reports based on their content. - Software Engineering Research: Understanding trends in bug reporting, issue resolution, and localization challenges. - Automated Bug Triage: Developing AI-driven solutions for assigning and prioritizing bug reports in multilingual repositories.

Benchmarks archive 2025-07-28

All 1 leaderboard whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

First row (archive order)PaperCode
Machine Translation Multi Lingual Bug Reports ChatGPT BERTScore 79 English Please: Evaluating Machine Translation with... av9ash/English-Please 1 Compare

Papers archive 2025-07-28

1 shown of 1 paper with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 1. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
English Please: Evaluating Machine Translation with Large Language Models for Multilingual Bug Reports 1 1 20 Feb 2025 not harvested

Dataset loaders archive 2025-07-28

No loader listed in the archive.

Tasks archive 2025-07-28

License archive 2025-07-28

CC BY 4.0

Modalities archive 2025-07-28

Languages archive 2025-07-28

Variants archive 2025-07-28

  • Multi Lingual Bug Reports

1 variant name, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections