Datasets › Business Scene Dialogue

Business Scene Dialogue

Introduced by Matīss Rikters et al. in Designing the Business Conversation Corpus5 Aug 2020 archive 2025-07-28

The Japanese-English business conversation corpus, namely Business Scene Dialogue corpus, was constructed in 3 steps:

  1. selecting business scenes,
  2. writing monolingual conversation scenarios according to the selected scenes, and
  3. translating the scenarios into the other language.

Half of the monolingual scenarios were written in Japanese and the other half were written in English. The whole construction process was supervised by a person who satisfies the following conditions to guarantee the conversations to be natural:

  • has the experience of being engaged in language learning programs, especially for business conversations
  • is able to smoothly communicate with others in various business scenes both in Japanese and English
  • has the experience of being involved in business

The BSD corpus is split into balanced training, development and evaluation sets. The documents in these sets are balanced in terms of scenes and original languages. In this repository we publicly share the full development and evaluation sets and a part of the training data set.

Source: BSD

Benchmarks archive 2025-07-28

All 2 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

First row (archive order)PaperCode
Machine Translation Business Scene Dialogue JA-EN Transformer-base BLEU 12.88 Designing the Business Conversation Corpus tsuruoka-lab/BSD 1 Compare
Machine Translation Business Scene Dialogue EN-JA Transformer-base BLEU 13.53 Designing the Business Conversation Corpus tsuruoka-lab/BSD 1 Compare

Papers archive 2025-07-28

1 shown of 1 paper with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 9. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
Designing the Business Conversation Corpus 1 2 5 Aug 2020 not harvested

Dataset loaders archive 2025-07-28

3 loaders as listed in the archive; links are outbound and not re-checked here.

Tasks archive 2025-07-28

License archive 2025-07-28

No licence recorded in the archive. Absence here is not a statement about the dataset's terms.

Modalities archive 2025-07-28

Languages archive 2025-07-28

Variants archive 2025-07-28

  • Business Scene Dialogue JA-EN
  • Business Scene Dialogue EN-JA
  • Business Scene Dialogue

3 variant names, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections