Datasets › PhotoChat

PhotoChat

Introduced by Xiaoxue Zang et al. in PhotoChat: A Human-Human Dialogue Dataset with Photo Sharing Behavior for Joint Image-Text Modeling6 Jul 2021 archive 2025-07-28

PhotoChat, the first dataset that casts light on the photo sharing behavior in online messaging. PhotoChat contains 12k dialogues, each of which is paired with a user photo that is shared during the conversation. Based on this dataset, we propose two tasks to facilitate research on image-text modeling: a photo-sharing intent prediction task that predicts whether one intends to share a photo in the next conversation turn, and a photo retrieval task that retrieves the most relevant photo according to the dialogue context.

Benchmarks archive 2025-07-28

All 2 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

Papers archive 2025-07-28

8 shown of 8 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 20. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
PaCE: Unified Multi-modal Dialogue Pre-training with Progressive and Compositional Experts 1 2 24 May 2023 ran 0 of 1 samples (1 unverified)
VLMo: Unified Vision-Language Pre-Training with Mixture-of-Modality-Experts 2 1 3 Nov 2021 not harvested
PhotoChat: A Human-Human Dialogue Dataset with Photo Sharing Behavior for Joint Image-Text Modeling 0 1 6 Jul 2021 not harvested
ViLT: Vision-and-Language Transformer Without Convolution or Region Supervision 6 2 5 Feb 2021 ran 1 of 4 samples (3 unverified; 1 pointer-only for licence)
Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer 57 2 23 Oct 2019 ran 2 of 31 samples (29 unverified)
ALBERT: A Lite BERT for Self-supervised Learning of Language Representations 48 1 26 Sep 2019 ran 46 of 126 samples (80 unverified; 22 pointer-only for licence)
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding 534 1 11 Oct 2018 ran 204 of 659 samples (455 unverified; 149 pointer-only for licence)
Stacked Cross Attention for Image-Text Matching 6 1 21 Mar 2018 ran 7 of 16 samples (9 unverified; 1 pointer-only for licence)

Dataset loaders archive 2025-07-28

No loader listed in the archive.

Tasks archive 2025-07-28

License archive 2025-07-28

https://github.com/google-research/google-research/tree/master/multimodalchat

Modalities archive 2025-07-28

Languages archive 2025-07-28

Variants archive 2025-07-28

  • PhotoChat

1 variant name, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections