Datasets › InpaintCOCO

InpaintCOCO

16 Aug 2024 archive 2025-07-28

InpaintCOCO is a benchmark to understand fine-grained concepts in multimodal models (vision-language) similar to Winoground. To our knowledge InpaintCOCO is the first benchmark, which consists of image pairs with minimum differences, so that the visual representation can be analyzed in a more standardized setting.

A data sample contains 2 images and 2 corresponding captions that differ only in one object, the color of an object, or the size of an object.

The metric used in the paper compares if true image-text pairs are more similar than wrong image-text combinations and that for both image-text pairs: sim(i_(COCO),t_(COCO)) > sim(iᵢₙₚ,t_(COCO)) sim(iᵢₙₚ,tᵢₙₚ) > sim(i_(COCO),tᵢₙₚ)

InpaintCOCO is published in Proceedings of the 3rd Workshop on Advances in Language and Vision Research (ALVR) at ACL 2024.

Benchmarks archive 2025-07-28

No leaderboard in the archive resolves to this dataset.

Papers archive 2025-07-28

No paper in the archive has a leaderboard row on this dataset.

Dataset loaders archive 2025-07-28

No loader listed in the archive.

Tasks archive 2025-07-28

License archive 2025-07-28

diverse licenses

Modalities archive 2025-07-28

Languages archive 2025-07-28

Variants archive 2025-07-28

  • InpaintCOCO

1 variant name, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections