Datasets › InpaintCOCO
InpaintCOCO
InpaintCOCO is a benchmark to understand fine-grained concepts in multimodal models (vision-language) similar to Winoground. To our knowledge InpaintCOCO is the first benchmark, which consists of image pairs with minimum differences, so that the visual representation can be analyzed in a more standardized setting.
A data sample contains 2 images and 2 corresponding captions that differ only in one object, the color of an object, or the size of an object.
The metric used in the paper compares if true image-text pairs are more similar than wrong image-text combinations and that for both image-text pairs: sim(i_(COCO),t_(COCO)) > sim(iᵢₙₚ,t_(COCO)) sim(iᵢₙₚ,tᵢₙₚ) > sim(i_(COCO),tᵢₙₚ)
InpaintCOCO is published in Proceedings of the 3rd Workshop on Advances in Language and Vision Research (ALVR) at ACL 2024.
Benchmarks archive 2025-07-28
No leaderboard in the archive resolves to this dataset.
Papers archive 2025-07-28
No paper in the archive has a leaderboard row on this dataset.
Dataset loaders archive 2025-07-28
No loader listed in the archive.
Tasks archive 2025-07-28
License archive 2025-07-28
Modalities archive 2025-07-28
Languages archive 2025-07-28
Variants archive 2025-07-28
- InpaintCOCO
1 variant name, as the archive lists them.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections