Papers › Im2Text: Describing Images Using 1 Million Captioned Photographs

Im2Text: Describing Images Using 1 Million Captioned Photographs

1 Dec 2011NeurIPS 2011 12archive 2025-07-28

Vicente Ordonez, Girish Kulkarni, Tamara L. Berg

We develop and demonstrate automatic image description methods using a large captioned photo collection. One contribution is our technique for the automatic collection of this new dataset -- performing a huge number of Flickr queries and then filtering the noisy results down to 1 million images with associated visually relevant captions. Such a collection allows us to approach the extremely challenging problem of description generation using relatively simple non-parametric methods and produces surprisingly effective results. We also develop methods incorporating many state of the art, but fairly noisy, estimates of image content to produce even more pleasing results. Finally we introduce a new objective performance measure for image captioning.

PaperPDF

Code

No code repository is listed for this paper in the archive or in Syntology's graph.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Image Captioning

1 archive task tag without a task page not shown.

Datasets

Introduced by this paper, per the archive.

SBU Captions Dataset

Results from the paper archive 2025-07-28

No leaderboard rows for this paper in the archive.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections