Browse › Computer Vision › Image Captioning › foundation-multimodal-models/DetailCaps-4870

Image Captioning archive 2025-07-28

foundation-multimodal-models/DetailCaps-4870 Benchmark (Image Captioning)

0 rows 0 with code listed 1 metric

Image Captioning is the task of describing the content of an image in words. This task lies at the intersection of computer vision and natural language processing. Most image captioning systems use an encoder-decoder framework, where an input image is encoded into an intermediate representation of the information in the image, and then decoded into a descriptive text sequence. The most popular benchmarks are nocaps and COCO, and models are typically evaluated according to a BLEU or CIDER metric.

The archive carries no text for this table; the description above is the archive's text for the task Image Captioning. archive 2025-07-28

Results archive 2025-07-28

No rows in the archive for this table at snapshot 2025-07-28. It declares 1 metric (value) but no result was ever recorded against it. That says nothing about whether results exist elsewhere.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections