Papers › Look Deeper See Richer: Depth-aware Image Paragraph Captioning

Look Deeper See Richer: Depth-aware Image Paragraph Captioning

15 Oct 2018ACM International Conference on Multimedia 2018 10archive 2025-07-28

Ziwei Wang, Yadan Luo, Yang Li, Zi Huang, Hongzhi Yin

With the widespread availability of image captioning at a sentence level, how to automatically generate image paragraphs is yet well explored. Describing an image by a full paragraph involves organising sentences orderly, coherently and diversely, inevitably leading higher complexity than by a single sentence. Existing image paragraph captioning methods give a series of sentences to represent the objects and regions of interests, where the descriptions are essentially generated by feeding the image fragments containing objects and regions into conventional image single-sentence captioning models. This strategy is difficult to generate the descriptions that guarantee the stereoscopic hierarchy and non-overlapping objects. In this paper, we propose a Depth-aware Attention Model (\textitDAM ) to generate paragraph captions for images. The depths of image areas are firstly estimated in order to discriminate objects in a range of spatial locations, which can further guide the linguistic decoder to reveal spatial relationships among objects. This model completes the paragraph in a logical and coherent manner. By incorporating the attention mechanism, the learned model swiftly shifts the sentence focus during paragraph generation, whilst avoiding verbose descriptions on a same object. Extensive quantitative experiments and the user study have been conducted on the Visual Genome dataset, which demonstrate the effectiveness and the interpretability of the proposed model.

PaperPDF

Code

No code repository is listed for this paper in the archive or in Syntology's graph.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

DecoderImage CaptioningImage Paragraph CaptioningSentence

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Image Paragraph Captioning Image Paragraph Captioning Depth-aware Attention Model (DAM) BLEU-4 6.7 #9 of 10 Archive leaderboard report
Image Paragraph Captioning Image Paragraph Captioning Depth-aware Attention Model (DAM) CIDEr 17.3 #9 of 10 Archive leaderboard report
Image Paragraph Captioning Image Paragraph Captioning Depth-aware Attention Model (DAM) METEOR 13.9 #9 of 10 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections