Papers › What is the Role of Recurrent Neural Networks (RNNs) in an Image Caption Generator?

What is the Role of Recurrent Neural Networks (RNNs) in an Image Caption Generator?

7 Aug 2017WS 2017 9arXiv:1708.02043archive 2025-07-28

Marc Tanti, Albert Gatt, Kenneth P. Camilleri

In neural image captioning systems, a recurrent neural network (RNN) is typically viewed as the primary `generation' component. This view suggests that the image features should be `injected' into the RNN. This is in fact the dominant view in the literature. Alternatively, the RNN can instead be viewed as only encoding the previously generated words. This view suggests that the RNN should only be used to encode linguistic features and that only the final representation should be `merged' with the image features at a later stage. This paper compares these two architectures. We find that, in general, late merging outperforms injection, suggesting that RNNs are better viewed as encoders, rather than generators.

PaperPDFConference PDFCode

Code

mtanti/rnn-role officialmentioned in papertf report
sinazarriess/zero_shot_reg mentioned on GitHubtf report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Image Captioning

Results from the paper archive 2025-07-28

No leaderboard rows for this paper in the archive.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections