Papers › Exploiting Multiple Sequence Lengths in Fast End to End Training for Image Captioning

Exploiting Multiple Sequence Lengths in Fast End to End Training for Image Captioning

13 Aug 2022arXiv:2208.06551archive 2025-07-28

Jia Cheng Hu, Roberto Cavicchioli, Alessandro Capotondi

We introduce a method called the Expansion mechanism that processes the input unconstrained by the number of elements in the sequence. By doing so, the model can learn more effectively compared to traditional attention-based approaches. To support this claim, we design a novel architecture ExpansionNet v2 that achieved strong results on the MS COCO 2014 Image Captioning challenge and the State of the Art in its respective category, with a score of 143.7 CIDErD in the offline test split, 140.8 CIDErD in the online evaluation server and 72.9 AllCIDEr on the nocaps validation set. Additionally, we introduce an End to End training algorithm up to 2.8 times faster than established alternatives. Source code available at: https://github.com/jchenghu/ExpansionNet_v2

PaperPDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

jchenghu/expansionnet_v2 officialmentioned in papermentioned on GitHubpytorch report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Image Captioning

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Image Captioning COCO (Common Objects in Context) ExpansionNet v2 CIDEr 143.7 #1 of 17 Archive leaderboard report
Image Captioning COCO Captions ExpansionNet v2 (No VL pretraining) BLEU-1 83.5 #6 of 41 Archive leaderboard report
Image Captioning COCO Captions ExpansionNet v2 (No VL pretraining) BLEU-4 42.7 #6 of 41 Archive leaderboard report
Image Captioning COCO Captions ExpansionNet v2 (No VL pretraining) CIDER 143.7 #6 of 41 Archive leaderboard report
Image Captioning COCO Captions ExpansionNet v2 (No VL pretraining) METEOR 30.6 #6 of 41 Archive leaderboard report
Image Captioning COCO Captions ExpansionNet v2 (No VL pretraining) ROUGE-L 61.1 #6 of 41 Archive leaderboard report
Image Captioning COCO Captions ExpansionNet v2 (No VL pretraining) SPICE 24.7 #6 of 41 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

Test

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections