Papers › Exploiting Multiple Sequence Lengths in Fast End to End Training for Image Captioning
Exploiting Multiple Sequence Lengths in Fast End to End Training for Image Captioning
Jia Cheng Hu, Roberto Cavicchioli, Alessandro Capotondi
We introduce a method called the Expansion mechanism that processes the input unconstrained by the number of elements in the sequence. By doing so, the model can learn more effectively compared to traditional attention-based approaches. To support this claim, we design a novel architecture ExpansionNet v2 that achieved strong results on the MS COCO 2014 Image Captioning challenge and the State of the Art in its respective category, with a score of 143.7 CIDErD in the offline test split, 140.8 CIDErD in the online evaluation server and 72.9 AllCIDEr on the nocaps validation set. Additionally, we introduce an End to End training algorithm up to 2.8 times faster than established alternatives. Source code available at: https://github.com/jchenghu/ExpansionNet_v2
In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.
Code
Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.
Code Syntology ran Syntology
Not run by Syntology. Nothing on this page verifies that the listed code works.
Tasks
Results from the paper archive 2025-07-28
| Task | Dataset | Model | Metric | Value | Rank at snapshot | Leaderboard | Report |
|---|---|---|---|---|---|---|---|
| Image Captioning | COCO (Common Objects in Context) | ExpansionNet v2 | CIDEr | 143.7 | #1 of 17 | Archive leaderboard | report |
| Image Captioning | COCO Captions | ExpansionNet v2 (No VL pretraining) | BLEU-1 | 83.5 | #6 of 41 | Archive leaderboard | report |
| Image Captioning | COCO Captions | ExpansionNet v2 (No VL pretraining) | BLEU-4 | 42.7 | #6 of 41 | Archive leaderboard | report |
| Image Captioning | COCO Captions | ExpansionNet v2 (No VL pretraining) | CIDER | 143.7 | #6 of 41 | Archive leaderboard | report |
| Image Captioning | COCO Captions | ExpansionNet v2 (No VL pretraining) | METEOR | 30.6 | #6 of 41 | Archive leaderboard | report |
| Image Captioning | COCO Captions | ExpansionNet v2 (No VL pretraining) | ROUGE-L | 61.1 | #6 of 41 | Archive leaderboard | report |
| Image Captioning | COCO Captions | ExpansionNet v2 (No VL pretraining) | SPICE | 24.7 | #6 of 41 | Archive leaderboard | report |
Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.
Methods
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections