Papers › Decoupled Box Proposal and Featurization with Ultrafine-Grained Semantic Labels...
Decoupled Box Proposal and Featurization with Ultrafine-Grained Semantic Labels Improve Image Captioning and Visual Question Answering
Soravit Changpinyo, Bo Pang, Piyush Sharma, Radu Soricut
Object detection plays an important role in current solutions to vision and language tasks like image captioning and visual question answering. However, popular models like Faster R-CNN rely on a costly process of annotating ground-truths for both the bounding boxes and their corresponding semantic labels, making it less amenable as a primitive task for transfer learning. In this paper, we examine the effect of decoupling box proposal and featurization for down-stream tasks. The key insight is that this allows us to leverage a large amount of labeled annotations that were previously unavailable for standard object detection benchmarks. Empirically, we demonstrate that this leads to effective transfer learning and improved image captioning and visual question answering models, as measured on publicly available benchmarks.
In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.
Code
No code repository is listed for this paper in the archive or in Syntology's graph.
Code Syntology ran Syntology
Not run by Syntology. Nothing on this page verifies that the listed code works.
Tasks
Results from the paper archive 2025-07-28
| Task | Dataset | Model | Metric | Value | Rank at snapshot | Leaderboard | Report |
|---|---|---|---|---|---|---|---|
| Visual Question Answering (VQA) | VizWiz 2018 | B-Ultra | number | 28.81 | #4 of 10 | Archive leaderboard | report |
| Visual Question Answering (VQA) | VizWiz 2018 | B-Ultra | other | 35.41 | #4 of 10 | Archive leaderboard | report |
| Visual Question Answering (VQA) | VizWiz 2018 | B-Ultra | overall | 53.68 | #4 of 10 | Archive leaderboard | report |
| Visual Question Answering (VQA) | VizWiz 2018 | B-Ultra | unanswerable | 84.03 | #4 of 10 | Archive leaderboard | report |
| Visual Question Answering (VQA) | VizWiz 2018 | B-Ultra | yes/no | 68.12 | #4 of 10 | Archive leaderboard | report |
Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.
Methods
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections