Papers › MoVie: Revisiting Modulated Convolutions for Visual Counting and Beyond

MoVie: Revisiting Modulated Convolutions for Visual Counting and Beyond

24 Apr 2020ICLR 2021 1arXiv:2004.11883archive 2025-07-28

Duy-Kien Nguyen, Vedanuj Goswami, Xinlei Chen

This paper focuses on visual counting, which aims to predict the number of occurrences given a natural image and a query (e.g. a question or a category). Unlike most prior works that use explicit, symbolic models which can be computationally expensive and limited in generalization, we propose a simple and effective alternative by revisiting modulated convolutions that fuse the query and the image locally. Following the design of residual bottleneck, we call our method MoVie, short for Modulated conVolutional bottlenecks. Notably, MoVie reasons implicitly and holistically and only needs a single forward-pass during inference. Nevertheless, MoVie showcases strong performance for counting: 1) advancing the state-of-the-art on counting-specific VQA tasks while being more efficient; 2) outperforming prior-art on difficult benchmarks like COCO for common object counting; 3) helped us secure the first place of 2020 VQA challenge when integrated as a module for 'number' related questions in generic VQA models. Finally, we show evidence that modulated convolutions such as MoVie can serve as a general mechanism for reasoning tasks beyond counting.

PaperPDFConference PDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

facebookresearch/mmf officialpytorch report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Object CountingQuestion AnsweringVisual Question Answering (VQA)

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Object Counting HowMany-QA MoVie-ResNeXt Accuracy 64 #1 of 3 Archive leaderboard report
Object Counting HowMany-QA MoVie-ResNeXt RMSE 2.3 #1 of 3 Archive leaderboard report
Object Counting HowMany-QA MoVie Accuracy 61.2 #2 of 3 Archive leaderboard report
Object Counting HowMany-QA MoVie RMSE 2.36 #2 of 3 Archive leaderboard report
Object Counting TallyQA-Complex MoVie-ResNeXt Accuracy 56.8 #4 of 6 Archive leaderboard report
Object Counting TallyQA-Complex MoVie-ResNeXt RMSE 1.43 #4 of 6 Archive leaderboard report
Object Counting TallyQA-Complex MoVie Accuracy 54.1 #6 of 6 Archive leaderboard report
Object Counting TallyQA-Complex MoVie RMSE 1.52 #6 of 6 Archive leaderboard report
Object Counting TallyQA-Simple MoVie-ResNeXt Accuracy 74.9 #4 of 6 Archive leaderboard report
Object Counting TallyQA-Simple MoVie-ResNeXt RMSE 1 #4 of 6 Archive leaderboard report
Object Counting TallyQA-Simple MoVie Accuracy 70.8 #6 of 6 Archive leaderboard report
Object Counting TallyQA-Simple MoVie RMSE 1.09 #6 of 6 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections