Papers › Image and Text fusion for UPMC Food-101 \\using BERT and CNNs

Image and Text fusion for UPMC Food-101 \\using BERT and CNNs

17 Dec 2020archive 2025-07-28

Ignazio Gallo, Gianmarco Ria, Nicola Landro, and Riccardo La Grassa

The modern digital world is becoming more and more multimodal. Looking on the internet, images are often associated with the text, so classification problems with these two modalities are very common. In this paper, we examine multimodal classification using textual information and visual representations of the same concept. We investigate two main basic methods to perform multimodal fusion and adapt them with stacking techniques to better handle this type of problem. Here, we use UPMC Food-101, which is a difficult and noisy multimodal dataset that well represents this category of multimodal problems. Our results show that the proposed early fusion technique combined with a stacking-based approach exceeds the state of the art on the dataset used.

PaperPDFCode

Code

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

ClassificationDocument Text ClassificationGeneral ClassificationImage ClassificationMulti-Modal Document ClassificationMultimodal Deep LearningMultimodal Text and Image Classification

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Document Text Classification Food-101 Bert Accuracy (%) 84.41 #1 of 1 Archive leaderboard report
Image Classification Food-101 Inception V3 Accuracy (%) 71.67 #11 of 11 Archive leaderboard report
Multimodal Text and Image Classification Food-101 Early Fusion (Bert + InceptionV3) Accuracy (%) 92.5 #1 of 2 Archive leaderboard report
Multimodal Text and Image Classification Food-101 Late Fusion (Bert + InceptionV3) Accuracy (%) 84.59 #2 of 2 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections