Papers › Learning Multi-Subset of Classes for Fine-Grained Food Recognition
Learning Multi-Subset of Classes for Fine-Grained Food Recognition
Javier Ródenas, Bhalaji Nagarajan, Marc Bolaños, Petia Radeva
Food image recognition is a complex computer vision task, because of the large number of fine-grained food classes. Fine-grained recognition tasks focus on learning subtle discriminative details to distinguish similar classes. In this paper, we introduce a new method to improve the classification of classes that are more difficult to discriminate based on Multi-Subsets learning. Using a pre-trained network, we organize classes in multiple subsets using a clustering technique. Later, we embed these subsets in a multi-head model structure. This structure has three distinguishable parts. First, we use several shared blocks to learn the generalized representation of the data. Second, we use multiple specialized blocks focusing on specific subsets that are difficult to distinguish. Lastly, we use a fully connected layer to weight the different subsets in an end-to-end manner by combining the neuron outputs. We validated our proposed method using two recent state-of-the-art vision transformers on three public food recognition datasets. Our method was successful in learning the confused classes better and we outperformed the state-of-the-art on the three datasets.
Code
Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.
Code Syntology ran Syntology
Not run by Syntology. Nothing on this page verifies that the listed code works.
Tasks
Results from the paper archive 2025-07-28
| Task | Dataset | Model | Metric | Value | Rank at snapshot | Leaderboard | Report |
|---|---|---|---|---|---|---|---|
| Fine-Grained Image Classification | Food-101 | CSWin-L | Accuracy | 93.81 | #5 of 15 | Archive leaderboard | report |
| Fine-Grained Image Classification | Food-101 | VOLO-D5 | Accuracy | 93.66 | #7 of 15 | Archive leaderboard | report |
| Fine-Grained Image Classification | FoodX-251 | CSWin-L | Accuracy (%) | 79.90 | #1 of 2 | Archive leaderboard | report |
| Fine-Grained Image Classification | FoodX-251 | VOLO-D5 | Accuracy (%) | 79.15 | #2 of 2 | Archive leaderboard | report |
Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections