Papers › Fine-grained Image Classification and Retrieval by Combining Visual and Locally Pooled...

Fine-grained Image Classification and Retrieval by Combining Visual and Locally Pooled Textual Features

14 Jan 2020arXiv:2001.04732archive 2025-07-28

Andres Mafla, Sounak Dey, Ali Furkan Biten, Lluis Gomez, Dimosthenis Karatzas

Text contained in an image carries high-level semantics that can be exploited to achieve richer image understanding. In particular, the mere presence of text provides strong guiding content that should be employed to tackle a diversity of computer vision tasks such as image retrieval, fine-grained classification, and visual question answering. In this paper, we address the problem of fine-grained classification and image retrieval by leveraging textual information along with visual cues to comprehend the existing intrinsic relation between the two modalities. The novelty of the proposed model consists of the usage of a PHOC descriptor to construct a bag of textual words along with a Fisher Vector Encoding that captures the morphology of text. This approach provides a stronger multimodal representation for this task and as our experiments demonstrate, it achieves state-of-the-art results on two different tasks, fine-grained classification and image retrieval.

PaperPDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

DreadPiratePsyopus/Fine_Grained_Clf officialmentioned in papermentioned on GitHubpytorch report
AndresPMD/Fine_Grained_Clf mentioned on GitHubpytorch report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

ClassificationDiversityFine-Grained Image ClassificationGeneral ClassificationImage ClassificationImage RetrievalQuestion AnsweringRetrievalVisual Question AnsweringVisual Question Answering (VQA)image-classification

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Fine-Grained Image Classification Bottles PHOC descriptor + Fisher Vector Encoding mAP 77.4 #1 of 1 Archive leaderboard report
Fine-Grained Image Classification Con-Text PHOC descriptor + Fisher Vector Encoding mAP 80.2 #1 of 1 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections