Papers › Beyond Specialization: Assessing the Capabilities of MLLMs in Age and Gender Estimation

Beyond Specialization: Assessing the Capabilities of MLLMs in Age and Gender Estimation

4 Mar 2024arXiv:2403.02302archive 2025-07-28

Maksim Kuprashevich, Grigorii Alekseenko, Irina Tolstykh

Multimodal Large Language Models (MLLMs) have recently gained immense popularity. Powerful commercial models like ChatGPT-4V and Gemini, as well as open-source ones such as LLaVA, are essentially general-purpose models and are applied to solve a wide variety of tasks, including those in computer vision. These neural networks possess such strong general knowledge and reasoning abilities that they have proven capable of working even on tasks for which they were not specifically trained. We compared the capabilities of the most powerful MLLMs to date: ShareGPT4V, ChatGPT, LLaVA-Next in a specialized task of age and gender estimation with our state-of-the-art specialized model, MiVOLO. We also updated MiVOLO and provide details and new metrics in this article. This comparison has yielded some interesting results and insights about the strengths and weaknesses of the participating models. Furthermore, we attempted various ways to fine-tune the ShareGPT4V model for this specific task, aiming to achieve state-of-the-art results in this particular challenge. Although such a model would not be practical in production, as it is incredibly expensive compared to a specialized model like MiVOLO, it could be very useful in some tasks, like data annotation.

PaperPDFConference PDFCode

Code

wildchlamydia/mivolo officialmentioned on GitHubpytorch report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Age And Gender ClassificationAge EstimationAge and Gender EstimationFacial Attribute ClassificationGender PredictionGeneral Knowledge

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Age And Gender Classification Adience Age MiVOLO-V2 Accuracy (5-fold) 69.43 #2 of 16 Archive leaderboard report
Age And Gender Classification Adience Gender MiVOLO-V2 Accuracy (5-fold) 97.39 #1 of 10 Archive leaderboard report
Age Estimation CACD MiVOLO-V2 MAE 3.89 #1 of 13 Archive leaderboard report
Age Estimation IMDB-Clean MiVOLO-V2 Average mean absolute error 3.97 #1 of 4 Archive leaderboard report
Age Estimation LAGENDA MiVOLO-V2 MAE 3.65 #1 of 2 Archive leaderboard report
Age and Gender Estimation LAGENDA age MiVOLO-V2 CS@5 74.48 #1 of 2 Archive leaderboard report
Age and Gender Estimation LAGENDA age MiVOLO-V2 MAE 3.65 #1 of 2 Archive leaderboard report
Age and Gender Estimation LAGENDA gender MiVOLO-V2 CS@5 74.48 #2 of 2 Archive leaderboard report
Facial Attribute Classification FairFace MiVOLO-V2 age-top1 62.28 #1 of 3 Archive leaderboard report
Facial Attribute Classification FairFace MiVOLO-V2 gender-top1 97.5 #1 of 3 Archive leaderboard report
Gender Prediction LAGENDA MiVOLO-V2 Accuracy 97.99 #1 of 2 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections