{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/unem-unrolled-generalized-em-for-transductive","title":"UNEM: UNrolled Generalized EM for Transductive Few-Shot Learning","arxiv_id":"2412.16739","date":"2024-12-21","proceeding":"CVPR 2025 1","authors":["Long Zhou","Fereshteh Shakeri","Aymen Sadraoui","Mounir Kaaniche","Jean-Christophe Pesquet","Ismail Ben Ayed"],"abstract":"Transductive few-shot learning has recently triggered wide attention in computer vision. Yet, current methods introduce key hyper-parameters, which control the prediction statistics of the test batches, such as the level of class balance, affecting performances significantly. Such hyper-parameters are empirically grid-searched over validation data, and their configurations may vary substantially with the target dataset and pre-training model, making such empirical searches both sub-optimal and computationally intractable. In this work, we advocate and introduce the unrolling paradigm, also referred to as \"learning to optimize\", in the context of few-shot learning, thereby learning efficiently and effectively a set of optimized hyper-parameters. Specifically, we unroll a generalization of the ubiquitous Expectation-Maximization (EM) optimizer into a neural network architecture, mapping each of its iterates to a layer and learning a set of key hyper-parameters over validation data. Our unrolling approach covers various statistical feature distributions and pre-training paradigms, including recent foundational vision-language models and standard vision-only classifiers. We report comprehensive experiments, which cover a breadth of fine-grained downstream image classification tasks, showing significant gains brought by the proposed unrolled EM algorithm over iterative variants. The achieved improvements reach up to 10% and 7.5% on vision-only and vision-language benchmarks, respectively.","url_abs":"https://arxiv.org/abs/2412.16739v1","url_pdf":"https://arxiv.org/pdf/2412.16739v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"unem-unrolled-generalized-em-for-transductive","repo_url":"https://github.com/zhoulong0/unem-transductive","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"few-shot-learning","task_name":"Few-Shot Learning"},{"task_slug":"few-shot-learning-4-shots","task_name":"Few-Shot Learning - 4 shots"},{"task_slug":"image-classification","task_name":"Image Classification"},{"task_slug":"image-classification","task_name":"image-classification"}],"methods":[{"method_slug":"attention","method_name":"Attention"},{"method_slug":"set","method_name":"SET"},{"method_slug":"softmax","method_name":"Softmax"}],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/few-shot-learning-on-mini-imagenet-10-shot","task":"Few-Shot Learning","dataset":"Mini-ImageNet - 10-Shot Learning","model":"UNEM-Gaussian","rank_in_archive_order":1,"of":1,"metrics":{"Accuracy":"65.7"},"uses_additional_data":false},{"leaderboard":"/sota/few-shot-learning-on-mini-imagenet-20-shot","task":"Few-Shot Learning","dataset":"Mini-ImageNet - 20-Shot Learning","model":"UNEM-Gaussian","rank_in_archive_order":1,"of":1,"metrics":{"Accuracy":"73.2"},"uses_additional_data":false},{"leaderboard":"/sota/few-shot-learning-on-mini-imagenet-5-shot","task":"Few-Shot Learning","dataset":"Mini-ImageNet - 5-Shot Learning","model":"UNEM-Gaussian","rank_in_archive_order":3,"of":3,"metrics":{"Accuracy":"66.4%"},"uses_additional_data":false},{"leaderboard":"/sota/few-shot-learning-on-tieredimagenet-5-shot","task":"Few-Shot Learning","dataset":"tieredImageNet - 5-shot","model":"UNEM-Gaussian","rank_in_archive_order":1,"of":1,"metrics":{"Accuracy":"52.3"},"uses_additional_data":false}],"syntology":{"syntology_url":null,"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}