{"url":"/task/zero-shot-learning","name":"Zero-Shot Learning","slug":"zero-shot-learning","description_markdown":"**Zero-shot learning (ZSL)** is a model's ability to detect classes never seen during training. The condition is that the classes are not known during supervised learning. \r\n\r\nEarlier work in zero-shot learning use attributes in a two-step approach to infer unknown classes. In the computer vision context, more recent advances learn mappings from image feature space to semantic space. Other approaches learn non-linear multimodal embeddings. In the modern NLP context, language models can be evaluated on downstream tasks without fine tuning. \r\n\r\nBenchmark datasets for zero-shot learning include [aPY](/dataset/apy), [AwA](/dataset/awa2-1), and [CUB](/dataset/cub-200-2011), among others. \r\n\r\n( Image credit: [Prototypical Networks for Few shot Learning in PyTorch\r\n](https://github.com/orobix/Prototypical-Networks-for-Few-shot-Learning-PyTorch) )\r\n\r\nFurther readings:  \r\n\r\n- [Zero-Shot Learning -- A Comprehensive Evaluation of the Good, the Bad and the Ugly](https://paperswithcode.com/paper/zero-shot-learning-a-comprehensive-evaluation)\r\n- [Zero-Shot Learning in Modern NLP](https://joeddav.github.io/blog/2020/05/29/ZSL.html)\r\n- [Zero-Shot Learning for Text Classification](https://amitness.com/2020/05/zero-shot-text-classification/)","categories":[{"name":"Computer Vision","url":"/area/computer-vision"},{"name":"Methodology","url":"/area/methodology"}],"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","slug_source":"archive_url"},"counts":{"papers_tagged":1864,"papers_with_code":787,"benchmarks":33,"benchmark_tables_in_archive":33,"benchmark_tables_shown":33,"benchmark_tables_withheld_as_spam":0,"benchmark_definition":"a leaderboard table with at least one row; benchmark_tables_shown also counts the zero-row tables; benchmark_tables_in_archive adds the tables withheld as spam","datasets":43,"subtasks":7,"parent_tasks":0},"benchmarks":[{"leaderboard":"/sota/zero-shot-learning-on-cub-200-2011","slug":"zero-shot-learning-on-cub-200-2011","dataset":"CUB-200-2011","dataset_url":"/dataset/cub-200-2011","rows_in_archive":14,"metrics":["average top-1 classification accuracy","Accuracy Seen","Accuracy Unseen","H","Accuracy"],"first_row_in_archive_order":{"model":"ZeroDiff","paper_title":"Exploring Data Efficiency in Zero-Shot Learning with Diffusion Models","paper_url":"/paper/exploring-data-efficiency-in-zero-shot","paper_date":"2024-06-05","arxiv_id":"2406.02929","code_links":[],"syntology":null}},{"leaderboard":"/sota/zero-shot-learning-on-medconceptsqa","slug":"zero-shot-learning-on-medconceptsqa","dataset":"MedConceptsQA","dataset_url":"/dataset/medconceptsqa","rows_in_archive":13,"metrics":["Accuracy"],"first_row_in_archive_order":{"model":"gpt-4-0125-preview","paper_title":"GPT-4 Technical Report","paper_url":"/paper/gpt-4-technical-report-1","paper_date":"2023-03-15","arxiv_id":"2303.08774","code_links":[{"title":"openai/evals","url":"https://github.com/openai/evals"},{"title":"shmsw25/factscore","url":"https://github.com/shmsw25/factscore"},{"title":"unispac/visual-adversarial-examples-jailbreak-large-language-models","url":"https://github.com/unispac/visual-adversarial-examples-jailbreak-large-language-models"},{"title":"gpt4life/alpagasus","url":"https://github.com/gpt4life/alpagasus"},{"title":"emrgnt-cmplxty/zero-shot-replication","url":"https://github.com/emrgnt-cmplxty/zero-shot-replication"},{"title":"ethz-privsec/superhuman-ai-consistency","url":"https://github.com/ethz-privsec/superhuman-ai-consistency"},{"title":"ethz-spylab/superhuman-ai-consistency","url":"https://github.com/ethz-spylab/superhuman-ai-consistency"},{"title":"eternityyw/tram-benchmark","url":"https://github.com/eternityyw/tram-benchmark"},{"title":"AUCOHL/RTL-Repo","url":"https://github.com/AUCOHL/RTL-Repo"},{"title":"zach-zhiling-zheng/reticular_chemist","url":"https://github.com/zach-zhiling-zheng/reticular_chemist"},{"title":"lflage/openfactscore","url":"https://github.com/lflage/openfactscore"}],"syntology":{"n":5,"n_ran":2,"n_unverified":3,"n_pointer_only":1}}},{"leaderboard":"/sota/zero-shot-learning-on-sun-attribute","slug":"zero-shot-learning-on-sun-attribute","dataset":"SUN Attribute","dataset_url":"/dataset/sun-attribute","rows_in_archive":9,"metrics":["average top-1 classification accuracy","Accuracy Seen","Accuracy Unseen","H"],"first_row_in_archive_order":{"model":"ZeroDiff","paper_title":"Exploring Data Efficiency in Zero-Shot Learning with Diffusion Models","paper_url":"/paper/exploring-data-efficiency-in-zero-shot","paper_date":"2024-06-05","arxiv_id":"2406.02929","code_links":[],"syntology":null}},{"leaderboard":"/sota/zero-shot-learning-on-awa2","slug":"zero-shot-learning-on-awa2","dataset":"AwA2","dataset_url":"/dataset/awa2-1","rows_in_archive":4,"metrics":["average top-1 classification accuracy","Accuracy Seen","Accuracy Unseen","H"],"first_row_in_archive_order":{"model":"ZeroDiff","paper_title":"Exploring Data Efficiency in Zero-Shot Learning with Diffusion Models","paper_url":"/paper/exploring-data-efficiency-in-zero-shot","paper_date":"2024-06-05","arxiv_id":"2406.02929","code_links":[],"syntology":null}},{"leaderboard":"/sota/zero-shot-learning-on-caltech-101","slug":"zero-shot-learning-on-caltech-101","dataset":"Caltech-101","dataset_url":"/dataset/caltech-101","rows_in_archive":2,"metrics":["Accuracy"],"first_row_in_archive_order":{"model":"ZLaP","paper_title":"Label Propagation for Zero-shot Classification with Vision-Language Models","paper_url":"/paper/label-propagation-for-zero-shot","paper_date":"2024-04-05","arxiv_id":"2404.04072","code_links":[{"title":"vladan-stojnic/zlap","url":"https://github.com/vladan-stojnic/zlap"}],"syntology":{"n":2,"n_ran":1,"n_unverified":1,"n_pointer_only":0}}},{"leaderboard":"/sota/zero-shot-learning-on-cifar-10","slug":"zero-shot-learning-on-cifar-10","dataset":"CIFAR-10","dataset_url":"/dataset/cifar-10","rows_in_archive":2,"metrics":["Accuracy"],"first_row_in_archive_order":{"model":"ZLaP*","paper_title":"Label Propagation for Zero-shot Classification with Vision-Language Models","paper_url":"/paper/label-propagation-for-zero-shot","paper_date":"2024-04-05","arxiv_id":"2404.04072","code_links":[{"title":"vladan-stojnic/zlap","url":"https://github.com/vladan-stojnic/zlap"}],"syntology":{"n":2,"n_ran":1,"n_unverified":1,"n_pointer_only":0}}},{"leaderboard":"/sota/zero-shot-learning-on-cifar-100","slug":"zero-shot-learning-on-cifar-100","dataset":"CIFAR-100","dataset_url":"/dataset/cifar-100","rows_in_archive":2,"metrics":["Accuracy"],"first_row_in_archive_order":{"model":"ZLaP*","paper_title":"Label Propagation for Zero-shot Classification with Vision-Language Models","paper_url":"/paper/label-propagation-for-zero-shot","paper_date":"2024-04-05","arxiv_id":"2404.04072","code_links":[{"title":"vladan-stojnic/zlap","url":"https://github.com/vladan-stojnic/zlap"}],"syntology":{"n":2,"n_ran":1,"n_unverified":1,"n_pointer_only":0}}},{"leaderboard":"/sota/zero-shot-learning-on-coco-mlt","slug":"zero-shot-learning-on-coco-mlt","dataset":"COCO-MLT","dataset_url":"/dataset/coco-mlt","rows_in_archive":2,"metrics":["Average mAP"],"first_row_in_archive_order":{"model":"ResNet-50","paper_title":"Learning Transferable Visual Models From Natural Language Supervision","paper_url":"/paper/learning-transferable-visual-models-from","paper_date":"2021-02-26","arxiv_id":"2103.00020","code_links":[{"title":"openai/CLIP","url":"https://github.com/openai/CLIP"},{"title":"mlfoundations/open_clip","url":"https://github.com/mlfoundations/open_clip"},{"title":"towhee-io/towhee","url":"https://github.com/towhee-io/towhee"},{"title":"facebookresearch/vissl","url":"https://github.com/facebookresearch/vissl"},{"title":"alibaba/EasyNLP","url":"https://github.com/alibaba/EasyNLP"},{"title":"apple/ml-mobileclip","url":"https://github.com/apple/ml-mobileclip"},{"title":"OML-Team/open-metric-learning","url":"https://github.com/OML-Team/open-metric-learning"},{"title":"FreddeFrallan/Multilingual-CLIP","url":"https://github.com/FreddeFrallan/Multilingual-CLIP"},{"title":"eps696/aphantasia","url":"https://github.com/eps696/aphantasia"},{"title":"muzairkhattak/multimodal-prompt-learning","url":"https://github.com/muzairkhattak/multimodal-prompt-learning"},{"title":"moein-shariatnia/OpenAI-CLIP","url":"https://github.com/moein-shariatnia/OpenAI-CLIP"},{"title":"facebookresearch/brainmagick","url":"https://github.com/facebookresearch/brainmagick"},{"title":"PaddlePaddle/PASSL","url":"https://github.com/PaddlePaddle/PASSL/blob/main/docs/Train_CLIP_model.md"},{"title":"taited/clip-score","url":"https://github.com/taited/clip-score"},{"title":"azshue/TPT","url":"https://github.com/azshue/TPT"},{"title":"clip-italian/clip-italian","url":"https://github.com/clip-italian/clip-italian"},{"title":"dhansmair/flamingo-mini","url":"https://github.com/dhansmair/flamingo-mini"},{"title":"ylqi/count-anything","url":"https://github.com/ylqi/count-anything"},{"title":"ml-jku/cloob","url":"https://github.com/ml-jku/cloob"},{"title":"sberbank-ai/ru-clip","url":"https://github.com/sberbank-ai/ru-clip"},{"title":"ai-forever/ru-clip","url":"https://github.com/ai-forever/ru-clip"},{"title":"Kaushalya/medclip","url":"https://github.com/Kaushalya/medclip"},{"title":"ajayjain/vectorascent","url":"https://github.com/ajayjain/vectorascent"},{"title":"linjieli222/hero_video_feature_extractor","url":"https://github.com/linjieli222/hero_video_feature_extractor"},{"title":"borisdayma/clip-jax","url":"https://github.com/borisdayma/clip-jax"},{"title":"sajjjadayobi/CLIPfa","url":"https://github.com/sajjjadayobi/CLIPfa"},{"title":"mertyg/post-hoc-cbm","url":"https://github.com/mertyg/post-hoc-cbm"},{"title":"mlbio-epfl/turtle","url":"https://github.com/mlbio-epfl/turtle"},{"title":"rinnakk/japanese-clip","url":"https://github.com/rinnakk/japanese-clip"},{"title":"salesforce/pb-ovd","url":"https://github.com/salesforce/pb-ovd"},{"title":"facebookresearch/clip-rocket","url":"https://github.com/facebookresearch/clip-rocket"},{"title":"sincerass/mvlpt","url":"https://github.com/sincerass/mvlpt"},{"title":"ericyinyzy/vlattack","url":"https://github.com/ericyinyzy/vlattack"},{"title":"redcaps-dataset/redcaps-downloader","url":"https://github.com/redcaps-dataset/redcaps-downloader"},{"title":"bespontaneous/proteus-pytorch","url":"https://github.com/bespontaneous/proteus-pytorch"},{"title":"shunk031/simple-aesthetics-predictor","url":"https://github.com/shunk031/simple-aesthetics-predictor"},{"title":"giantseaweed/decree","url":"https://github.com/giantseaweed/decree"},{"title":"sithu31296/simple-object-tracking","url":"https://github.com/sithu31296/simple-object-tracking"},{"title":"Gahyeonkim09/AAPL","url":"https://github.com/Gahyeonkim09/AAPL"},{"title":"filipbasara0/simple-clip","url":"https://github.com/filipbasara0/simple-clip"},{"title":"michi-3000/eyeclip","url":"https://github.com/michi-3000/eyeclip"},{"title":"baskargroup/Arboretum","url":"https://github.com/baskargroup/Arboretum"},{"title":"baskargroup/biotrove","url":"https://github.com/baskargroup/biotrove"},{"title":"kynkaat/role-of-imagenet-classes-in-fid","url":"https://github.com/kynkaat/role-of-imagenet-classes-in-fid"},{"title":"SforAiDl/CountCLIP","url":"https://github.com/SforAiDl/CountCLIP"},{"title":"mainaksingha01/applenet","url":"https://github.com/mainaksingha01/applenet"},{"title":"zhangxu0963/npc","url":"https://github.com/zhangxu0963/npc"},{"title":"mainaksingha01/odg-clip","url":"https://github.com/mainaksingha01/odg-clip"},{"title":"jhaprince/multibully","url":"https://github.com/jhaprince/multibully"},{"title":"klemens-floege/oneprot","url":"https://github.com/klemens-floege/oneprot"},{"title":"leolee99/CLIP_ITM","url":"https://github.com/leolee99/CLIP_ITM"},{"title":"madrylab/pretraining-distribution-shift-robustness","url":"https://github.com/madrylab/pretraining-distribution-shift-robustness"},{"title":"AndresPMD/Clip_CMR","url":"https://github.com/AndresPMD/Clip_CMR"},{"title":"fastscience-ai/medflamingo","url":"https://github.com/fastscience-ai/medflamingo"},{"title":"buyeah1109/KEN","url":"https://github.com/buyeah1109/KEN"},{"title":"pseulki/rococo","url":"https://github.com/pseulki/rococo"},{"title":"shkarupa-alex/tfclip","url":"https://github.com/shkarupa-alex/tfclip"},{"title":"IMvision12/keras-vision-models","url":"https://github.com/IMvision12/keras-vision-models"},{"title":"brown-palm/ObjectPrompt","url":"https://github.com/brown-palm/ObjectPrompt"},{"title":"YvanG/VQGAN-CLIP","url":"https://github.com/YvanG/VQGAN-CLIP"},{"title":"ramanakshay/clip","url":"https://github.com/ramanakshay/clip"},{"title":"NYU-DICE-Lab/open_clip","url":"https://github.com/NYU-DICE-Lab/open_clip"},{"title":"shivammehta25/clip","url":"https://github.com/shivammehta25/clip"},{"title":"minhanh151/respro","url":"https://github.com/minhanh151/respro"},{"title":"nopperl/clip_arxiv_pmc","url":"https://github.com/nopperl/clip_arxiv_pmc"},{"title":"s-a-malik/multi-few","url":"https://github.com/s-a-malik/multi-few"},{"title":"prabhupad26/100daysofML","url":"https://github.com/prabhupad26/100daysofML"},{"title":"armaank/archlectures","url":"https://github.com/armaank/archlectures"},{"title":"minhanh151/pre","url":"https://github.com/minhanh151/pre"},{"title":"yuuun/clip_pytorch","url":"https://github.com/yuuun/clip_pytorch"},{"title":"2024-MindSpore-1/Code2","url":"https://github.com/2024-MindSpore-1/Code2/tree/main/model-1/clip"},{"title":"lunaproject22/rpa","url":"https://github.com/lunaproject22/rpa"},{"title":"fiabdu/Commonly-Interesting-Images","url":"https://github.com/fiabdu/Commonly-Interesting-Images"},{"title":"iejMac/ScriptWriter","url":"https://github.com/iejMac/ScriptWriter"},{"title":"ZackPashkin/text2cartoon-pytorch-CLIP","url":"https://github.com/ZackPashkin/text2cartoon-pytorch-CLIP"},{"title":"bruthyu/bpt-vlm","url":"https://github.com/bruthyu/bpt-vlm"},{"title":"buyeah1109/finc","url":"https://github.com/buyeah1109/finc"},{"title":"pwc-1/Paper-8","url":"https://github.com/pwc-1/Paper-8/tree/main/clip"},{"title":"eify/open_clip","url":"https://github.com/eify/open_clip"},{"title":"2023-MindSpore-4/Code12","url":"https://github.com/2023-MindSpore-4/Code12/tree/main/MindFormers/clip"},{"title":"a736875071/clip-vit-large-patch14","url":"https://github.com/a736875071/clip-vit-large-patch14"},{"title":"nahidalam/open_clip","url":"https://github.com/nahidalam/open_clip"}],"syntology":{"n":20,"n_ran":16,"n_unverified":4,"n_pointer_only":16}}},{"leaderboard":"/sota/zero-shot-learning-on-dtd","slug":"zero-shot-learning-on-dtd","dataset":"DTD","dataset_url":"/dataset/dtd","rows_in_archive":2,"metrics":["Accuracy"],"first_row_in_archive_order":{"model":"ZLaP","paper_title":"Label Propagation for Zero-shot Classification with Vision-Language Models","paper_url":"/paper/label-propagation-for-zero-shot","paper_date":"2024-04-05","arxiv_id":"2404.04072","code_links":[{"title":"vladan-stojnic/zlap","url":"https://github.com/vladan-stojnic/zlap"}],"syntology":{"n":2,"n_ran":1,"n_unverified":1,"n_pointer_only":0}}},{"leaderboard":"/sota/zero-shot-learning-on-fgvc-aircraft","slug":"zero-shot-learning-on-fgvc-aircraft","dataset":"FGVC-Aircraft","dataset_url":"/dataset/fgvc-aircraft-1","rows_in_archive":2,"metrics":["Accuracy"],"first_row_in_archive_order":{"model":"ZLaP","paper_title":"Label Propagation for Zero-shot Classification with Vision-Language Models","paper_url":"/paper/label-propagation-for-zero-shot","paper_date":"2024-04-05","arxiv_id":"2404.04072","code_links":[{"title":"vladan-stojnic/zlap","url":"https://github.com/vladan-stojnic/zlap"}],"syntology":{"n":2,"n_ran":1,"n_unverified":1,"n_pointer_only":0}}},{"leaderboard":"/sota/zero-shot-learning-on-flowers-102","slug":"zero-shot-learning-on-flowers-102","dataset":"Flowers-102","dataset_url":"/dataset/oxford-102-flower","rows_in_archive":2,"metrics":["Accuracy"],"first_row_in_archive_order":{"model":"ZLaP","paper_title":"Label Propagation for Zero-shot Classification with Vision-Language Models","paper_url":"/paper/label-propagation-for-zero-shot","paper_date":"2024-04-05","arxiv_id":"2404.04072","code_links":[{"title":"vladan-stojnic/zlap","url":"https://github.com/vladan-stojnic/zlap"}],"syntology":{"n":2,"n_ran":1,"n_unverified":1,"n_pointer_only":0}}},{"leaderboard":"/sota/zero-shot-learning-on-food-101","slug":"zero-shot-learning-on-food-101","dataset":"Food-101","dataset_url":"/dataset/food-101","rows_in_archive":2,"metrics":["Accuracy"],"first_row_in_archive_order":{"model":"ZLaP*","paper_title":"Label Propagation for Zero-shot Classification with Vision-Language Models","paper_url":"/paper/label-propagation-for-zero-shot","paper_date":"2024-04-05","arxiv_id":"2404.04072","code_links":[{"title":"vladan-stojnic/zlap","url":"https://github.com/vladan-stojnic/zlap"}],"syntology":{"n":2,"n_ran":1,"n_unverified":1,"n_pointer_only":0}}},{"leaderboard":"/sota/zero-shot-learning-on-imagenet","slug":"zero-shot-learning-on-imagenet","dataset":"ImageNet","dataset_url":"/dataset/imagenet","rows_in_archive":2,"metrics":["Top 1 Accuracy"],"first_row_in_archive_order":{"model":"ZLaP","paper_title":"Label Propagation for Zero-shot Classification with Vision-Language Models","paper_url":"/paper/label-propagation-for-zero-shot","paper_date":"2024-04-05","arxiv_id":"2404.04072","code_links":[{"title":"vladan-stojnic/zlap","url":"https://github.com/vladan-stojnic/zlap"}],"syntology":{"n":2,"n_ran":1,"n_unverified":1,"n_pointer_only":0}}},{"leaderboard":"/sota/zero-shot-learning-on-oxford-102-flower","slug":"zero-shot-learning-on-oxford-102-flower","dataset":"Oxford 102 Flower","dataset_url":"/dataset/oxford-102-flower","rows_in_archive":2,"metrics":["average top-1 classification accuracy"],"first_row_in_archive_order":{"model":"SPOT","paper_title":"Synthetic Sample Selection for Generalized Zero-Shot Learning","paper_url":"/paper/synthetic-sample-selection-for-generalized","paper_date":"2023-04-06","arxiv_id":"2304.02846","code_links":[],"syntology":null}},{"leaderboard":"/sota/zero-shot-learning-on-oxford-iiit-pets","slug":"zero-shot-learning-on-oxford-iiit-pets","dataset":"Oxford-IIIT Pets","dataset_url":"/dataset/oxford-iiit-pets-1","rows_in_archive":2,"metrics":["Accuracy"],"first_row_in_archive_order":{"model":"ZLaP","paper_title":"Label Propagation for Zero-shot Classification with Vision-Language Models","paper_url":"/paper/label-propagation-for-zero-shot","paper_date":"2024-04-05","arxiv_id":"2404.04072","code_links":[{"title":"vladan-stojnic/zlap","url":"https://github.com/vladan-stojnic/zlap"}],"syntology":{"n":2,"n_ran":1,"n_unverified":1,"n_pointer_only":0}}},{"leaderboard":"/sota/zero-shot-learning-on-stanford-cars","slug":"zero-shot-learning-on-stanford-cars","dataset":"Stanford Cars","dataset_url":"/dataset/stanford-cars","rows_in_archive":2,"metrics":["Accuracy"],"first_row_in_archive_order":{"model":"ZLaP*","paper_title":"Label Propagation for Zero-shot Classification with Vision-Language Models","paper_url":"/paper/label-propagation-for-zero-shot","paper_date":"2024-04-05","arxiv_id":"2404.04072","code_links":[{"title":"vladan-stojnic/zlap","url":"https://github.com/vladan-stojnic/zlap"}],"syntology":{"n":2,"n_ran":1,"n_unverified":1,"n_pointer_only":0}}},{"leaderboard":"/sota/zero-shot-learning-on-sun397","slug":"zero-shot-learning-on-sun397","dataset":"SUN397","dataset_url":"/dataset/sun397","rows_in_archive":2,"metrics":["Accuracy"],"first_row_in_archive_order":{"model":"ZLaP*","paper_title":"Label Propagation for Zero-shot Classification with Vision-Language Models","paper_url":"/paper/label-propagation-for-zero-shot","paper_date":"2024-04-05","arxiv_id":"2404.04072","code_links":[{"title":"vladan-stojnic/zlap","url":"https://github.com/vladan-stojnic/zlap"}],"syntology":{"n":2,"n_ran":1,"n_unverified":1,"n_pointer_only":0}}},{"leaderboard":"/sota/zero-shot-learning-on-ucf101","slug":"zero-shot-learning-on-ucf101","dataset":"UCF101","dataset_url":"/dataset/ucf101","rows_in_archive":2,"metrics":["Accuracy"],"first_row_in_archive_order":{"model":"ZLaP","paper_title":"Label Propagation for Zero-shot Classification with Vision-Language Models","paper_url":"/paper/label-propagation-for-zero-shot","paper_date":"2024-04-05","arxiv_id":"2404.04072","code_links":[{"title":"vladan-stojnic/zlap","url":"https://github.com/vladan-stojnic/zlap"}],"syntology":{"n":2,"n_ran":1,"n_unverified":1,"n_pointer_only":0}}},{"leaderboard":"/sota/zero-shot-learning-on-voc-mlt","slug":"zero-shot-learning-on-voc-mlt","dataset":"VOC-MLT","dataset_url":"/dataset/voc-mlt","rows_in_archive":2,"metrics":["Average mAP"],"first_row_in_archive_order":{"model":"CLIP(ResNet-50)","paper_title":"Learning Transferable Visual Models From Natural Language Supervision","paper_url":"/paper/learning-transferable-visual-models-from","paper_date":"2021-02-26","arxiv_id":"2103.00020","code_links":[{"title":"openai/CLIP","url":"https://github.com/openai/CLIP"},{"title":"mlfoundations/open_clip","url":"https://github.com/mlfoundations/open_clip"},{"title":"towhee-io/towhee","url":"https://github.com/towhee-io/towhee"},{"title":"facebookresearch/vissl","url":"https://github.com/facebookresearch/vissl"},{"title":"alibaba/EasyNLP","url":"https://github.com/alibaba/EasyNLP"},{"title":"apple/ml-mobileclip","url":"https://github.com/apple/ml-mobileclip"},{"title":"OML-Team/open-metric-learning","url":"https://github.com/OML-Team/open-metric-learning"},{"title":"FreddeFrallan/Multilingual-CLIP","url":"https://github.com/FreddeFrallan/Multilingual-CLIP"},{"title":"eps696/aphantasia","url":"https://github.com/eps696/aphantasia"},{"title":"muzairkhattak/multimodal-prompt-learning","url":"https://github.com/muzairkhattak/multimodal-prompt-learning"},{"title":"moein-shariatnia/OpenAI-CLIP","url":"https://github.com/moein-shariatnia/OpenAI-CLIP"},{"title":"facebookresearch/brainmagick","url":"https://github.com/facebookresearch/brainmagick"},{"title":"PaddlePaddle/PASSL","url":"https://github.com/PaddlePaddle/PASSL/blob/main/docs/Train_CLIP_model.md"},{"title":"taited/clip-score","url":"https://github.com/taited/clip-score"},{"title":"azshue/TPT","url":"https://github.com/azshue/TPT"},{"title":"clip-italian/clip-italian","url":"https://github.com/clip-italian/clip-italian"},{"title":"dhansmair/flamingo-mini","url":"https://github.com/dhansmair/flamingo-mini"},{"title":"ylqi/count-anything","url":"https://github.com/ylqi/count-anything"},{"title":"ml-jku/cloob","url":"https://github.com/ml-jku/cloob"},{"title":"sberbank-ai/ru-clip","url":"https://github.com/sberbank-ai/ru-clip"},{"title":"ai-forever/ru-clip","url":"https://github.com/ai-forever/ru-clip"},{"title":"Kaushalya/medclip","url":"https://github.com/Kaushalya/medclip"},{"title":"ajayjain/vectorascent","url":"https://github.com/ajayjain/vectorascent"},{"title":"linjieli222/hero_video_feature_extractor","url":"https://github.com/linjieli222/hero_video_feature_extractor"},{"title":"borisdayma/clip-jax","url":"https://github.com/borisdayma/clip-jax"},{"title":"sajjjadayobi/CLIPfa","url":"https://github.com/sajjjadayobi/CLIPfa"},{"title":"mertyg/post-hoc-cbm","url":"https://github.com/mertyg/post-hoc-cbm"},{"title":"mlbio-epfl/turtle","url":"https://github.com/mlbio-epfl/turtle"},{"title":"rinnakk/japanese-clip","url":"https://github.com/rinnakk/japanese-clip"},{"title":"salesforce/pb-ovd","url":"https://github.com/salesforce/pb-ovd"},{"title":"facebookresearch/clip-rocket","url":"https://github.com/facebookresearch/clip-rocket"},{"title":"sincerass/mvlpt","url":"https://github.com/sincerass/mvlpt"},{"title":"ericyinyzy/vlattack","url":"https://github.com/ericyinyzy/vlattack"},{"title":"redcaps-dataset/redcaps-downloader","url":"https://github.com/redcaps-dataset/redcaps-downloader"},{"title":"bespontaneous/proteus-pytorch","url":"https://github.com/bespontaneous/proteus-pytorch"},{"title":"shunk031/simple-aesthetics-predictor","url":"https://github.com/shunk031/simple-aesthetics-predictor"},{"title":"giantseaweed/decree","url":"https://github.com/giantseaweed/decree"},{"title":"sithu31296/simple-object-tracking","url":"https://github.com/sithu31296/simple-object-tracking"},{"title":"Gahyeonkim09/AAPL","url":"https://github.com/Gahyeonkim09/AAPL"},{"title":"filipbasara0/simple-clip","url":"https://github.com/filipbasara0/simple-clip"},{"title":"michi-3000/eyeclip","url":"https://github.com/michi-3000/eyeclip"},{"title":"baskargroup/Arboretum","url":"https://github.com/baskargroup/Arboretum"},{"title":"baskargroup/biotrove","url":"https://github.com/baskargroup/biotrove"},{"title":"kynkaat/role-of-imagenet-classes-in-fid","url":"https://github.com/kynkaat/role-of-imagenet-classes-in-fid"},{"title":"SforAiDl/CountCLIP","url":"https://github.com/SforAiDl/CountCLIP"},{"title":"mainaksingha01/applenet","url":"https://github.com/mainaksingha01/applenet"},{"title":"zhangxu0963/npc","url":"https://github.com/zhangxu0963/npc"},{"title":"mainaksingha01/odg-clip","url":"https://github.com/mainaksingha01/odg-clip"},{"title":"jhaprince/multibully","url":"https://github.com/jhaprince/multibully"},{"title":"klemens-floege/oneprot","url":"https://github.com/klemens-floege/oneprot"},{"title":"leolee99/CLIP_ITM","url":"https://github.com/leolee99/CLIP_ITM"},{"title":"madrylab/pretraining-distribution-shift-robustness","url":"https://github.com/madrylab/pretraining-distribution-shift-robustness"},{"title":"AndresPMD/Clip_CMR","url":"https://github.com/AndresPMD/Clip_CMR"},{"title":"fastscience-ai/medflamingo","url":"https://github.com/fastscience-ai/medflamingo"},{"title":"buyeah1109/KEN","url":"https://github.com/buyeah1109/KEN"},{"title":"pseulki/rococo","url":"https://github.com/pseulki/rococo"},{"title":"shkarupa-alex/tfclip","url":"https://github.com/shkarupa-alex/tfclip"},{"title":"IMvision12/keras-vision-models","url":"https://github.com/IMvision12/keras-vision-models"},{"title":"brown-palm/ObjectPrompt","url":"https://github.com/brown-palm/ObjectPrompt"},{"title":"YvanG/VQGAN-CLIP","url":"https://github.com/YvanG/VQGAN-CLIP"},{"title":"ramanakshay/clip","url":"https://github.com/ramanakshay/clip"},{"title":"NYU-DICE-Lab/open_clip","url":"https://github.com/NYU-DICE-Lab/open_clip"},{"title":"shivammehta25/clip","url":"https://github.com/shivammehta25/clip"},{"title":"minhanh151/respro","url":"https://github.com/minhanh151/respro"},{"title":"nopperl/clip_arxiv_pmc","url":"https://github.com/nopperl/clip_arxiv_pmc"},{"title":"s-a-malik/multi-few","url":"https://github.com/s-a-malik/multi-few"},{"title":"prabhupad26/100daysofML","url":"https://github.com/prabhupad26/100daysofML"},{"title":"armaank/archlectures","url":"https://github.com/armaank/archlectures"},{"title":"minhanh151/pre","url":"https://github.com/minhanh151/pre"},{"title":"yuuun/clip_pytorch","url":"https://github.com/yuuun/clip_pytorch"},{"title":"2024-MindSpore-1/Code2","url":"https://github.com/2024-MindSpore-1/Code2/tree/main/model-1/clip"},{"title":"lunaproject22/rpa","url":"https://github.com/lunaproject22/rpa"},{"title":"fiabdu/Commonly-Interesting-Images","url":"https://github.com/fiabdu/Commonly-Interesting-Images"},{"title":"iejMac/ScriptWriter","url":"https://github.com/iejMac/ScriptWriter"},{"title":"ZackPashkin/text2cartoon-pytorch-CLIP","url":"https://github.com/ZackPashkin/text2cartoon-pytorch-CLIP"},{"title":"bruthyu/bpt-vlm","url":"https://github.com/bruthyu/bpt-vlm"},{"title":"buyeah1109/finc","url":"https://github.com/buyeah1109/finc"},{"title":"pwc-1/Paper-8","url":"https://github.com/pwc-1/Paper-8/tree/main/clip"},{"title":"eify/open_clip","url":"https://github.com/eify/open_clip"},{"title":"2023-MindSpore-4/Code12","url":"https://github.com/2023-MindSpore-4/Code12/tree/main/MindFormers/clip"},{"title":"a736875071/clip-vit-large-patch14","url":"https://github.com/a736875071/clip-vit-large-patch14"},{"title":"nahidalam/open_clip","url":"https://github.com/nahidalam/open_clip"}],"syntology":{"n":20,"n_ran":16,"n_unverified":4,"n_pointer_only":16}}},{"leaderboard":"/sota/zero-shot-learning-on-apy-0-shot","slug":"zero-shot-learning-on-apy-0-shot","dataset":"aPY - 0-Shot","dataset_url":"/dataset/apy","rows_in_archive":1,"metrics":["Top-1"],"first_row_in_archive_order":{"model":"ZSL-KG","paper_title":"Zero-Shot Learning with Common Sense Knowledge Graphs","paper_url":"/paper/zero-shot-learning-with-common-sense","paper_date":"2020-06-18","arxiv_id":"2006.10713","code_links":[{"title":"BatsResearch/zsl-kg","url":"https://github.com/BatsResearch/zsl-kg"},{"title":"BatsResearch/nayak-arxiv20-code","url":"https://github.com/BatsResearch/nayak-arxiv20-code"},{"title":"batsresearch/nayak-tmlr22-code","url":"https://github.com/batsresearch/nayak-tmlr22-code"}],"syntology":null}},{"leaderboard":"/sota/zero-shot-learning-on-cub-200-0-shot-learning-1","slug":"zero-shot-learning-on-cub-200-0-shot-learning-1","dataset":"CUB-200 - 0-Shot Learning","dataset_url":"/dataset/cub-200-2011","rows_in_archive":1,"metrics":["Average Per-Class Accuracy"],"first_row_in_archive_order":{"model":"zsl_ADA","paper_title":"A Generative Framework for Zero-Shot Learning with Adversarial Domain Adaptation","paper_url":"/paper/a-generative-framework-for-zero-shot-learning","paper_date":"2019-06-07","arxiv_id":"1906.03038","code_links":[{"title":"vkkhare/ZSL-ADA","url":"https://github.com/vkkhare/ZSL-ADA"}],"syntology":null}},{"leaderboard":"/sota/zero-shot-learning-on-eurosat","slug":"zero-shot-learning-on-eurosat","dataset":"EuroSAT","dataset_url":"/dataset/eurosat","rows_in_archive":1,"metrics":["Accuracy"],"first_row_in_archive_order":{"model":"ZLaP*","paper_title":"Label Propagation for Zero-shot Classification with Vision-Language Models","paper_url":"/paper/label-propagation-for-zero-shot","paper_date":"2024-04-05","arxiv_id":"2404.04072","code_links":[{"title":"vladan-stojnic/zlap","url":"https://github.com/vladan-stojnic/zlap"}],"syntology":{"n":2,"n_ran":1,"n_unverified":1,"n_pointer_only":0}}},{"leaderboard":"/sota/zero-shot-learning-on-gdscv2","slug":"zero-shot-learning-on-gdscv2","dataset":"GDSCv2","dataset_url":null,"rows_in_archive":1,"metrics":["Pearson correlation coefficient (PCC)"],"first_row_in_archive_order":{"model":"MSDA","paper_title":"Zero-shot Learning of Drug Response Prediction for Preclinical Drug Screening","paper_url":"/paper/zero-shot-learning-of-drug-response","paper_date":"2023-10-05","arxiv_id":"2310.12996","code_links":[{"title":"drugd/msda","url":"https://github.com/drugd/msda"}],"syntology":null}},{"leaderboard":"/sota/zero-shot-learning-on-how2qa","slug":"zero-shot-learning-on-how2qa","dataset":"How2QA","dataset_url":"/dataset/how2qa","rows_in_archive":1,"metrics":["Accuracy"],"first_row_in_archive_order":{"model":"SeViLA","paper_title":null,"paper_url":null,"paper_date":"","arxiv_id":null,"code_links":[],"syntology":null}},{"leaderboard":"/sota/zero-shot-learning-on-imagenet-cn","slug":"zero-shot-learning-on-imagenet-cn","dataset":"ImageNet_CN","dataset_url":"/dataset/imagenet-cn","rows_in_archive":1,"metrics":["Accuracy"],"first_row_in_archive_order":{"model":"$M^2$-Encoder","paper_title":"M2-Encoder: Advancing Bilingual Image-Text Understanding by Large-scale Efficient Pretraining","paper_url":"/paper/boldsymbol-m-2-encoder-advancing-bilingual","paper_date":"2024-01-29","arxiv_id":"2401.15896","code_links":[{"title":"alipay/Ant-Multi-Modal-Framework","url":"https://github.com/alipay/Ant-Multi-Modal-Framework/tree/main/prj/M2_Encoder"}],"syntology":null}},{"leaderboard":"/sota/zero-shot-learning-on-ivqa","slug":"zero-shot-learning-on-ivqa","dataset":"iVQA","dataset_url":"/dataset/ivqa","rows_in_archive":1,"metrics":["Accuracy"],"first_row_in_archive_order":{"model":"FrozenBiLM","paper_title":"Zero-Shot Video Question Answering via Frozen Bidirectional Language Models","paper_url":"/paper/zero-shot-video-question-answering-via-frozen","paper_date":"2022-06-16","arxiv_id":"2206.08155","code_links":[{"title":"antoyang/FrozenBiLM","url":"https://github.com/antoyang/FrozenBiLM"},{"title":"klauscc/dam","url":"https://github.com/klauscc/dam"},{"title":"sts-vlcc/sts-vlcc","url":"https://github.com/sts-vlcc/sts-vlcc"}],"syntology":{"n":34,"n_ran":14,"n_unverified":20,"n_pointer_only":1}}},{"leaderboard":"/sota/zero-shot-learning-on-lsmdc","slug":"zero-shot-learning-on-lsmdc","dataset":"LSMDC","dataset_url":"/dataset/lsmdc","rows_in_archive":1,"metrics":["Accuracy"],"first_row_in_archive_order":{"model":"FrozenBiLM","paper_title":"Zero-Shot Video Question Answering via Frozen Bidirectional Language Models","paper_url":"/paper/zero-shot-video-question-answering-via-frozen","paper_date":"2022-06-16","arxiv_id":"2206.08155","code_links":[{"title":"antoyang/FrozenBiLM","url":"https://github.com/antoyang/FrozenBiLM"},{"title":"klauscc/dam","url":"https://github.com/klauscc/dam"},{"title":"sts-vlcc/sts-vlcc","url":"https://github.com/sts-vlcc/sts-vlcc"}],"syntology":{"n":34,"n_ran":14,"n_unverified":20,"n_pointer_only":1}}},{"leaderboard":"/sota/zero-shot-learning-on-mit-states-1","slug":"zero-shot-learning-on-mit-states-1","dataset":"MIT-States","dataset_url":"/dataset/mit-states","rows_in_archive":1,"metrics":["A-acc"],"first_row_in_archive_order":{"model":"CZSL","paper_title":"LOCL: Learning Object-Attribute Composition using Localization","paper_url":"/paper/locl-learning-object-attribute-composition","paper_date":"2022-10-07","arxiv_id":"2210.03780","code_links":[{"title":"satish1901/LOCL-Learning-Object-Attribute-Composition-using-Localization","url":"https://github.com/satish1901/LOCL-Learning-Object-Attribute-Composition-using-Localization"}],"syntology":null}},{"leaderboard":"/sota/zero-shot-learning-on-msrvtt-qa","slug":"zero-shot-learning-on-msrvtt-qa","dataset":"MSRVTT-QA","dataset_url":"/dataset/msrvtt-qa","rows_in_archive":1,"metrics":["Accuracy"],"first_row_in_archive_order":{"model":"HiTeA","paper_title":"HiTeA: Hierarchical Temporal-Aware Video-Language Pre-training","paper_url":"/paper/hitea-hierarchical-temporal-aware-video","paper_date":"2022-12-30","arxiv_id":"2212.14546","code_links":[],"syntology":null}},{"leaderboard":"/sota/zero-shot-learning-on-msvd-qa","slug":"zero-shot-learning-on-msvd-qa","dataset":"MSVD-QA","dataset_url":"/dataset/msvd-qa","rows_in_archive":1,"metrics":["Accuracy"],"first_row_in_archive_order":{"model":"HiTeA","paper_title":"HiTeA: Hierarchical Temporal-Aware Video-Language Pre-training","paper_url":"/paper/hitea-hierarchical-temporal-aware-video","paper_date":"2022-12-30","arxiv_id":"2212.14546","code_links":[],"syntology":null}},{"leaderboard":"/sota/zero-shot-learning-on-pascal-context","slug":"zero-shot-learning-on-pascal-context","dataset":"PASCAL Context","dataset_url":"/dataset/pascal-context","rows_in_archive":1,"metrics":["k=10 mIOU"],"first_row_in_archive_order":{"model":"ZS3Net","paper_title":"Zero-Shot Semantic Segmentation","paper_url":"/paper/190600817","paper_date":"2019-06-03","arxiv_id":"1906.00817","code_links":[{"title":"valeoai/ZS3","url":"https://github.com/valeoai/ZS3"},{"title":"bcmi/CaGNetv2-Zero-Shot-Semantic-Segmentation","url":"https://github.com/bcmi/CaGNetv2-Zero-Shot-Semantic-Segmentation"}],"syntology":null}},{"leaderboard":"/sota/zero-shot-learning-on-snips","slug":"zero-shot-learning-on-snips","dataset":"SNIPS","dataset_url":"/dataset/snips","rows_in_archive":1,"metrics":["Accuracy"],"first_row_in_archive_order":{"model":"ZSL-KG","paper_title":"Zero-Shot Learning with Common Sense Knowledge Graphs","paper_url":"/paper/zero-shot-learning-with-common-sense","paper_date":"2020-06-18","arxiv_id":"2006.10713","code_links":[{"title":"BatsResearch/zsl-kg","url":"https://github.com/BatsResearch/zsl-kg"},{"title":"BatsResearch/nayak-arxiv20-code","url":"https://github.com/BatsResearch/nayak-arxiv20-code"},{"title":"batsresearch/nayak-tmlr22-code","url":"https://github.com/batsresearch/nayak-tmlr22-code"}],"syntology":null}},{"leaderboard":"/sota/zero-shot-learning-on-tvqa","slug":"zero-shot-learning-on-tvqa","dataset":"TVQA","dataset_url":"/dataset/tvqa","rows_in_archive":1,"metrics":["Accuracy"],"first_row_in_archive_order":{"model":"VideoChat2","paper_title":"MVBench: A Comprehensive Multi-modal Video Understanding Benchmark","paper_url":"/paper/mvbench-a-comprehensive-multi-modal-video","paper_date":"2023-11-28","arxiv_id":"2311.17005","code_links":[{"title":"opengvlab/ask-anything","url":"https://github.com/opengvlab/ask-anything"},{"title":"magic-research/PLLaVA","url":"https://github.com/magic-research/PLLaVA"},{"title":"bytedance/tarsier","url":"https://github.com/bytedance/tarsier"}],"syntology":{"n":10,"n_ran":7,"n_unverified":3,"n_pointer_only":0}}}],"datasets":[{"url":"/dataset/cifar-10","name":"CIFAR-10","full_name":"CIFAR-10","num_papers_in_archive":16145},{"url":"/dataset/imagenet","name":"ImageNet","full_name":"","num_papers_in_archive":15430},{"url":"/dataset/cifar-100","name":"CIFAR-100","full_name":"","num_papers_in_archive":9045},{"url":"/dataset/cub-200-2011","name":"CUB-200-2011","full_name":"Caltech-UCSD Birds-200-2011","num_papers_in_archive":2235},{"url":"/dataset/ucf101","name":"UCF101","full_name":"UCF101 Human Actions dataset","num_papers_in_archive":1863},{"url":"/dataset/oxford-102-flower","name":"Oxford 102 Flower","full_name":"102 Category Flower Dataset","num_papers_in_archive":1307},{"url":"/dataset/dtd","name":"DTD","full_name":"Describable Textures Dataset","num_papers_in_archive":870},{"url":"/dataset/food-101","name":"Food-101","full_name":"","num_papers_in_archive":805},{"url":"/dataset/stanford-cars","name":"Stanford Cars","full_name":"","num_papers_in_archive":790},{"url":"/dataset/caltech-101","name":"Caltech-101","full_name":"","num_papers_in_archive":709},{"url":"/dataset/eurosat","name":"EuroSAT","full_name":"EuroSAT","num_papers_in_archive":687},{"url":"/dataset/fgvc-aircraft-1","name":"FGVC-Aircraft","full_name":"","num_papers_in_archive":520},{"url":"/dataset/pascal-context","name":"PASCAL Context","full_name":"","num_papers_in_archive":323},{"url":"/dataset/awa-1","name":"AwA","full_name":"Animals with Attributes","num_papers_in_archive":264},{"url":"/dataset/snips","name":"SNIPS","full_name":"SNIPS Natural Language Understanding benchmark","num_papers_in_archive":256},{"url":"/dataset/awa2-1","name":"AwA2","full_name":"Animals with Attributes 2","num_papers_in_archive":231},{"url":"/dataset/apy","name":"aPY","full_name":"Attribute Pascal and Yahoo","num_papers_in_archive":147},{"url":"/dataset/tvqa","name":"TVQA","full_name":"TVQA","num_papers_in_archive":146},{"url":"/dataset/lsmdc","name":"LSMDC","full_name":"Large Scale Movie Description Challenge","num_papers_in_archive":126},{"url":"/dataset/mit-states","name":"MIT-States","full_name":"","num_papers_in_archive":91},{"url":"/dataset/msrvtt-qa","name":"MSRVTT-QA","full_name":"","num_papers_in_archive":66},{"url":"/dataset/msvd-qa","name":"MSVD-QA","full_name":"","num_papers_in_archive":61},{"url":"/dataset/tvqa-1","name":"TVQA+","full_name":"","num_papers_in_archive":60},{"url":"/dataset/oxford-iiit-pets-1","name":"Oxford-IIIT Pets","full_name":"","num_papers_in_archive":59},{"url":"/dataset/sun397","name":"SUN397","full_name":"SUN397","num_papers_in_archive":52},{"url":"/dataset/ocnli","name":"OCNLI","full_name":"Original Chinese Natural Language Inference","num_papers_in_archive":44},{"url":"/dataset/sun-attribute","name":"SUN Attribute","full_name":"SUN Attribute","num_papers_in_archive":38},{"url":"/dataset/how2qa","name":"How2QA","full_name":"How2QA","num_papers_in_archive":28},{"url":"/dataset/eurlex57k","name":"EURLEX57K","full_name":"","num_papers_in_archive":23},{"url":"/dataset/ivqa","name":"iVQA","full_name":"Instructional Video Question Answering","num_papers_in_archive":22},{"url":"/dataset/medconceptsqa","name":"MedConceptsQA","full_name":"","num_papers_in_archive":13},{"url":"/dataset/coco-mlt","name":"COCO-MLT","full_name":"","num_papers_in_archive":12},{"url":"/dataset/voc-mlt","name":"VOC-MLT","full_name":"","num_papers_in_archive":12},{"url":"/dataset/lad","name":"LAD","full_name":"Large-scale Attribute Dataset","num_papers_in_archive":10},{"url":"/dataset/ao-clevr","name":"AO-CLEVr","full_name":null,"num_papers_in_archive":6},{"url":"/dataset/tasty-videos","name":"Tasty Videos","full_name":"","num_papers_in_archive":5},{"url":"/dataset/imagenet-cn","name":"ImageNet_CN","full_name":"Chinese ImageNet Classification","num_papers_in_archive":3},{"url":"/dataset/xl-r2r","name":"XL-R2R","full_name":"Cross-lingual Room-to-Room","num_papers_in_archive":2},{"url":"/dataset/edge-map-345c","name":"Edge-Map-345C","full_name":null,"num_papers_in_archive":1},{"url":"/dataset/goz","name":"GOZ","full_name":"Generic Object ZSL Dataset","num_papers_in_archive":1},{"url":"/dataset/pubchem18","name":"PubChem18","full_name":"PubChem 2018","num_papers_in_archive":1},{"url":"/dataset/sequence-consistency-evaluation-sce-tests","name":"Sequence Consistency Evaluation (SCE) tests","full_name":"","num_papers_in_archive":1},{"url":"/dataset/events-classification-biotech","name":"Events classification - Biotech news","full_name":"","num_papers_in_archive":0}],"subtasks":[{"url":"/task/action-recognition","name":"Temporal Action Localization"},{"url":"/task/compositional-zero-shot-learning","name":"Compositional Zero-Shot Learning"},{"url":"/task/generalized-zero-shot-learning","name":"Generalized Zero-Shot Learning"},{"url":"/task/gzsl-video-classification","name":"GZSL Video Classification"},{"url":"/task/multi-label-zero-shot-learning","name":"Multi-label zero-shot learning"},{"url":"/task/transductive-zero-shot-classification","name":"Transductive Zero-Shot Classification"},{"url":"/task/zero-shot-gesture-recognition","name":"Zero-shot gesture recognition"}],"parent_tasks":[],"papers":{"order":"repositories listed in the archive (desc), then date (desc); the archive holds no stars","population":"papers tagged with this task that list at least one repository in the archive","shown":30,"of":787,"tagged_in_all":1864,"items":[{"url":"/paper/learning-transferable-visual-models-from","title":"Learning Transferable Visual Models From Natural Language Supervision","date":"2021-02-26","arxiv_id":"2103.00020","repositories_listed":82,"syntology":{"n":20,"n_ran":16,"n_unverified":4,"n_pointer_only":16}},{"url":"/paper/language-models-are-few-shot-learners","title":"Language Models are Few-Shot Learners","date":"2020-05-28","arxiv_id":"2005.14165","repositories_listed":67,"syntology":{"n":65,"n_ran":15,"n_unverified":50,"n_pointer_only":4}},{"url":"/paper/llama-open-and-efficient-foundation-language-1","title":"LLaMA: Open and Efficient Foundation Language Models","date":"2023-02-27","arxiv_id":"2302.13971","repositories_listed":57,"syntology":{"n":58,"n_ran":26,"n_unverified":32,"n_pointer_only":4}},{"url":"/paper/prototypical-networks-for-few-shot-learning","title":"Prototypical Networks for Few-shot Learning","date":"2017-03-15","arxiv_id":"1703.05175","repositories_listed":43,"syntology":{"n":64,"n_ran":49,"n_unverified":15,"n_pointer_only":18}},{"url":"/paper/biobert-a-pre-trained-biomedical-language","title":"BioBERT: a pre-trained biomedical language representation model for biomedical text mining","date":"2019-01-25","arxiv_id":"1901.08746","repositories_listed":19,"syntology":{"n":25,"n_ran":4,"n_unverified":21,"n_pointer_only":1}},{"url":"/paper/learning-to-compare-relation-network-for-few","title":"Learning to Compare: Relation Network for Few-Shot Learning","date":"2017-11-16","arxiv_id":"1711.06025","repositories_listed":13,"syntology":{"n":2,"n_ran":1,"n_unverified":1,"n_pointer_only":1}},{"url":"/paper/gpt-4-technical-report-1","title":"GPT-4 Technical Report","date":"2023-03-15","arxiv_id":"2303.08774","repositories_listed":11,"syntology":{"n":5,"n_ran":2,"n_unverified":3,"n_pointer_only":1}},{"url":"/paper/cpm-a-large-scale-generative-chinese-pre","title":"CPM: A Large-scale Generative Chinese Pre-trained Language Model","date":"2020-12-01","arxiv_id":"2012.00413","repositories_listed":10,"syntology":null},{"url":"/paper/zero-shot-learning-a-comprehensive-evaluation","title":"Zero-Shot Learning -- A Comprehensive Evaluation of the Good, the Bad and the Ugly","date":"2017-07-03","arxiv_id":"1707.00600","repositories_listed":10,"syntology":null},{"url":"/paper/learning-deep-representations-of-fine-grained","title":"Learning Deep Representations of Fine-grained Visual Descriptions","date":"2016-05-17","arxiv_id":"1605.05395","repositories_listed":9,"syntology":null},{"url":"/paper/finetuned-language-models-are-zero-shot","title":"Finetuned Language Models Are Zero-Shot Learners","date":"2021-09-03","arxiv_id":"2109.01652","repositories_listed":8,"syntology":{"n":1,"n_ran":0,"n_unverified":1,"n_pointer_only":0}},{"url":"/paper/sampling-matters-in-deep-embedding-learning","title":"Sampling Matters in Deep Embedding Learning","date":"2017-06-23","arxiv_id":"1706.07567","repositories_listed":6,"syntology":{"n":2,"n_ran":2,"n_unverified":0,"n_pointer_only":0}},{"url":"/paper/reproducible-scaling-laws-for-contrastive","title":"Reproducible scaling laws for contrastive language-image learning","date":"2022-12-14","arxiv_id":"2212.07143","repositories_listed":5,"syntology":{"n":3,"n_ran":3,"n_unverified":0,"n_pointer_only":3}},{"url":"/paper/laion-5b-an-open-large-scale-dataset-for-1","title":"LAION-5B: An open large-scale dataset for training next generation image-text models","date":"2022-10-16","arxiv_id":"2210.08402","repositories_listed":5,"syntology":{"n":18,"n_ran":4,"n_unverified":14,"n_pointer_only":3}},{"url":"/paper/flamingo-a-visual-language-model-for-few-shot-1","title":"Flamingo: a Visual Language Model for Few-Shot Learning","date":"2022-04-29","arxiv_id":"2204.14198","repositories_listed":5,"syntology":{"n":24,"n_ran":18,"n_unverified":6,"n_pointer_only":7}},{"url":"/paper/improving-zero-shot-learning-by-mitigating","title":"Improving zero-shot learning by mitigating the hubness problem","date":"2014-12-20","arxiv_id":"1412.6568","repositories_listed":5,"syntology":null},{"url":"/paper/llms-as-zero-shot-graph-learners-alignment-of","title":"LLMs as Zero-shot Graph Learners: Alignment of GNN Representations with LLM Token Embeddings","date":"2024-08-25","arxiv_id":"2408.14512","repositories_listed":4,"syntology":{"n":9,"n_ran":1,"n_unverified":8,"n_pointer_only":0}},{"url":"/paper/bioscan-clip-bridging-vision-and-genomics-for","title":"CLIBD: Bridging Vision and Genomics for Biodiversity Monitoring at Scale","date":"2024-05-27","arxiv_id":"2405.17537","repositories_listed":4,"syntology":{"n":12,"n_ran":4,"n_unverified":8,"n_pointer_only":0}},{"url":"/paper/your-diffusion-model-is-secretly-a-zero-shot","title":"Your Diffusion Model is Secretly a Zero-Shot Classifier","date":"2023-03-28","arxiv_id":"2303.16203","repositories_listed":4,"syntology":{"n":2,"n_ran":2,"n_unverified":0,"n_pointer_only":1}},{"url":"/paper/prompt-injection-parameterization-of-fixed","title":"Prompt Injection: Parameterization of Fixed Inputs","date":"2022-05-31","arxiv_id":"2206.11349","repositories_listed":4,"syntology":{"n":3,"n_ran":1,"n_unverified":2,"n_pointer_only":0}},{"url":"/paper/supervision-exists-everywhere-a-data-1","title":"Supervision Exists Everywhere: A Data Efficient Contrastive Language-Image Pre-training Paradigm","date":"2021-10-11","arxiv_id":"2110.05208","repositories_listed":4,"syntology":{"n":3,"n_ran":3,"n_unverified":0,"n_pointer_only":3}},{"url":"/paper/zero-shot-user-intent-detection-via-capsule","title":"Zero-shot User Intent Detection via Capsule Neural Networks","date":"2018-09-02","arxiv_id":"1809.00385","repositories_listed":4,"syntology":null},{"url":"/paper/feature-generating-networks-for-zero-shot","title":"Feature Generating Networks for Zero-Shot Learning","date":"2017-12-04","arxiv_id":"1712.00981","repositories_listed":4,"syntology":null},{"url":"/paper/semantic-autoencoder-for-zero-shot-learning","title":"Semantic Autoencoder for Zero-Shot Learning","date":"2017-04-26","arxiv_id":"1704.08345","repositories_listed":4,"syntology":null},{"url":"/paper/learning-a-deep-embedding-model-for-zero-shot","title":"Learning a Deep Embedding Model for Zero-Shot Learning","date":"2016-11-15","arxiv_id":"1611.05088","repositories_listed":4,"syntology":null},{"url":"/paper/benchmark-evaluations-applications-and","title":"A Survey of State of the Art Large Vision Language Models: Alignment, Benchmark, Evaluations and Challenges","date":"2025-01-04","arxiv_id":"2501.02189","repositories_listed":3,"syntology":null},{"url":"/paper/mobileclip-fast-image-text-models-through","title":"MobileCLIP: Fast Image-Text Models through Multi-Modal Reinforced Training","date":"2023-11-28","arxiv_id":"2311.17049","repositories_listed":3,"syntology":{"n":1,"n_ran":1,"n_unverified":0,"n_pointer_only":1}},{"url":"/paper/mvbench-a-comprehensive-multi-modal-video","title":"MVBench: A Comprehensive Multi-modal Video Understanding Benchmark","date":"2023-11-28","arxiv_id":"2311.17005","repositories_listed":3,"syntology":{"n":10,"n_ran":7,"n_unverified":3,"n_pointer_only":0}},{"url":"/paper/anomalyclip-object-agnostic-prompt-learning","title":"AnomalyCLIP: Object-agnostic Prompt Learning for Zero-shot Anomaly Detection","date":"2023-10-29","arxiv_id":"2310.18961","repositories_listed":3,"syntology":{"n":32,"n_ran":14,"n_unverified":18,"n_pointer_only":0}},{"url":"/paper/time-llm-time-series-forecasting-by","title":"Time-LLM: Time Series Forecasting by Reprogramming Large Language Models","date":"2023-10-03","arxiv_id":"2310.01728","repositories_listed":3,"syntology":{"n":9,"n_ran":7,"n_unverified":2,"n_pointer_only":0}}],"syntology_records":21,"syntology_note":"a paper without a record is not a recorded non-run: it may lack an arXiv id or simply be absent from the graph layer"},"description_links":{"kept":4,"unwrapped_to_text":0,"bare_urls_linked":0,"relative_images_dropped":0,"rule":"internal links are kept only when the target slug exists in the catalog"},"syntology":{"read_at":"2026-09-24T18:15:14+00:00","claim":"Per-sample execution status on synthesized fixtures ('ran N of M samples'); not a correctness claim and not a ranking signal.","status_vocabulary":{"ran_honours":"ran, honoured the contract we drafted","ran_violates":"ran, violated the contract we drafted","ran_draft_wrong":"ran; our contract draft was wrong, not the code","ran_fixture":"ran; our fixture could not drive it","ran":"ran on a synthesized input","unverified":"unverified (harvested, no recorded run)"}},"not_shown":{"libraries":"the archive has no per-task library table","trend_sparklines":"the Trend column of the benchmarks table was a rendered image; it is not in the archive","social_and_latest_sorts":"stars and social signals are not in the archive"}}