{"url":"/sota/image-classification-on-resisc45","task":{"name":"Image Classification","url":"/task/image-classification","note":null},"dataset":{"name":"RESISC45","url":"/dataset/resisc45"},"category":"Computer Vision","categories":["Adversarial","Computer Vision"],"category_note":null,"description":"**Image Classification** is a fundamental task in vision recognition that aims to understand and categorize an image as a whole under a specific label. Unlike [object detection](/task/object-detection), which involves classification and location of multiple objects within an image, image classification typically pertains to single-object images. When the classification becomes highly detailed or reaches instance-level, it is often referred to as [image retrieval](/task/image-retrieval), which also involves finding similar images in a large database.\r\n\r\n\r\n<span class=\"description-source\">Source: [Metamorphic Testing for Object Detection Systems ](https://arxiv.org/abs/1912.12162)</span>","description_from":"task","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","rank":"the archive's row order at snapshot; not re-ranked","rows_end_at":"2025-07-28","rows_withheld_as_spam":0,"metric_values":"the archive's strings, untouched"},"metrics":["Top 1 Accuracy","F1","zero-shot Acc"],"metric_direction":{"note":"inferred from the metric name only (the archive records no direction); null = not inferred, chart draws points only","by_metric":{"Top 1 Accuracy":"higher","F1":"higher","zero-shot Acc":"higher"}},"counts":{"rows":20,"rows_with_code":19,"rows_with_paper_page":19,"rows_dated":19,"rows_using_additional_data":11},"rows":[{"rank_in_archive_order":1,"model":"ResNet50","metrics":{"Top 1 Accuracy":"96.83"},"uses_additional_data":true,"paper_date":"2019-11-15","paper":"/paper/in-domain-representation-learning-for-remote-1","paper_url":"https://arxiv.org/abs/1911.06721v1","paper_title":"In-domain representation learning for remote sensing","code":"https://github.com/google-research/google-research/tree/master/remote_sensing_representations","n_code_links":1,"syntology":null},{"rank_in_archive_order":2,"model":"LWGANet L2","metrics":{"Top 1 Accuracy":"96.17"},"uses_additional_data":false,"paper_date":"2025-01-17","paper":"/paper/lwganet-a-lightweight-group-attention","paper_url":"https://arxiv.org/abs/2501.10040v1","paper_title":"LWGANet: A Lightweight Group Attention Backbone for Remote Sensing Visual Tasks","code":"https://github.com/lwcver/lwganet","n_code_links":1,"syntology":null},{"rank_in_archive_order":3,"model":"DecoupleNet D2","metrics":{"Top 1 Accuracy":"95.87"},"uses_additional_data":false,"paper_date":"2024-09-23","paper":"/paper/decouplenet-a-lightweight-backbone-network","paper_url":"https://ieeexplore.ieee.org/document/10685518","paper_title":"DecoupleNet: A Lightweight Backbone Network With Efficient Feature Decoupling for Remote Sensing Visual Tasks","code":"https://github.com/lwCVer/DecoupleNet","n_code_links":1,"syntology":null},{"rank_in_archive_order":4,"model":"LWGANet L1","metrics":{"Top 1 Accuracy":"95.70"},"uses_additional_data":false,"paper_date":"2025-01-17","paper":"/paper/lwganet-a-lightweight-group-attention","paper_url":"https://arxiv.org/abs/2501.10040v1","paper_title":"LWGANet: A Lightweight Group Attention Backbone for Remote Sensing Visual Tasks","code":"https://github.com/lwcver/lwganet","n_code_links":1,"syntology":null},{"rank_in_archive_order":5,"model":"SEER (RegNet10B)","metrics":{"Top 1 Accuracy":"95.61"},"uses_additional_data":true,"paper_date":"2022-02-16","paper":"/paper/vision-models-are-more-robust-and-fair-when","paper_url":"https://arxiv.org/abs/2202.08360v2","paper_title":"Vision Models Are More Robust And Fair When Pretrained On Uncurated Images Without Supervision","code":"https://github.com/facebookresearch/vissl","n_code_links":1,"syntology":null},{"rank_in_archive_order":6,"model":"LWGANet L0","metrics":{"Top 1 Accuracy":"95.49"},"uses_additional_data":false,"paper_date":"2025-01-17","paper":"/paper/lwganet-a-lightweight-group-attention","paper_url":"https://arxiv.org/abs/2501.10040v1","paper_title":"LWGANet: A Lightweight Group Attention Backbone for Remote Sensing Visual Tasks","code":"https://github.com/lwcver/lwganet","n_code_links":1,"syntology":null},{"rank_in_archive_order":7,"model":"AGOS","metrics":{"Top 1 Accuracy":"94.91"},"uses_additional_data":false,"paper_date":"2022-05-06","paper":"/paper/all-grains-one-scheme-agos-learning-multi","paper_url":"https://arxiv.org/abs/2205.03371v1","paper_title":"All Grains, One Scheme (AGOS): Learning Multi-grain Instance Representation for Aerial Scene Classification","code":"https://github.com/biqiwhu/agos","n_code_links":1,"syntology":null},{"rank_in_archive_order":8,"model":"SwAV (ResNet50-w5)","metrics":{"Top 1 Accuracy":"94.73"},"uses_additional_data":true,"paper_date":"2022-02-16","paper":"/paper/vision-models-are-more-robust-and-fair-when","paper_url":"https://arxiv.org/abs/2202.08360v2","paper_title":"Vision Models Are More Robust And Fair When Pretrained On Uncurated Images Without Supervision","code":"https://github.com/facebookresearch/vissl","n_code_links":1,"syntology":null},{"rank_in_archive_order":9,"model":"DINO (DeiT-B/16)","metrics":{"Top 1 Accuracy":"93.97"},"uses_additional_data":true,"paper_date":"2022-02-16","paper":"/paper/vision-models-are-more-robust-and-fair-when","paper_url":"https://arxiv.org/abs/2202.08360v2","paper_title":"Vision Models Are More Robust And Fair When Pretrained On Uncurated Images Without Supervision","code":"https://github.com/facebookresearch/vissl","n_code_links":1,"syntology":null},{"rank_in_archive_order":10,"model":"LSENet","metrics":{"Top 1 Accuracy":"93.49"},"uses_additional_data":false,"paper_date":"2021-07-08","paper":"/paper/local-semantic-enhanced-convnet-for-aerial","paper_url":"https://drive.google.com/file/d/1c1dM43l24mchg8Pcy52mxRaeTg_kfYzY/view","paper_title":"Local semantic enhanced convnet for aerial scene recognition","code":"https://github.com/BiQiWHU/LSENet","n_code_links":1,"syntology":null},{"rank_in_archive_order":11,"model":"MoCo-v3 (ViT-B/16)","metrics":{"Top 1 Accuracy":"93.35"},"uses_additional_data":true,"paper_date":"2022-02-16","paper":"/paper/vision-models-are-more-robust-and-fair-when","paper_url":"https://arxiv.org/abs/2202.08360v2","paper_title":"Vision Models Are More Robust And Fair When Pretrained On Uncurated Images Without Supervision","code":"https://github.com/facebookresearch/vissl","n_code_links":1,"syntology":null},{"rank_in_archive_order":12,"model":"CLIP (ViT-B/16)","metrics":{"Top 1 Accuracy":"92.7"},"uses_additional_data":true,"paper_date":"2022-02-16","paper":"/paper/vision-models-are-more-robust-and-fair-when","paper_url":"https://arxiv.org/abs/2202.08360v2","paper_title":"Vision Models Are More Robust And Fair When Pretrained On Uncurated Images Without Supervision","code":"https://github.com/facebookresearch/vissl","n_code_links":1,"syntology":null},{"rank_in_archive_order":13,"model":"BYOL (ResNet200-w2)","metrics":{"Top 1 Accuracy":"92.53"},"uses_additional_data":true,"paper_date":null,"paper":null,"paper_url":null,"paper_title":"","code":null,"n_code_links":0,"syntology":null},{"rank_in_archive_order":14,"model":"DeiT-B/16","metrics":{"Top 1 Accuracy":"92.48"},"uses_additional_data":true,"paper_date":"2022-02-16","paper":"/paper/vision-models-are-more-robust-and-fair-when","paper_url":"https://arxiv.org/abs/2202.08360v2","paper_title":"Vision Models Are More Robust And Fair When Pretrained On Uncurated Images Without Supervision","code":"https://github.com/facebookresearch/vissl","n_code_links":1,"syntology":null},{"rank_in_archive_order":15,"model":"SimCLR-v2 (ResNet152-w3 + SK)","metrics":{"Top 1 Accuracy":"89.77"},"uses_additional_data":true,"paper_date":"2022-02-16","paper":"/paper/vision-models-are-more-robust-and-fair-when","paper_url":"https://arxiv.org/abs/2202.08360v2","paper_title":"Vision Models Are More Robust And Fair When Pretrained On Uncurated Images Without Supervision","code":"https://github.com/facebookresearch/vissl","n_code_links":1,"syntology":null},{"rank_in_archive_order":16,"model":"ResNet50 (ImageNet-supervised)","metrics":{"Top 1 Accuracy":"88.56"},"uses_additional_data":true,"paper_date":"2022-02-16","paper":"/paper/vision-models-are-more-robust-and-fair-when","paper_url":"https://arxiv.org/abs/2202.08360v2","paper_title":"Vision Models Are More Robust And Fair When Pretrained On Uncurated Images Without Supervision","code":"https://github.com/facebookresearch/vissl","n_code_links":1,"syntology":null},{"rank_in_archive_order":17,"model":"MIDC-Net","metrics":{"Top 1 Accuracy":"87.99"},"uses_additional_data":false,"paper_date":"2020-03-03","paper":"/paper/a-multiple-instance-densely-connected-convnet","paper_url":"https://drive.google.com/file/u/0/d/1a0q-lXSCCrCeoIG_0Cx4cS7ulsg5dLr3/view","paper_title":"A multiple-instance densely-connected ConvNet for aerial scene classification","code":"https://github.com/BiQiWHU/Attention-based-Multi-instance-CNN","n_code_links":1,"syntology":null},{"rank_in_archive_order":18,"model":"MoCo-v2 (ResNet50)","metrics":{"Top 1 Accuracy":"85.4"},"uses_additional_data":true,"paper_date":"2022-02-16","paper":"/paper/vision-models-are-more-robust-and-fair-when","paper_url":"https://arxiv.org/abs/2202.08360v2","paper_title":"Vision Models Are More Robust And Fair When Pretrained On Uncurated Images Without Supervision","code":"https://github.com/facebookresearch/vissl","n_code_links":1,"syntology":null},{"rank_in_archive_order":19,"model":"SAG-ViT","metrics":{"F1":"95.49"},"uses_additional_data":false,"paper_date":"2024-11-14","paper":"/paper/sag-vit-a-scale-aware-high-fidelity-patching","paper_url":"https://arxiv.org/abs/2411.09420v3","paper_title":"SAG-ViT: A Scale-Aware, High-Fidelity Patching Approach with Graph Attention for Vision Transformers","code":"https://github.com/shravan-18/SAG-ViT","n_code_links":1,"syntology":null},{"rank_in_archive_order":20,"model":"SkySense-O","metrics":{"zero-shot Acc":"83.28"},"uses_additional_data":false,"paper_date":"2023-12-15","paper":"/paper/skysense-a-multi-modal-remote-sensing","paper_url":"https://arxiv.org/abs/2312.10115v2","paper_title":"SkySense: A Multi-Modal Remote Sensing Foundation Model Towards Universal Interpretation for Earth Observation Imagery","code":"https://github.com/jack-bo1220/awesome-remote-sensing-foundation-models","n_code_links":1,"syntology":null}],"since_archive":{"present":false,"note":"No Syntology-extracted rows are published in this build."},"syntology":{"read_at":"2026-09-24T18:15:14+00:00","claim":"Per row: N of M harvested code samples from that row's paper executed on a synthesized fixture; the other M-N are unverified. Not a reproduction of the row's number; not a correctness claim. n_pointer_only_licence counts samples the site points at rather than redistributes (a licence axis, independent of ran/unverified).","rows_with_graph_line":0,"rows_with_any_sample_ran":0,"distinct_papers_with_graph_line":0,"distinct_papers_with_any_sample_ran":0,"samples_over_distinct_papers":{"n_ran":0,"n_unverified":0,"n_samples":0,"n_pointer_only_licence":0,"note":"each paper (arXiv id) counted once, however many rows it is behind; this is the page-level figure"},"samples_row_weighted":{"n_ran":0,"n_unverified":0,"n_samples":0,"n_pointer_only_licence":0,"note":"row-weighted: a paper behind several rows is counted once per row; inflated relative to samples_over_distinct_papers by design, kept for readers summing the per-row syntology blocks"}}}