{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/hsi-bert-hyperspectral-image-classification","title":"HSI-BERT: Hyperspectral Image Classification Using the Bidirectional Encoder Representation From Transformers","arxiv_id":null,"date":"2019-09-04","proceeding":"IEEE Transactions on Geoscience and Remote Sensing 2019 9","authors":["Ji He","Lina Zhao","HongWei Yang","Mengmeng Zhang","Wei Li"],"abstract":"Deep learning methods have been widely used in hyperspectral image classification and have achieved state-of-the-art performance. Nonetheless, the existing deep learning methods are restricted by a limited receptive field, inflexibility, and difficult generalization problems in hyperspectral image classification. To solve these problems, we propose HSI-BERT, where BERT stands for bidirectional encoder representations from transformers and HSI stands for hyperspectral imagery. The proposed HSI-BERT has a global receptive field that captures the global dependence among pixels regardless of their spatial distance. HSI-BERT is very flexible and enables the flexible and dynamic input regions. Furthermore, HSI-BERT has good generalization ability because the jointly trained HSI-BERT can be generalized from regions with different shapes without retraining. HSI-BERT is primarily built on a multihead self-attention (MHSA) mechanism in an MHSA layer. Moreover, several attentions are learned by different heads, and each head of the MHSA layer encodes the semantic context-aware representation to obtain discriminative features. Because all head-encoded features are merged, the resulting features exhibit spatial-spectral information that is essential for accurate pixel-level classification. Quantitative and qualitative results demonstrate that HSI-BERT outperforms any other CNN-based model in terms of both classification accuracy and computational time and achieves state-of-the-art performance on three widely used hyperspectral image data sets.","url_abs":"https://doi.org/10.1109/TGRS.2019.2934760","url_pdf":"https://doi.org/10.1109/TGRS.2019.2934760","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"classification-1","task_name":"Classification"},{"task_slug":"few-shot-image-classification","task_name":"Few-Shot Image Classification"},{"task_slug":"hyperspectral-image-classification","task_name":"Hyperspectral Image Classification"},{"task_slug":"image-classification","task_name":"Image Classification"},{"task_slug":"image-classification","task_name":"image-classification"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/hyperspectral-image-classification-on-indian","task":"Hyperspectral Image Classification","dataset":"Indian Pines","model":"HSI-BERT","rank_in_archive_order":12,"of":34,"metrics":{"OA@15perclass":"58.50±1.56"},"uses_additional_data":false},{"leaderboard":"/sota/hyperspectral-image-classification-on-kennedy","task":"Hyperspectral Image Classification","dataset":"Kennedy Space Center","model":"HSI-BERT","rank_in_archive_order":9,"of":14,"metrics":{"OA@15perclass":"82.93±0.94"},"uses_additional_data":false},{"leaderboard":"/sota/hyperspectral-image-classification-on-pavia","task":"Hyperspectral Image Classification","dataset":"Pavia University","model":"HSI-BERT","rank_in_archive_order":11,"of":33,"metrics":{"OA@15perclass":"75.31±1.59"},"uses_additional_data":false}],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}