Papers › SkySense: A Multi-Modal Remote Sensing Foundation Model Towards Universal...
SkySense: A Multi-Modal Remote Sensing Foundation Model Towards Universal Interpretation for Earth Observation Imagery
Xin Guo, Jiangwei Lao, Bo Dang, Yingying Zhang, Lei Yu, Lixiang Ru, Liheng Zhong, Ziyuan Huang, Kang Wu, Dingxiang Hu, Huimei He, Jian Wang, Jingdong Chen, Ming Yang, Yongjun Zhang, Yansheng Li
Prior studies on Remote Sensing Foundation Model (RSFM) reveal immense potential towards a generic model for Earth Observation. Nevertheless, these works primarily focus on a single modality without temporal and geo-context modeling, hampering their capabilities for diverse tasks. In this study, we present SkySense, a generic billion-scale model, pre-trained on a curated multi-modal Remote Sensing Imagery (RSI) dataset with 21.5 million temporal sequences. SkySense incorporates a factorized multi-modal spatiotemporal encoder taking temporal sequences of optical and Synthetic Aperture Radar (SAR) data as input. This encoder is pre-trained by our proposed Multi-Granularity Contrastive Learning to learn representations across different modal and spatial granularities. To further enhance the RSI representations by the geo-context clue, we introduce Geo-Context Prototype Learning to learn region-aware prototypes upon RSI's multi-modal spatiotemporal features. To our best knowledge, SkySense is the largest Multi-Modal RSFM to date, whose modules can be flexibly combined or used individually to accommodate various tasks. It demonstrates remarkable generalization capabilities on a thorough evaluation encompassing 16 datasets over 7 tasks, from single- to multi-modal, static to temporal, and classification to localization. SkySense surpasses 18 recent RSFMs in all test scenarios. Specifically, it outperforms the latest models such as GFM, SatLas and Scale-MAE by a large margin, i.e., 2.76%, 3.67% and 3.61% on average respectively. We will release the pre-trained weights to facilitate future research and Earth Observation applications.
In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.
Code
Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.
Code Syntology ran Syntology
Not run by Syntology. Nothing on this page verifies that the listed code works.
Tasks
Results from the paper archive 2025-07-28
| Task | Dataset | Model | Metric | Value | Rank at snapshot | Leaderboard | Report |
|---|---|---|---|---|---|---|---|
| Image Classification | RESISC45 | SkySense-O | zero-shot Acc | 83.28 | #20 of 20 | Archive leaderboard | report |
| Open Vocabulary Semantic Segmentation | FAST | SkySense-O | mIoU | 8.3 | #1 of 1 | Archive leaderboard | report |
| Open Vocabulary Semantic Segmentation | ISPRS Potsdam | SkySense-O | mIoU | 54.1 | #1 of 1 | Archive leaderboard | report |
| Open Vocabulary Semantic Segmentation | SIOR | SkySense-O | mIoU | 30.89 | #1 of 1 | Archive leaderboard | report |
| Open Vocabulary Semantic Segmentation | SOTA | SkySense-O | mIoU | 32.12 | #1 of 1 | Archive leaderboard | report |
| Open Vocabulary Semantic Segmentation | iSAID | SkySense-O | mIoU- | 43.9 | #1 of 2 | Archive leaderboard | report |
| Visual Question Answering | AID-VQA | SkySense-O | Acc. (test) | 94.10 | #1 of 1 | Archive leaderboard | report |
| Visual Question Answering | RSVQA-HR | SkySense-O | zero-shot Acc | 78.09 | #1 of 1 | Archive leaderboard | report |
| Visual Question Answering | SIRI-WHU | SkySense-O | Acc. (test) | 74.79 | #1 of 1 | Archive leaderboard | report |
Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.
Methods
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections