Papers › A Vision-Language Foundation Model for Leaf Disease Identification

A Vision-Language Foundation Model for Leaf Disease Identification

11 May 2025arXiv:2505.07019archive 2025-07-28

Khang Nguyen Quoc, Lan Le Thi Thu, Luyl-Da Quach

Leaf disease identification plays a pivotal role in smart agriculture. However, many existing studies still struggle to integrate image and textual modalities to compensate for each other's limitations. Furthermore, many of these approaches rely on pretraining with constrained datasets such as ImageNet, which lack domain-specific information. We propose SCOLD (Soft-target COntrastive learning for Leaf Disease identification), a context-aware vision-language foundation model tailored to address these challenges for agricultural tasks. SCOLD is developed using a diverse corpus of plant leaf images and corresponding symptom descriptions, comprising over 186,000 image-caption pairs aligned with 97 unique concepts. Through task-agnostic pretraining, SCOLD leverages contextual soft targets to mitigate overconfidence in contrastive learning by smoothing labels, thereby improving model generalization and robustness on fine-grained classification tasks. Experimental results demonstrate that SCOLD outperforms existing vision-language models such as OpenAI-CLIP-L, BioCLIP, and SigLIP2 across several benchmarks, including zero-shot and few-shot classification, image-text retrieval, and image classification, while maintaining a competitive parameter footprint. Ablation studies further highlight SCOLD's effectiveness in contrast to its counterparts. The proposed approach significantly advances the agricultural vision-language foundation model, offering strong performance with minimal or no supervised fine-tuning. This work lays a solid groundwork for future research on models trained with long-form and simplified contexts, tasks involving class ambiguity, and multi-modal systems for intelligent plant disease diagnostics. The code for this study is available at https://huggingface.co/enalis/scold

PaperPDFCode

Code

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Contrastive LearningImage ClassificationImage-text RetrievalText Retrievalimage-classification

Datasets

Introduced by this paper, per the archive.

LeafNet

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Image Classification LeafNet SCOLD Accuracy (Top-1) 95.49 #1 of 1 Archive leaderboard report
Image Classification PlantDoc SCOLD Accuracy 99.69 #1 of 2 Archive leaderboard report
Image Classification PlantDoc SCOLD PARAMS .171M #1 of 2 Archive leaderboard report
Image Classification PlantVillage SCOLD Accuracy 99.96 #2 of 2 Archive leaderboard report
Image Classification PlantVillage SCOLD F1 99.95 #2 of 2 Archive leaderboard report
Image Classification PlantVillage SCOLD Testing Ratio 10% #2 of 2 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

Contrastive Learning

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections