Benchmarking deep vision architectures for fruit-tree sapling bark classification on SapBark-64
Remote Sensing Letters, cilt.17, sa.12, ss.1660-1674, 2026 (SCI-Expanded, Scopus)
- Yayın Türü: Makale / Tam Makale
- Cilt numarası: 17 Sayı: 12
- Basım Tarihi: 2026
- Doi Numarası: 10.1080/2150704x.2026.2720794
- Dergi Adı: Remote Sensing Letters
- Derginin Tarandığı İndeksler: Science Citation Index Expanded (SCI-EXPANDED), Scopus, Aerospace Database, Applied Science & Technology Source, BIOSIS, Compendex, Environment Index, Geobase, INSPEC, Academic Search Ultimate (EBSCO), Natural Science Collection (ProQuest), Earth, Atmospheric, & Aquatic Science Collection (ProQuest), Technology Collection (ProQuest)
- Sayfa Sayıları: ss.1660-1674
- Anahtar Kelimeler: bark image classification, convolutional neural networks, deep learning benchmark, explainable artificial intelligence, fine-grained visual recognition, Fruit-tree sapling classification, SapBark-64, vision transformers
- Karadeniz Teknik Üniversitesi Adresli: Evet
Özet
Accurate identification of fruit-tree saplings is important for nursery management and varietal authentication, but related species and cultivars can be visually similar at early growth stages. This study benchmarks seven ImageNet-pretrained deep vision architectures for bark-based classification on the public SapBark-64 dataset: ResNet18, ResNet50, DenseNet121, EfficientNet-B0, MobileNetV3-Large, ConvNeXt-Tiny and Swin-Tiny. All models were evaluated using a unified stratified 10-fold cross-validation protocol with identical preprocessing, training, model-selection and evaluation procedures. ConvNeXt-Tiny achieved the highest accuracy (0.9527 ± 0.0087) and macro-F1 (0.9499 ± 0.0084), significantly outperforming all other models after multiple-comparison correction. Swin-Tiny achieved 0.9188 ± 0.0136 accuracy and 0.9146 ± 0.0147 macro-F1. ImageNet-pretrained full fine-tuning performed best in the three-fold diagnostic analysis of DenseNet121 and EfficientNet-B0. ConvNeXt-Tiny was the most stable of the three models evaluated under synthetic perturbations, although blur and especially Gaussian noise caused substantial degradation. Grad-CAM provided qualitative evidence of attention to trunk and bark-texture regions. Because no external validation was performed and SapBark-64 lacks sapling and capture-session identifiers, possible identity-level leakage may make image-level estimates optimistic. Results should therefore be interpreted as a reproducible within-dataset benchmark rather than evidence of generalization across nurseries, devices or environments.