Similar Accuracy, Different Explanations in YOLO-Based Palm Fruit Ripeness Detection
DOI:
https://doi.org/10.31436/iiumej.v27i3.4693Keywords:
Explainable AI, Class Activation Mapping, Object Detection, YOLO, Oil Palm FruitAbstract
Modern palm-fruit detectors can achieve similarly high benchmark accuracy, yet class activation mapping (CAM) methods are often applied without establishing whether their explanatory behavior transfers across detector configurations. This study is a controlled exploratory within-dataset comparison: model scale and target layer were selected on the same 416-image test partition used for final CAM evaluation, so the reported values are not independent estimates of explanation generalization. We evaluated 15 YOLOv8, YOLO11, and YOLO26 variants under a matched training protocol and assessed representative medium models across three random seeds. Backbone representations were then audited on the 416 test images using class-agnostic EigenCAM, backbone-adapted Grad-CAM++ (bGrad-CAM++), and backbone-adapted Score-CAM (bScore-CAM). Explanation behavior was quantified with CAM-box intersection over union (IoU) and Pointing Game for spatial localization and with Mean Absolute Confidence Drop (MACD) and Increase in Confidence as image-level confidence-retention diagnostics. Repeated-run mean mAP50–95 was narrowly separated at 0.903 ± 0.002, 0.899 ± 0.002, and 0.901 ± 0.003 for YOLOv8m, YOLO11m, and YOLO26m, respectively. In contrast, Pointing Game varied by 68.03 percentage points and MACD by 48.62 percentage points across model–explainer configurations. EigenCAM showed the strongest observed profile on YOLOv8m, whereas bScore-CAM achieved the highest localization on YOLO11m; no explainer dominated across all evaluated backbones. The ranking reversal was also present within the difficult underripe, ripe, and overripe subset, and the recorded layer search showed that layer 8 was the strongest compatible backbone target across the layer 6–10 window in all three detectors. Similar detection accuracy therefore did not imply transferable explanation behavior. CAM explanations should be validated jointly with the explanatory target, layer, and detector rather than treated as interchangeable post-hoc visualizations.
ABSTRAK: Pengesan buah sawit moden boleh mencapai ketepatan penanda aras tinggi dan hampir setara, namun kaedah pemetaan pengaktifan kelas (CAM) sering diguna tanpa memastikan sama ada tingkah laku penjelasannya boleh dipindah merentas konfigurasi pengesan. Kajian ini merupakan perbandingan penerokaan terkawal dalam set data: skala model dan lapisan sasaran dipilih menggunakan pecahan ujian 416 imej sama seperti penilaian CAM akhir, di mana dapatan yang dilaporkan bukan anggaran bebas bagi generalisasi penjelasan. Sebanyak 15 varian YOLOv8, YOLO11 dan YOLO26 dinilai menggunakan protokol latihan sepadan, manakala model sederhana yang mewakili setiap keluarga dinilai menggunakan tiga benih rawak. Perwakilan tulang belakang kemudiannya diaudit pada 416 imej ujian menggunakan EigenCAM yang bebas kelas, Grad-CAM++ teradaptasi tulang belakang (bGrad-CAM++), dan Score-CAM teradaptasi tulang belakang (bScore-CAM). Tingkah laku penjelasan dinilai menggunakan CAM-box IoU dan Pengarahan Perburuan bagi menempat spatial, serta Purata Mutlak Keyakinan Menurun (MACD) dan Keyakinan Bertambah sebagai diagnostik pengekalan keyakinan pada aras imej. Purata mAP50–95 bagi larian berulang masing-masing adalah 0.903 ± 0.002, 0.899 ± 0.002 dan 0.901 ± 0.003 untuk YOLOv8m, YOLO11m dan YOLO26m. Sebaliknya, Pengarahan Pemburuan berbeza sebanyak 68.03 mata peratusan dan MACD sebanyak 48.62 mata merentas konfigurasi model–penerang. EigenCAM menunjukkan profil pemerhatian terkuat pada YOLOv8m, manakala bScore-CAM mencapai penyetempatan tertinggi pada YOLO11m; tiada penerang yang mendominasi semua tulang belakang yang dinilai. Pembalikan kedudukan ini turut hadir dalam subset sukar seperti kurang masak, masak dan terlebih masak, manakala carian lapisan yang direkod menunjukkan bahawa lapisan 8 ialah sasaran tulang belakang serasi terkuat dalam tetingkap lapisan 6–10 bagi ketiga-tiga pengesan. Dapatan ini menunjukkan bahawa ketepatan pengesanan yang hampir setara tidak semestinya menghasilkan tingkah laku penjelasan yang boleh dipindahkan. Oleh itu, penjelasan CAM perlu disahkan secara bersama dengan sasaran penjelasan, lapisan dan pengesan.
Downloads
References
T. S. Gunawan, M. Kartiwi, and A. Jamali, "Drone and deep learning-based instrumentation for palm fruit detection and yield measurement: A Southeast Asian roadmap," IEEE Instrumentation & Measurement Magazine, pp. 1-17, 2026, doi: 10.1109/MIM.2026.11224873.
G. K. A. Parveez et al., "Oil palm economic performance in Malaysia and R&D progress in 2024," Journal of Oil Palm Research, vol. 37, no. 2, pp. 187-208, 2025, doi: https://doi.org/10.21894/jopr.2025.0031.
T. S. Gunawan, M. Kartiwi, H. Mansor, and N. M. Yusoff, "Palm fruit ripeness detection and classification using various yolov8 models," in 2023 IEEE 9th International Conference on Smart Instrumentation, Measurement and Applications (ICSIMA), 2023: IEEE, pp. 193-198.
S. J. Purba, W. F. Tandion, and E. Irwansyah, "Comparison of the latest version of deep learning YOLO model in automatically detecting the ripeness level of oil palm fruit in PTPN IV, North Sumatra, Indonesia," in 2024 International Conference on Information Technology and Computing (ICITCOM), 2024: IEEE, pp. 295-300.
J. Y. Goh, Y. Md Yunos, and M. S. Mohamed Ali, "Fresh fruit bunch ripeness classification methods: A review," Food and Bioprocess Technology, vol. 18, no. 1, pp. 183-206, 2025.
J. W. Lai, H. R. Ramli, L. I. Ismail, and W. Z. Wan Hasan, "Oil palm fresh fruit bunch ripeness detection methods: A systematic review," Agriculture, vol. 13, no. 1, p. 156, 2023.
J. Y. Goh, M. S. Mohamed Ali, Y. Md Yunos, U. U. Sheikh, and M. S. Khan, "Outdoor RGB and Point Cloud Depth Dataset for Palm Oil Fresh Fruit Bunch Ripeness Classification and Localization," Scientific Data, vol. 12, no. 1, p. 687, 2025.
P. Chotikawanid, P. Saeleung, Y. Pianroj, S. Jumrat, T. Punvichai, and J. Muangprathub, "Optimizing an Object Detection Algorithm for Detecting Oil Palm Fruit Bunches and Their Ripeness," Applied Computational Intelligence and Soft Computing, vol. 2025, no. 1, p. 6263757, 2025.
J. Josdaan, V. C. Tamsil, J. Harefa, and K. Jingga, "Revolutionizing palm oil ripeness classification: Utilizing YOLOv8 for ultra-precise ripeness detection," Procedia Computer Science, vol. 245, pp. 700-709, 2024.
L. T. Ramos and A. D. Sappa, "A comprehensive analysis of YOLO architectures for tomato leaf disease identification," Scientific Reports, vol. 15, no. 1, p. 26890, 2025.
J. Terven, D.-M. Córdova-Esparza, and J.-A. Romero-González, "A comprehensive review of YOLO architectures in computer vision: From YOLOv1 to YOLOv8 and YOLO-NAS," Machine Learning and Knowledge Extraction, vol. 5, no. 4, pp. 1680-1716, 2023.
Y. Zhao et al., "DETRs beat YOLOs on real-time object detection," in 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2024: IEEE, pp. 16965-16974.
M. Tornero-Soria, A.-J. Sánchez-Salmerón, and E. Vendrell Vidal, "Toward a Deeper Understanding of YOLO26: Block-Level Architectural Analysis and Ablation Studies," Applied Sciences, vol. 16, no. 13, p. 6758, 2026.
A. Chattopadhay, A. Sarkar, P. Howlader, and V. N. Balasubramanian, "Grad-cam++: Generalized gradient-based visual explanations for deep convolutional networks," in 2018 IEEE winter conference on applications of computer vision (WACV), 2018: IEEE, pp. 839-847.
R. R. Selvaraju, M. Cogswell, A. Das, R. Vedantam, D. Parikh, and D. Batra, "Grad-cam: Visual explanations from deep networks via gradient-based localization," in Proceedings of the IEEE International Conference on Computer Vision, 2017, pp. 618-626.
M. Bany Muhammad and M. Yeasin, "Eigen-CAM: Visual explanations for deep convolutional neural networks," SN Computer Science, vol. 2, no. 1, p. 47, 2021.
H. Wang et al., "Score-CAM: Score-weighted visual explanations for convolutional neural networks," in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops, 2020, pp. 24-25.
V. Petsiuk et al., "Black-box explanation of object detectors via saliency maps," in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2021, pp. 11443-11452.
J. Adebayo, J. Gilmer, M. Muelly, I. Goodfellow, M. Hardt, and B. Kim, "Sanity checks for saliency maps," Advances in Neural Information Processing Systems, vol. 31, 2018.
A. Andres, A. Martinez-Seras, I. Laña, and J. Del Ser, "On the black-box explainability of object detection models for safe and trustworthy industrial applications," Results in Engineering, vol. 24, p. 103498, 2024.
D. G. Pai, M. Balachandra, and R. Kamath, "Explainable AI in agriculture: Review of applications, methodologies, and future directions," Engineering Research Express, vol. 7, no. 3, p. 032202, 2025.
Suharjito et al., "Annotated datasets of oil palm fruit bunch piles for ripeness grading using deep learning," Scientific data, vol. 10, no. 1, p. 72, 2023.
P.-T. Jiang, C.-B. Zhang, Q. Hou, M.-M. Cheng, and Y. Wei, "LayerCAM: Exploring hierarchical class activation maps for localization," IEEE Transactions on Image Processing, vol. 30, pp. 5875-5888, 2021.
J. Choe, S. J. Oh, S. Lee, S. Chun, Z. Akata, and H. Shim, "Evaluating weakly supervised object localization methods right," in 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2020: IEEE, pp. 3133-3142.
T.-Y. Lin et al., "Microsoft COCO: Common objects in context," in European Conference on Computer Vision, 2014: Springer, pp. 740-755.
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 IIUM Press

This work is licensed under a Creative Commons Attribution-NonCommercial 4.0 International License.








