Efficient Edge-Based Object Detection via Multi-Signal Teacher–Student Knowledge Distillation
DOI:
https://doi.org/10.31436/iiumej.v27i3.4344Keywords:
object detection, edge computing, anchor-assignment consistency, knowledge distillationAbstract
A key obstacle to bringing computer vision everywhere, from robots and drones to smart cities and wearable tech, is the difficulty of deploying real-time object detectors on resource-limited edge hardware such as the NVIDIA Jetson Nano, Raspberry Pi 5, and ARM-based mobile processors. State-of-the-art detectors such as YOLOv8-L achieve outstanding accuracy on server GPUs but often deliver only single-digit or low FPS on embedded hardware. This paper presents EdgeKD-Net, a multi-signal Knowledge Distillation (KD) framework that transfers the representational knowledge of a high-capacity teacher to a compact edge-deployable student. The student employs a MobileNetV3-Large backbone augmented with a quantization- and distillation-aware Dual-Path Feature Pyramid (DPFP). Three complementary distillation signals are introduced: (i) temperature-scaled response-level KL divergence (?=4); (ii) Channel-Affinity Feature Alignment (CAFA), transferring structural channel co-activation patterns via cosine-affinity matrices; and (iii) Detection Fidelity Loss (DFL), applied exclusively on the formally defined consistent anchor set, eliminating noisy gradients from anchor-assignment mismatch. Across three independent runs, EdgeKD-Net achieves a mAP of 38.6 ± 0.14 mAP@[0.50:0.95] on MS-COCO 2017, only 1.4 mAP below its YOLOv8-L teacher while operating at 61 FPS (FP16) and 89 FPS (INT8) on NVIDIA Jetson Nano with 8.77 mAP/W energy efficiency. The framework is also evaluated on the VisDrone 2023 dataset, which contains dense aerial small objects. It achieves 14.1 ± 0.22 mAP@[0.50:0.95], compared with 12.4 for the best distillation baseline. This improvement is larger than the 0.9-point gain on MS-COCO, showing that the proposed method is effective on both datasets. The evaluation is limited to these two datasets, and we do not claim broader generalization.
ABSTRAK: Halangan utama dalam memperluas penggunaan penglihatan komputer, daripada robot dan dron kepada bandar pintar serta teknologi boleh pakai, ialah kesukaran melaksanakan pengesan objek masa nyata pada perkakasan pengkomputeran pinggir yang mempunyai sumber terhad, seperti NVIDIA Jetson Nano, Raspberry Pi 5 dan pemproses mudah alih berasaskan ARM. Pengesan tercanggih seperti YOLOv8-L mencapai ketepatan yang sangat baik pada unit pemprosesan grafik (GPU) pelayan, tetapi lazimnya hanya menghasilkan kadar bingkai sesaat (FPS) satu digit atau rendah apabila dilaksanakan pada perkakasan terbenam. Makalah ini memperkenalkan EdgeKD-Net, iaitu kerangka Penyulingan Pengetahuan (*Knowledge Distillation*, KD) berbilang isyarat yang memindahkan pengetahuan representasi daripada model pengajar berkapasiti tinggi kepada model pelajar kompak yang sesuai dilaksanakan pada peranti pinggir. Model pelajar menggunakan tulang belakang MobileNetV3-Large yang dipertingkatkan dengan Piramid Ciri Dwialiran (*Dual-Path Feature Pyramid*, DPFP) yang peka terhadap pengkuantuman dan penyulingan. Tiga isyarat penyulingan yang saling melengkapi diperkenalkan: (i) pencapahan Kullback–Leibler (KL) pada aras respons dengan penskalaan suhu (? = 4); (ii) Penjajaran Ciri Afiniti Saluran (*Channel-Affinity Feature Alignment*, CAFA), yang memindahkan pola struktur pengaktifan bersama saluran melalui matriks afiniti kosinus; dan (iii) Kehilangan Kesetiaan Pengesanan (*Detection Fidelity Loss*, DFL), yang digunakan secara eksklusif pada set sauh tekal yang ditakrifkan secara formal bagi menghapuskan kecerunan hingar akibat ketidakpadanan penetapan sauh. Berdasarkan tiga larian bebas, EdgeKD-Net mencapai mAP sebanyak 38.6 ± 0.14 bagi mAP@[0.50:0.95] pada set data MS-COCO 2017, iaitu hanya 1.4 mata mAP lebih rendah daripada model pengajarnya, YOLOv8-L. Pada NVIDIA Jetson Nano, model ini beroperasi pada 61 FPS menggunakan FP16 dan 89 FPS menggunakan INT8, dengan kecekapan tenaga sebanyak 8.77 mAP/W. Kerangka ini turut dinilai menggunakan set data VisDrone 2023 yang mengandungi objek udara bersaiz kecil dan padat. EdgeKD-Net mencapai 14.1 ± 0.22 mAP@[0.50:0.95], berbanding 12.4 yang dicatatkan oleh kaedah asas penyulingan terbaik. Peningkatan ini lebih besar daripada peningkatan 0.9 mata pada MS-COCO, sekali gus menunjukkan bahawa kaedah yang dicadangkan berkesan pada kedua-dua set data. Walau bagaimanapun, penilaian kajian ini terhad kepada kedua-dua set data tersebut dan tiada dakwaan dibuat mengenai keupayaan generalisasi yang lebih luas.
Downloads
References
Jocher G, Chaurasia A, Qiu J (2023). Ultralytics YOLOv8. GitHub. Available: https://github.com/ultralytics/ultralytics
Lv Y, Xu Y, Zhao W, Wang G, Wei J, Cui Y. (2024) DETRs Beat YOLOs on Real-time Object Detection. Proc. IEEE Conf. Comput. Vis. Pattern Recognit. (CVPR); Seattle, USA, 16965–16974. doi: 10.1109/CVPR52733.2024.01605
Zhang H, Li F, Liu S, Zhang L, Su H, Zhu J, Ni LM, Shum H (2022). DINO: DETR with Improved DeNoising Anchor Boxes for End-to-End Object Detection. arXiv preprint arXiv:2203.03605. Available: https://arxiv.org/abs/2203.03605
Hinton G, Vinyals O, Dean J (2014) Distilling the Knowledge in a Neural Network. Proc. NeurIPS Workshop on Deep Learning; Montreal, Canada.
Romero A, Ballas N, Kahou SE, Chassang A, Gatta C, Bengio Y. (2015) FitNets: Hints for Thin Deep Nets. Proc. Int. Conf. Learn. Represent. (ICLR); San Diego, USA.
Zagoruyko S, Komodakis N. (2017) Paying More Attention to Attention: Improving the Performance of Convolutional Neural Networks via Attention Transfer. Proc. Int. Conf. Learn. Represent. (ICLR); Toulon, France.
Tian Y, Krishnan D, Isola P. (2020) Contrastive representation distillation. In Proc. Int. Conf. Learning Representations (ICLR); Addis Ababa, Ethiopia. https://doi.org/10.48550/arXiv.1910.10699
Park W, Kim D, Lu Y, Cho M. (2019) Relational knowledge distillation. In Proc. IEEE Conf. Computer Vision and Pattern Recognition (CVPR); Long Beach, USA. pp 3962-3971. https://doi.org/10.1109/CVPR.2019.00409
Passalis N, Tefas A. (2018) Learning deep representations with probabilistic knowledge transfer. In Proc. European Conf. Computer Vision (ECCV); Munich, Germany. Lecture Notes in Computer Science, vol 11215, pp 268-284. https://doi.org/10.1007/978-3-030-01252-6_17
Chen G, Choi W, Yu X, Han T, Chandraker M. (2017) Learning efficient object detection models with knowledge distillation. In Proc. 31st Int. Conf. Neural Information Processing Systems (NIPS); Long Beach, USA. pp 742-751.
Yang Z, Li Z, Jiang X, Gong Y, Yuan Z, Zhao D, Yuan C. (2022) Focal and global knowledge distillation for detectors. In Proc. IEEE Conf. Computer Vision and Pattern Recognition (CVPR); New Orleans, USA. pp 4643-4652. https://doi.org/10.1109/CVPR52688.2022.00460
Yang Z, Li Z, Shao M, Shi D, Yuan Z, Yuan C. (2022) Masked generative distillation. In Proc. European Conf. Computer Vision (ECCV); Tel Aviv, Israel. Lecture Notes in Computer Science, vol 13671, pp 53-69. https://doi.org/10.1007/978-3-031-20083-0_4
Guo J, Han K, Wang Y, Wu H, Chen X, Xu C, Xu C. (2021) Distilling object detectors via decoupled features. In Proc. IEEE Conf. Computer Vision and Pattern Recognition (CVPR); Nashville, USA. pp 2154-2164. https://doi.org/10.1109/CVPR46437.2021.00219
Howard A, Sandler M, Chu G, Chen LC, Chen B, Tan M, Wang W, Zhu Y, Pang R, Vasudevan V, Le QV, Adam H. (2019) Searching for MobileNetV3. In Proc. IEEE Int. Conf. Computer Vision (ICCV); Seoul, Korea. pp 1314-1324. https://doi.org/10.1109/ICCV.2019.00140
Lin TY, Dollár P, Girshick R, He K, Hariharan B, Belongie S. (2017) Feature pyramid networks for object detection. In Proc. IEEE Conf. Computer Vision and Pattern Recognition (CVPR); Honolulu, USA. pp 936-944. https://doi.org/10.1109/CVPR.2017.106
Liu S, Qi L, Qin H, Shi J, Jia J. (2018) Path aggregation network for instance segmentation. In Proc. IEEE Conf. Computer Vision and Pattern Recognition (CVPR); Salt Lake City, USA. https://doi.org/10.1109/CVPR.2018.00913
Tan M, Pang R, Le QV. (2020) EfficientDet: scalable and efficient object detection. In Proc. IEEE Conf. Computer Vision and Pattern Recognition (CVPR); Seattle, USA. pp 10778-10787. https://doi.org/10.1109/CVPR42600.2020.01079
Zhou X, Wang D, Krähenbühl P. (2019) Objects as points. arXiv:1904.07850. https://doi.org/10.48550/arXiv.1904.07850
Tian Z, Shen C, Chen H, He T. (2019) FCOS: fully convolutional one-stage object detection. In Proc. IEEE Int. Conf. Computer Vision (ICCV); Seoul, Korea. pp 9626-9635. https://doi.org/10.1109/ICCV.2019.00972
Gevorgyan Z. (2022) SIoU loss: more powerful learning for bounding box regression. arXiv:2205.12740. https://doi.org/10.48550/arXiv.2205.12740
Xu S, Wang L, Wen G (2023). LiteFPN: A Lightweight Feature Pyramid Network for Real-Time Object Detection on Edge Devices. IEEE Trans. Circuits Syst. Video Technol., 33(8), 4021–4033.
Han K, Wang Y, Tian Q, Guo J, Xu C, Xu C. (2020) GhostNet: more features from cheap operations. In Proc. IEEE Conf. Computer Vision and Pattern Recognition (CVPR); Seattle, USA. pp 1577-1586. https://doi.org/10.1109/CVPR42600.2020.00165
Dong Y, Gao P, Liu C, Zhao X. (2023) SlimFPN: Slimming Feature Pyramid Networks via Structured Pruning for Efficient Object Detection. Proc. Int. Joint Conf. Neural Networks (IJCNN); Gold Coast, Australia.
Zheng G, Ding X, Huang X, Feng J, Wei Y, Zhang Y, Liu J, Yuille A, Kong T. (2022) Localisation Distillation for Dense Object Detection. Proc. IEEE Conf. Comput. Vis. Pattern Recognit. (CVPR); New Orleans, USA, 9407–9416.
Zhu Y, Zhou Q, Liu N, Xu Z, Ou Z, Mou X, Tang J. (2023) ScaleKD: Distilling Scale-Aware Knowledge in Small Object Detector. Proc. IEEE Conf. Comput. Vis. Pattern Recognit. (CVPR); Vancouver, Canada, 19723–19733.
Jang Y, Shin W, Kim J, Woo S, Bae SH. (2022) GLAMD: global and local attention mask distillation for object detectors. In Proc. European Conf. Computer Vision (ECCV); Tel Aviv, Israel. Lecture Notes in Computer Science, vol 13670, pp 460-476. https://doi.org/10.1007/978-3-031-20080-9_27
Li Y, Li X, Li W, Hou Q, Liu L, Cheng MM, Yang J. (2024) SARDet-100K: towards open-source benchmark and toolkit for large-scale SAR object detection. arXiv:2403.06534. https://doi.org/10.48550/arXiv.2403.06534
Cai H, Li J, Hu M, Gan C, Han S. (2023) EfficientViT: lightweight multi-scale attention for high-resolution dense prediction. In Proc. IEEE Int. Conf. Computer Vision (ICCV); Paris, France. pp 17256-17267. https://doi.org/10.1109/ICCV51070.2023.01587
Mehta S, Rastegari M. (2022) Separable self-attention for mobile vision transformers. arXiv:2206.02680. https://doi.org/10.48550/arXiv.2206.02680.
Wang CY, Yeh IH, Liao HYM. (2024) YOLOv9: learning what you want to learn using programmable gradient information. In Proc. European Conf. Computer Vision (ECCV); Milan, Italy. Lecture Notes in Computer Science, vol 15089. https://doi.org/10.1007/978-3-031-72751-1_1
Lyu C, Zhang W, Huang H, Zhou Y, Wang Y, Liu Y, Zhang S, Chen K. (2022) RTMDet: an empirical study of designing real-time object detectors. arXiv:2212.07784. https://doi.org/10.48550/arXiv.2212.07784.
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 IIUM Press

This work is licensed under a Creative Commons Attribution-NonCommercial 4.0 International License.








