Optimizing Deep Learning Architectures for Resource-Constrained Edge Devices
Main Article Content
Abstract
Implementing large-scale transformer models on edge devices in precision agriculture causes high computation and memory cost. While traditional CNN has limited global contextual modeling capabilities, the large number of parameters required by a Vision Transformer makes it difficult to run on local agricultural hardware, without the use of clusters of graphics processing units (GPUs) in the cloud. This study aims at understanding how to optimize deep learning architectures for resource-constrained edge devices using sophisticated compression techniques in deep learning. In detail, the effectiveness of the Activation-aware Weight Quantization and structured pruning of Multi Head Self-Attention models is examined. These optimization techniques allow the use of a light-weight hybrid vision system, significantly reducing computing power while maintaining diagnostic performance for crop disease classification problems. The analytical results show that by using a 4-bit quantization scheme and a ½ structured pruning ratio, the memory usage and inference time in platforms like Raspberry Pi 5 and NVIDIA Jetson Orin Nano can be significantly reduced. Overall, the findings validate the feasibility of using heavy transformer models for real time, on-device crop disease detection in agriculture to enable offline and decentralized smart farming without relying on cloud connectivity at all times.