Hardware-Aware Deep Reinforcement Learning for Anticipatory Sleep Scheduling in Multi-Hop IoT Networks: An FPGA-based RTL Implementation
DOI:
https://doi.org/10.31838/jvcs/08.01.16Keywords:
Deep Reinforcement Learning (DRL); Internet of Things (IoT); Sleep Scheduling; Energy-efficient Communication; Multi-Hop Wireless Networks; Hardware–Software Co-Design; FPGA-based Implementation; VLSI-enabled IoT SystemsAbstract
In low-traffic and event-driven networks, energy efficiency in multi-hop Internet of Things (IoT) networks remains a significant challenge. Conventional duty-cycling and medium access control protocols rely on periodic wake-up or reactive activation mechanisms, causing nodes to remain partially active even when no communication is required. This behavior leads to excessive energy consumption and a reduced network lifetime. In this paper, we present a hardware-aware anticipatory sleep scheduling framework based on Deep Reinforcement Learning (DRL) that enables each IoT node to learn its communication patterns and optimize its sleep–wake schedule in advance. A lightweight DRL agent estimates the likelihood of future communication events based on locally observed traffic and channel dynamics, enabling each node to independently select an appropriate operating mode (Sleep, Listen, or Active) before communication is required. The proposed framework employs energy-efficient and communication-reliable predictive node activity control, distinguishing it from traditional reactive approaches. For practical implementation, a hardware–software co-design architecture is developed for deployment on embedded and FPGA platforms. The DRL model is trained offline on high-performance computing platforms, and only the lightweight fixed-point inference is performed in resource-limited IoT devices. Additionally, the proposed hardware architecture is synthesized on an FPGA, analyzed to verify timing constraints, and simulated for power consumption at the Register Transfer Level (RTL) using VHDL. The results of the simulations performed in a multi-hop wireless network environment show a considerable performance improvement that includes the ability to extend the network lifetime by up to 30%, to reduce the energy consumption per successful event by 35–45%, and to reduce the control overhead by about 20–25% without
affecting communication latency and reliability. Furthermore, the implementation results on FPGAs show that the proposed scheduler is well-suited to resource-constrained embedded IoT platforms due to its low area usage, high operating frequency and low power consumption. The results obtained in this research show that the designed learning-enabled node activity control with hardware-aware design enables a scalable, self-organizing, and energy-efficient communication environment for the next generation of intelligent IoT systems and VLSI-enabled embedded systems.



