Many embedded applications have strict energy, memory, and time constraints, making neural network (NN) inference particularly challenging. Recently, a novel NN architecture, called Fast Feedforward Networks (FFFs), has been proposed to achieve inference with extremely lightweight computational demands and minimal latency. Yet, compared to feedforward networks with similar sizes, FFFs still lag behind in terms of performance, indicating that they do not utilize all of their parameters effectively. In this article, we explore a possible reason for this performance gap: the uncertainty in how samples are assigned to the network’s leaves. We attempt to overcome this challenge by making FFFs’ training inference-aware, hence introducing Inference-Aware Fast Feedforward Networks (IAFFFs). We imitate FFFs’ inference during training by using a step activation function alongside the traditional sigmoid activation function. We test different aware scheduling methods, which we dub “awareness scheduler”, to adjust the balance between the two activation functions during training, and examine how different schedules impact the model’s performance. Additionally, we employ leaf-weight virtualization with inference-aware retraining to compress our models so they can fit onto edge devices. We further employ an iterative compression approach to find an optimal awareness scheduler for compression to minimize performance drop due to compression. We experiment with different model sizes on various microcontrollers (MCUs) with different memory constraints to observe the latency and energy consumption introduced by the compression algorithm.
INSTANT: Inference-Aware Fast Feedforward Networks / Kilic, R.B., Yildirim, K.S., Iacca, G.. - In: ACM TRANSACTIONS ON EMBEDDED COMPUTING SYSTEMS. - ISSN 1539-9087. - 2026:62(2026). [10.1145/3815117]
INSTANT: Inference-Aware Fast Feedforward Networks
Renan Beran Kilic;Kasim Sinan Yildirim
;Giovanni Iacca
2026-01-01
Abstract
Many embedded applications have strict energy, memory, and time constraints, making neural network (NN) inference particularly challenging. Recently, a novel NN architecture, called Fast Feedforward Networks (FFFs), has been proposed to achieve inference with extremely lightweight computational demands and minimal latency. Yet, compared to feedforward networks with similar sizes, FFFs still lag behind in terms of performance, indicating that they do not utilize all of their parameters effectively. In this article, we explore a possible reason for this performance gap: the uncertainty in how samples are assigned to the network’s leaves. We attempt to overcome this challenge by making FFFs’ training inference-aware, hence introducing Inference-Aware Fast Feedforward Networks (IAFFFs). We imitate FFFs’ inference during training by using a step activation function alongside the traditional sigmoid activation function. We test different aware scheduling methods, which we dub “awareness scheduler”, to adjust the balance between the two activation functions during training, and examine how different schedules impact the model’s performance. Additionally, we employ leaf-weight virtualization with inference-aware retraining to compress our models so they can fit onto edge devices. We further employ an iterative compression approach to find an optimal awareness scheduler for compression to minimize performance drop due to compression. We experiment with different model sizes on various microcontrollers (MCUs) with different memory constraints to observe the latency and energy consumption introduced by the compression algorithm.| File | Dimensione | Formato | |
|---|---|---|---|
|
3815117.pdf
accesso aperto
Tipologia:
Versione editoriale (Publisher’s layout)
Licenza:
Creative commons
Dimensione
776.55 kB
Formato
Adobe PDF
|
776.55 kB | Adobe PDF | Visualizza/Apri |
I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione



