State-of-the-art ML models deliver unprecedented performance across computer vision, natural language processing, and multi-modal tasks. These improvements have been driven by significantly growing computational costs, memory footprint, and consequently, energy consumption. Several model compression techniques and hyperparameter optimization strategies have been proposed to address these increases. Most of these approaches, however, target the training phase and often overlook the deployment context and the model usage. We propose a general-purpose framework for ML model optimization, focused on deployment and inference, to assist practitioners by quantifying trade-offs between model accuracy, power consumption, and other deployment-relevant metrics such as latency and memory usage. The proposed system is validated in three use cases: deploying a computer vision model on a resource-constrained edge device, optimizing LLM serving on data center-grade hardware, and running a large-scale benchmark on the Frontier HPC system. The system successfully identifies optimal configurations, demonstrating its versatility and applicability to practical scenarios.
A Deployment-Aware Multi-Objective Hyper-Parameter Optimization Framework for Machine Learning / Zanotto, M., Kienzler, R., Padovani, G., Fiore, S.. - ELETTRONICO. - (2026).
A Deployment-Aware Multi-Objective Hyper-Parameter Optimization Framework for Machine Learning
Zanotto, Matteo
Primo
;Padovani, GabrielePenultimo
;Fiore, Sandro
Ultimo
2026-01-01
Abstract
State-of-the-art ML models deliver unprecedented performance across computer vision, natural language processing, and multi-modal tasks. These improvements have been driven by significantly growing computational costs, memory footprint, and consequently, energy consumption. Several model compression techniques and hyperparameter optimization strategies have been proposed to address these increases. Most of these approaches, however, target the training phase and often overlook the deployment context and the model usage. We propose a general-purpose framework for ML model optimization, focused on deployment and inference, to assist practitioners by quantifying trade-offs between model accuracy, power consumption, and other deployment-relevant metrics such as latency and memory usage. The proposed system is validated in three use cases: deploying a computer vision model on a resource-constrained edge device, optimizing LLM serving on data center-grade hardware, and running a large-scale benchmark on the Frontier HPC system. The system successfully identifies optimal configurations, demonstrating its versatility and applicability to practical scenarios.| File | Dimensione | Formato | |
|---|---|---|---|
|
e2e_model_opt OPEN ACCESS.pdf
accesso aperto
Tipologia:
Post-print referato (Refereed author’s manuscript)
Licenza:
Tutti i diritti riservati (All rights reserved)
Dimensione
1.87 MB
Formato
Adobe PDF
|
1.87 MB | Adobe PDF | Visualizza/Apri |
I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione



