HPC users often misconfigure Power-Management Knobs (PMKs) for their job executions, leading to energy wastage and performance loss. Existing approaches for configuring PMKs automatically present limitations which make them incompatible with production environments. In this work we present MACK, an end-to-end software framework which addresses such limitations and performs automatic PMKs configuration, prior to job execution. Our framework leverages online machine learning techniques to predict job performance characteristics, and automatically rewrites the submission script to configure the optimal PMKs configuration. We deploy MACK for the end-users of Supercomputer Fugaku, and we define specific policies to configure PMKs based on prediction of job performance characteristics. We test the end-to-end pipeline of MACK on over 150 test jobs running classic HPC applications, showing that our approach allows to save up to 6% of computation time, and 18% of energy consumption, per job.
Antici, F., Proia, A., Ohara, R., Hanawa, T., Kiziltan, Z., Bartolini, A., et al. (2026). Automated Configuration of Power-Management Knobs for Optimal HPC Job Executions. Institute of Electrical and Electronics Engineers Inc. [10.1109/ccgrid68966.2026.00030].
Automated Configuration of Power-Management Knobs for Optimal HPC Job Executions
Antici, FrancescoPrimo
;Proia, Andrea;Kiziltan, Zeynep;Bartolini, Andrea;
2026
Abstract
HPC users often misconfigure Power-Management Knobs (PMKs) for their job executions, leading to energy wastage and performance loss. Existing approaches for configuring PMKs automatically present limitations which make them incompatible with production environments. In this work we present MACK, an end-to-end software framework which addresses such limitations and performs automatic PMKs configuration, prior to job execution. Our framework leverages online machine learning techniques to predict job performance characteristics, and automatically rewrites the submission script to configure the optimal PMKs configuration. We deploy MACK for the end-users of Supercomputer Fugaku, and we define specific policies to configure PMKs based on prediction of job performance characteristics. We test the end-to-end pipeline of MACK on over 150 test jobs running classic HPC applications, showing that our approach allows to save up to 6% of computation time, and 18% of energy consumption, per job.I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.



