Adaptive optimization methods like Adam excel in handling sparse gradients and varying parameter scales, but risk instability in regions of low curvature due to excessive step sizes. We introduce Eva (Escaping Via Adaptive gradients), a family of novel optimizers comprising two variants—Eva+ and Eva-—that dynamically switch between Adam and scaled gradient steps based on local gradient and curvature properties. Our method implements component-wise detection to identify updates where Adam step magnitude exceeds the historical gradient norm. Eva+ employs gradient ascent to escape from degenerate regions, while Eva- provides conservative descent with magnitude damping. We set a threshold for detecting unstable updates, supported by theoretical analysis and empirical validation. Extensive experiments prove that both Eva variants achieve faster convergence and improved generalization compared to popular optimizers.
Eva Optimizer: Escaping Low-Curvature Traps in Deep Learning / Di Cecco, A., Metta, C., Papini, A., Fantozzi, M., Galfré, S.G., Veglió, M., Bianchi, L.A., Parton, M., Morandin, F.. - 16816:(2026), pp. 112-127. (28th International Conference on Pattern Recognition, ICPR 2026 fra 2026) [10.1007/978-3-032-31666-0_8].
Eva Optimizer: Escaping Low-Curvature Traps in Deep Learning
Bianchi, Luigi Amedeo;Parton, Maurizio;
2026-01-01
Abstract
Adaptive optimization methods like Adam excel in handling sparse gradients and varying parameter scales, but risk instability in regions of low curvature due to excessive step sizes. We introduce Eva (Escaping Via Adaptive gradients), a family of novel optimizers comprising two variants—Eva+ and Eva-—that dynamically switch between Adam and scaled gradient steps based on local gradient and curvature properties. Our method implements component-wise detection to identify updates where Adam step magnitude exceeds the historical gradient norm. Eva+ employs gradient ascent to escape from degenerate regions, while Eva- provides conservative descent with magnitude damping. We set a threshold for detecting unstable updates, supported by theoretical analysis and empirical validation. Extensive experiments prove that both Eva variants achieve faster convergence and improved generalization compared to popular optimizers.I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione



