We develop a mixture model for non-negative, heavy-tailed data, such as losses in actuarial and risk management applications. The mixture has a lognormal component, which is usually appropriate for the body of the distribution, and a Pareto-type tail, aimed at accommodating the largest observations, since the lognormal often decays too fast. Given that the tail is modeled by a zero-location Generalized Pareto distribution, the model is fully unsupervised, i.e. no threshold needs to be chosen. We show that maximum likelihood estimation can be performed by means of the EM algorithm and that the model is quite flexible in fitting data from different data-generating processes. Simulation experiments and a real-data application to automobiles claims suggest that the approach is equivalent in terms of goodness-of-fit, but easier to estimate, with respect to two existing distributions with similar features. All the methods are implemented in the R package lognGPD, available on CRAN
Modeling and estimating skewed and heavy-tailed populations via unsupervised mixture models / Bee, M., Santi, F.. - In: ADVANCES IN DATA ANALYSIS AND CLASSIFICATION. - ISSN 1862-5347. - 2026:(2026). [10.1007/s11634-026-00703-7]
Modeling and estimating skewed and heavy-tailed populations via unsupervised mixture models
Bee, Marco
Primo
;Santi, FlavioUltimo
2026-01-01
Abstract
We develop a mixture model for non-negative, heavy-tailed data, such as losses in actuarial and risk management applications. The mixture has a lognormal component, which is usually appropriate for the body of the distribution, and a Pareto-type tail, aimed at accommodating the largest observations, since the lognormal often decays too fast. Given that the tail is modeled by a zero-location Generalized Pareto distribution, the model is fully unsupervised, i.e. no threshold needs to be chosen. We show that maximum likelihood estimation can be performed by means of the EM algorithm and that the model is quite flexible in fitting data from different data-generating processes. Simulation experiments and a real-data application to automobiles claims suggest that the approach is equivalent in terms of goodness-of-fit, but easier to estimate, with respect to two existing distributions with similar features. All the methods are implemented in the R package lognGPD, available on CRAN| File | Dimensione | Formato | |
|---|---|---|---|
|
BeeSanti2026-Advances_in_Data_Analysis_and_Classification.pdf
accesso aperto
Descrizione: PDF online-first
Tipologia:
Versione editoriale (Publisher’s layout)
Licenza:
Creative commons
Dimensione
3.14 MB
Formato
Adobe PDF
|
3.14 MB | Adobe PDF | Visualizza/Apri |
I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione



