Context: Handling imbalanced data poses a significant challenge in binary classification tasks across various domains. The low prevalence of rare events or minority classes in such datasets often results in biased models, negatively affecting their predictive performance and reliability. This issue becomes even more critical in extremely imbalanced databases, where the minority class represents <1 % of the total data. Objective: A research gap was identified in the literature concerning approaches for binary classification in extremely imbalanced datasets. Therefore, this Systematic Literature Review (SLR) aims to synthesize existing knowledge by providing insights into the applicability and effectiveness of different approaches—categorized into preprocessing techniques, classifiers, and ensemble methods—for addressing extreme class imbalance. Method: The development of this SLR followed a rigorous, phased protocol for gathering, reviewing, and synthesizing existing literature based on well-defined selection and quality criteria. Ultimately, 22 articles were selected, restricted to primary (experimental) studies focused exclusively on extremely imbalanced databases across various application domains. Results: The findings of this SLR highlight key directions based on summarized, quantitative data and the top-performing approaches identified across experiments, conducted on 52 databases. The use of combined approaches demonstrated superior, performance across multiple evaluation metrics. Particularly notable were, preprocessing techniques paired with ensemble methods—specifically, oversampling techniques combined with the Random Forest (RF) algorithm—which consistently achieved the best performance in extreme imbalance scenarios. Conclusion: In conclusion, the adoption of tailored and efficient approaches to address extreme imbalance in binary classification tasks, as presented in this SLR, can support the initial selection of methods based on database characteristics—thereby reducing time consumption, computational resources, and potential rework.
Efficient approaches for binary classification in extremely imbalanced databases: A systematic literature review / Duarte Pereira, L., Alves De Almeida, F., Melgani, F., Paulo Balestrassi, P.. - In: INFORMATION AND SOFTWARE TECHNOLOGY. - ISSN 0950-5849. - 187:(2025), pp. 107867-107867. [10.1016/j.infsof.2025.107867]
Efficient approaches for binary classification in extremely imbalanced databases: A systematic literature review
Farid Melgani;
2025-01-01
Abstract
Context: Handling imbalanced data poses a significant challenge in binary classification tasks across various domains. The low prevalence of rare events or minority classes in such datasets often results in biased models, negatively affecting their predictive performance and reliability. This issue becomes even more critical in extremely imbalanced databases, where the minority class represents <1 % of the total data. Objective: A research gap was identified in the literature concerning approaches for binary classification in extremely imbalanced datasets. Therefore, this Systematic Literature Review (SLR) aims to synthesize existing knowledge by providing insights into the applicability and effectiveness of different approaches—categorized into preprocessing techniques, classifiers, and ensemble methods—for addressing extreme class imbalance. Method: The development of this SLR followed a rigorous, phased protocol for gathering, reviewing, and synthesizing existing literature based on well-defined selection and quality criteria. Ultimately, 22 articles were selected, restricted to primary (experimental) studies focused exclusively on extremely imbalanced databases across various application domains. Results: The findings of this SLR highlight key directions based on summarized, quantitative data and the top-performing approaches identified across experiments, conducted on 52 databases. The use of combined approaches demonstrated superior, performance across multiple evaluation metrics. Particularly notable were, preprocessing techniques paired with ensemble methods—specifically, oversampling techniques combined with the Random Forest (RF) algorithm—which consistently achieved the best performance in extreme imbalance scenarios. Conclusion: In conclusion, the adoption of tailored and efficient approaches to address extreme imbalance in binary classification tasks, as presented in this SLR, can support the initial selection of methods based on database characteristics—thereby reducing time consumption, computational resources, and potential rework.| File | Dimensione | Formato | |
|---|---|---|---|
|
2025_INFSOF-Pedro.pdf
Solo gestori archivio
Tipologia:
Versione editoriale (Publisher’s layout)
Licenza:
Tutti i diritti riservati (All rights reserved)
Dimensione
4.63 MB
Formato
Adobe PDF
|
4.63 MB | Adobe PDF | Visualizza/Apri |
I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione



