Static code analyzers, such as deep learning-based techniques, have proven to be essential in the security testing process. These methods rely on training an effective and reliable deep learning model to detect weaknesses and vulnerabilities. In the domain of source code analysis, the use of an adequate dataset of vulnerable and non-vulnerable code is crucial. However, it is recognized that these datasets are difficult to obtain, contain noise such as duplications, and are often strongly imbalanced. We aim to investigate the impact of these datasets on performance across various models and settings. Our methodology involves implementing two editing strategies, which include both model-level adjustments via a redefined loss function and enhancements through dataset resembling-based balancing approaches. We analyze various learning-based models, including two transformer-based models, SVDet and LineVul, two graph-based models, IVDetect and JLineVD, and two large language models, LLama and Mistral. The experimental analysis shows that resampling strategies combining over-sampling and under-sampling can improve model performance by up to and yield gains of up to in terms of MCC, a suited metric for imbalanced data evaluation. Appropriate resampling methodologies, loss-function refinements, and retrieval-augmented generation techniques improve the performance of graph-based models, sequence-based models, and large language models, respectively, for vulnerability detection. However, despite these improvements, comparative analysis across models reveals that resampling and loss-function adjustments affect model stability and are not always adequate to fully address the challenges posed by severely imbalanced datasets.

Beyond the Data: Architectural Choices for Imbalanced Learning in Vulnerability Detection / Lekeufack Foulefack, R.Z., Djifack, U.D., Marchetto, A.. - In: SN COMPUTER SCIENCE. - ISSN 2661-8907. - 7:585(2026), pp. 1-18. [10.1007/s42979-026-05164-5]

Beyond the Data: Architectural Choices for Imbalanced Learning in Vulnerability Detection

Rosmael Zidane Lekeufack Foulefack;Ulrich Duplex Djifack;Alessandro Marchetto
2026-01-01

Abstract

Static code analyzers, such as deep learning-based techniques, have proven to be essential in the security testing process. These methods rely on training an effective and reliable deep learning model to detect weaknesses and vulnerabilities. In the domain of source code analysis, the use of an adequate dataset of vulnerable and non-vulnerable code is crucial. However, it is recognized that these datasets are difficult to obtain, contain noise such as duplications, and are often strongly imbalanced. We aim to investigate the impact of these datasets on performance across various models and settings. Our methodology involves implementing two editing strategies, which include both model-level adjustments via a redefined loss function and enhancements through dataset resembling-based balancing approaches. We analyze various learning-based models, including two transformer-based models, SVDet and LineVul, two graph-based models, IVDetect and JLineVD, and two large language models, LLama and Mistral. The experimental analysis shows that resampling strategies combining over-sampling and under-sampling can improve model performance by up to and yield gains of up to in terms of MCC, a suited metric for imbalanced data evaluation. Appropriate resampling methodologies, loss-function refinements, and retrieval-augmented generation techniques improve the performance of graph-based models, sequence-based models, and large language models, respectively, for vulnerability detection. However, despite these improvements, comparative analysis across models reveals that resampling and loss-function adjustments affect model stability and are not always adequate to fully address the challenges posed by severely imbalanced datasets.
2026
585
Lekeufack Foulefack, Rosmael Zidane; Djifack, Ulrich Duplex; Marchetto, Alessandro
Beyond the Data: Architectural Choices for Imbalanced Learning in Vulnerability Detection / Lekeufack Foulefack, R.Z., Djifack, U.D., Marchetto, A.. - In: SN COMPUTER SCIENCE. - ISSN 2661-8907. - 7:585(2026), pp. 1-18. [10.1007/s42979-026-05164-5]
File in questo prodotto:
File Dimensione Formato  
s42979-026-05164-5.pdf

accesso aperto

Tipologia: Versione editoriale (Publisher’s layout)
Licenza: Creative commons
Dimensione 5.58 MB
Formato Adobe PDF
5.58 MB Adobe PDF Visualizza/Apri

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/11572/492473
Citazioni
  • ???jsp.display-item.citation.pmc??? ND
  • Scopus 0
  • ???jsp.display-item.citation.isi??? ND
  • OpenAlex 0
social impact