Compound identification is the main hurdle in LC-HRMS-based metabolomics, given the high number of ‘unknown’ metabolites. In recent years, numerous in silico fragmentation simulators have been developed to simplify and improve mass spectral interpretation and compound annotation. Nevertheless, expert mass spectrometry users and chemists are still needed to select the correct entry from the numerous candidates proposed by automatic tools, especially in the plant kingdom due to the huge structural diversity of natural compounds occurring in plants. In this work, we propose the use of a supervised machine learning approach to predict molecular substructures from isotopic patterns, training the model on a large database of grape metabolites. This approach, called ‘Compounds Characteristics Comparison’ (CCC) emulates the experience of a plant chemist who ‘gains experience’ from a (proof-of-principle) dataset of grape compounds. The results show that the CCC approach is able to predict with good accuracy most of the sub-structures proposed. In addition, after querying MS/MS spectra in Metfrag 2.2 and applying CCC predictions as scoring terms with real data, the CCC approach helped to give a better ranking to the correct candidates, improving users’ confidence in candidate selection. Our results demonstrated that the proposed approach can complement current identification strategies based on fragmentation simulators and formula calculators, assisting compound identification. The CCC algorithm is freely available as R package (https://github.com/lucanard/CCC) which includes a seamless integration with Metfrag. The CCC package also permits uploading additional training data, which can be used to extend the proposed approach to other systems biological matrices.

The Compound Characteristics Comparison (CCC) approach: a tool for improving confidence in natural compound identification / Narduzzi, Luca; Stanstrup, Jan; Mattivi, Fulvio; Franceschi, Pietro. - In: FOOD ADDITIVES & CONTAMINANTS. PART A. CHEMISTRY, ANALYSIS, CONTROL, EXPOSURE & RISK ASSESSMENT. - ISSN 1944-0049. - ELETTRONICO. - 35:11(2018), pp. 2145-2157. [10.1080/19440049.2018.1523572]

The Compound Characteristics Comparison (CCC) approach: a tool for improving confidence in natural compound identification

Narduzzi, Luca;Mattivi, Fulvio;Franceschi, Pietro
2018-01-01

Abstract

Compound identification is the main hurdle in LC-HRMS-based metabolomics, given the high number of ‘unknown’ metabolites. In recent years, numerous in silico fragmentation simulators have been developed to simplify and improve mass spectral interpretation and compound annotation. Nevertheless, expert mass spectrometry users and chemists are still needed to select the correct entry from the numerous candidates proposed by automatic tools, especially in the plant kingdom due to the huge structural diversity of natural compounds occurring in plants. In this work, we propose the use of a supervised machine learning approach to predict molecular substructures from isotopic patterns, training the model on a large database of grape metabolites. This approach, called ‘Compounds Characteristics Comparison’ (CCC) emulates the experience of a plant chemist who ‘gains experience’ from a (proof-of-principle) dataset of grape compounds. The results show that the CCC approach is able to predict with good accuracy most of the sub-structures proposed. In addition, after querying MS/MS spectra in Metfrag 2.2 and applying CCC predictions as scoring terms with real data, the CCC approach helped to give a better ranking to the correct candidates, improving users’ confidence in candidate selection. Our results demonstrated that the proposed approach can complement current identification strategies based on fragmentation simulators and formula calculators, assisting compound identification. The CCC algorithm is freely available as R package (https://github.com/lucanard/CCC) which includes a seamless integration with Metfrag. The CCC package also permits uploading additional training data, which can be used to extend the proposed approach to other systems biological matrices.
2018
11
Narduzzi, Luca; Stanstrup, Jan; Mattivi, Fulvio; Franceschi, Pietro
The Compound Characteristics Comparison (CCC) approach: a tool for improving confidence in natural compound identification / Narduzzi, Luca; Stanstrup, Jan; Mattivi, Fulvio; Franceschi, Pietro. - In: FOOD ADDITIVES & CONTAMINANTS. PART A. CHEMISTRY, ANALYSIS, CONTROL, EXPOSURE & RISK ASSESSMENT. - ISSN 1944-0049. - ELETTRONICO. - 35:11(2018), pp. 2145-2157. [10.1080/19440049.2018.1523572]
File in questo prodotto:
File Dimensione Formato  
Narduzzi_et_al_The Compound Characteristics Comparison CCC approach_Food Additives & Contaminants A_2018.pdf

Solo gestori archivio

Descrizione: Articolo principale
Tipologia: Versione editoriale (Publisher’s layout)
Licenza: Tutti i diritti riservati (All rights reserved)
Dimensione 1.8 MB
Formato Adobe PDF
1.8 MB Adobe PDF   Visualizza/Apri

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/11572/218809
Citazioni
  • ???jsp.display-item.citation.pmc??? 1
  • Scopus 4
  • ???jsp.display-item.citation.isi??? 4
social impact