In Bacteria and Archaea, mobile genetic elements (MGEs) are extremely diverse in terms of size, structure and mobilization mechanism, ranging from minimal non-autonomous sequences (<100 bp) to complex elements (>100,000 bp) capable of mobilizing many passenger genes. MGEs drive genome evolution and horizontal gene transfer (HGT), shaping the dissemination of functional traits across species, including antimicrobial resistance genes (ARGs) and virulence factors. Despite their importance, the distribution and transmission of MGEs in complex microbial communities remains understudied, largely due to the lack of methods for their systematic identification in metagenomic data and the limited scope of existing databases. In this thesis, I developed a comparative genomics framework to systematically identify insertions flanked by conserved regions in 1,345,857 prokaryotic genomes. This approach reconstructed 9,008,672 insertion clusters (ICs), capturing 61.4% of MGEs in public databases while revealing a vast diversity of previously uncharacterized elements. Integrating homology annotations, structural motif detection and machine learning predictions, I assigned 29.2% of ICs to known MGE classes, expanding their diversity by 32-fold. Functional analyses showed that ICs are enriched in canonical MGE-associated traits, such as ARGs, virulence factors, secondary metabolic pathways and prokaryotic defense systems. Mapping ICs across isolate genomes revealed extensive HGT across the prokaryotic tree of life, providing direct evidence of mobilization, and showed that 89.5% of horizontally transferred ICs were previously unknown. Together, these results represent the most comprehensive characterization of the prokaryotic mobilome to date and reveal a large diversity of uncharacterized elements, providing a resource for the exploration of MGEs at the metagenomic scale. In the second part of this thesis, I leverage large metagenomic data to explore the diversity of programmable nucleases for the development of new genome editing tools. CRISPR-Cas systems, which provide adaptive immunity against MGEs, have been widely repurposed for genome editing. However, clinical applications of currently available Cas nucleases remain limited by several factors, including activity, specificity, targeting requirements and efficient in vivo delivery. In particular, the widely used SpCas9 is not compatible with single adeno-associated viral (AAV) vector delivery, due 8 to its size, and is restricted to targets flanked by an NGG protospacer adjacent motif (PAM). To address these limitations, I developed a computational pipeline to identify and characterize TnpB proteins, compact programmable nucleases encoded by widespread MGE families (IS200/605 and IS607). By analyzing 330,895 TnpB orthologs from 14,127 species, I selected 25 candidates for experimental characterization. This led to the identification of ISPmu1, a TnpB from Pasteurella multocida that is active in human cells and represents a promising candidate for the development of compact genome editors. In parallel, I developed PAMpredict, a computational tool to accurately predict the PAM sequence of Cas9 nucleases. Applying this tool at scale revealed that natural PAM diversity across prokaryotes is sufficient to target almost all disease-causing mutations in the human genome with allele specificity. Overall, this thesis demonstrates that large metagenomic data enables the systematic exploration and characterization of MGEs and programmable nucleases across prokaryotes.
Systematic investigation of prokaryotic mobile genetic elements and programmable nucleases using large metagenomic data / Ciciani, M.. - (2026 Aug 27).
Systematic investigation of prokaryotic mobile genetic elements and programmable nucleases using large metagenomic data
Ciciani, Matteo
2026-08-27
Abstract
In Bacteria and Archaea, mobile genetic elements (MGEs) are extremely diverse in terms of size, structure and mobilization mechanism, ranging from minimal non-autonomous sequences (<100 bp) to complex elements (>100,000 bp) capable of mobilizing many passenger genes. MGEs drive genome evolution and horizontal gene transfer (HGT), shaping the dissemination of functional traits across species, including antimicrobial resistance genes (ARGs) and virulence factors. Despite their importance, the distribution and transmission of MGEs in complex microbial communities remains understudied, largely due to the lack of methods for their systematic identification in metagenomic data and the limited scope of existing databases. In this thesis, I developed a comparative genomics framework to systematically identify insertions flanked by conserved regions in 1,345,857 prokaryotic genomes. This approach reconstructed 9,008,672 insertion clusters (ICs), capturing 61.4% of MGEs in public databases while revealing a vast diversity of previously uncharacterized elements. Integrating homology annotations, structural motif detection and machine learning predictions, I assigned 29.2% of ICs to known MGE classes, expanding their diversity by 32-fold. Functional analyses showed that ICs are enriched in canonical MGE-associated traits, such as ARGs, virulence factors, secondary metabolic pathways and prokaryotic defense systems. Mapping ICs across isolate genomes revealed extensive HGT across the prokaryotic tree of life, providing direct evidence of mobilization, and showed that 89.5% of horizontally transferred ICs were previously unknown. Together, these results represent the most comprehensive characterization of the prokaryotic mobilome to date and reveal a large diversity of uncharacterized elements, providing a resource for the exploration of MGEs at the metagenomic scale. In the second part of this thesis, I leverage large metagenomic data to explore the diversity of programmable nucleases for the development of new genome editing tools. CRISPR-Cas systems, which provide adaptive immunity against MGEs, have been widely repurposed for genome editing. However, clinical applications of currently available Cas nucleases remain limited by several factors, including activity, specificity, targeting requirements and efficient in vivo delivery. In particular, the widely used SpCas9 is not compatible with single adeno-associated viral (AAV) vector delivery, due 8 to its size, and is restricted to targets flanked by an NGG protospacer adjacent motif (PAM). To address these limitations, I developed a computational pipeline to identify and characterize TnpB proteins, compact programmable nucleases encoded by widespread MGE families (IS200/605 and IS607). By analyzing 330,895 TnpB orthologs from 14,127 species, I selected 25 candidates for experimental characterization. This led to the identification of ISPmu1, a TnpB from Pasteurella multocida that is active in human cells and represents a promising candidate for the development of compact genome editors. In parallel, I developed PAMpredict, a computational tool to accurately predict the PAM sequence of Cas9 nucleases. Applying this tool at scale revealed that natural PAM diversity across prokaryotes is sufficient to target almost all disease-causing mutations in the human genome with allele specificity. Overall, this thesis demonstrates that large metagenomic data enables the systematic exploration and characterization of MGEs and programmable nucleases across prokaryotes.I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione



