Deciphering Cell-Type-Specific Transcriptional Regulation in Tomato Leaves Through Ensemble Machine Learning and Single-Cell Transcriptomics
Abstract
High-throughput single-cell RNA sequencing (scRNA-seq) has substantially advanced plant transcriptional landscapes. However, decoding cell-type-specific transcriptional regulation in non-model crops like tomato (<i>Solanum lycopersicum</i>) remains challenging. An integrated computational pipeline was applied using high-dimensional weighted gene co-expression (hdWGCNA) and ensemble machine learning to analyze tomato leaf single-cell transcriptomes. Unsupervised clustering identified 19 cell subpopulations mapped to five major cell-types: mesophyll cells (50.6%), guard cells (31.0%), trichomes (8.3%), vascular cells (7.5%), and lamina epidermis (2.6%). hdWGCNA revealed eight cell-type-specific modules, linking mesophyll cells to photosynthesis and guard cells to redox homeostasis. Machine learning classifiers prioritized candidate transcription factors (TFs), with XGBoost achieving the highest accuracy (0.85) to define cell identity. A consensus of 33 core TFs was identified, from which four candidate TFs (<i>SlWRKY-78</i>, <i>SlWRKY-75</i>, <i>SlERF-57</i>, and <i>SlGLK-49</i>) were selected for in silico knockout (KO) analysis. The simulations predicted that these knockouts might dysregulate core functional pathways, such as serine-type endopeptidase inhibitor activity and protein binding. Furthermore, CellOracle simulations suggested that the virtual deletion of the guard-cell-associated <i>SlWRKY-78</i> and <i>SlWRKY-75</i> could induce a directional trajectory shift from the terminally differentiated guard cells back to the less differentiated mesophyll territory. These findings provide a promising computational framework for deciphering cell-type-specific regulatory programs in horticultural crops.