Loading...
Development of Chemometric Methods to Analyze Metabolomic Data Obtained from Imaging and Spectroscopic Technique for Untargeted Analysis of Cancer
Kashi, Maryam | 2025
69
Viewed
- Type of Document: Ph.D. Dissertation
- Language: Farsi
- Document No: 58706 (03)
- University: Sharif University of Technology
- Department: Chemistry
- Advisor(s): Parastar Shahri, Hadi
- Abstract:
- Metabolomics is a powerful approach for identifying breast cancer (BC) biomarkers, especially when combined with advanced spectroscopy/imaging and machine learning (ML) techniques. In this paper, proton nuclear magnetic resonance spectroscop (1H-NMR), visible-near infrared spectroscopy (Vis-NIR), and hyperspectral imaging (HSI) techniques are integrated with chemometrics and ML approaches for early detection of BC using serum, plasma, and dried plasma (DPS) samples. Various optimized extraction techniques, including liquid-liquid extraction, were used to enhance metabolite recovery, and all spectra were subjected to standard preprocessing before modeling. In the first study, kohonen self-organizing maps (SOM) and counter-propagation artificial neural networks (CP-ANN) were used to model 1H-NMR spectra from control groups and BC patients. Serum samples were prepared from 24 healthy individuals and 18 patients, and methanol and chloroform were used to extract metabolites under optimized conditions. 1H-NMR data were preprocessed for phase, baseline, and chemical shift correction. The preprocessed data were then analyzed with SOM and CP-ANN. The CP-ANN model distinguished the two classes with 100% accuracy in training set. The sensitivity was 96% for the control group and 100% for the patient group. In testing, CP-ANN achieved 90% sensitivity, 98% specificity, and 96% accuracy for the control group and 100% sensitivity, 90% specificity, and 96% accuracy for the patient group. Analysis of the resulting topology map revealed 14 important biomarkers (such as sarcosine, lysine, trehalose, tryptophan, betaine) that effectively distinguished healthy individuals from BC patients. The second study highlights metabolomics in early detection and monitoring of BC through plasma analysis. Using 1H-NMR and a supervised kohonen network (SKN), plasma samples from 72 BC patients and 75 healthy individuals were analyzed. Metabolites were extracted under optimal conditions using the response surface methodology (RSM) and spectra were preprocessed for phase, basis and shift corrections before modeling. The classification achieved 93% sensitivity and specificity with 96% accuracy for healthy individuals and 96% sensitivity, 96% specificity and 93% accuracy for BC patients. Topological map analysis identified key biomarkers (betaine, methionine, choline, histidine, tyrosine) that correctly distinguished patients from healthy individuals with 98% accuracy. In addition, SKN classified patients based on family history and highlighted metabolites such as aspartate, proline and leucine. Novel associations between ornithine and N-acetylglycoprotein were associated with BMI classification, a first in this field. This study reveals a set of metabolites associated with cancer occurrence and progression, providing insights into tumor biology and potential therapeutic targets. The third study also investigated portable Vis-NIR spectroscopy and hyperspectral imaging (HSI) in the 400–1000 nm range for the analysis of 143 dried plasma (DPS) samples (73 control and 70 BC samples) using ML for BC diagnosis. Plasma samples were dried and the variability between samples, drying methods, and contributing factors was investigated using ASCA. Vis-NIR spectroscopy and HSI imaging offer a safe, rapid, and cost-effective diagnostic approach suitable for screening. Given the complexity of HSI data, the multivariate curve-based least squares (MCR-ALS) algorithm was used as a feature extraction technique to obtain pure spatial and spectral profiles of the component present. Multivariate classification was performed on the spectroscopic and HSI data with artificial neural networks (ANN), k-nearest neighbor (kNN), random forest (RF), and support vector machines (SVM). The ANN achieved 86% accuracy in distinguishing healthy and diseased samples in the HSI data, while SVM modeling yielded 62% accuracy for the portable spectroscopic data. The results indicated changes in bilirubin, hemoglobin, porphyrins, proteins, and lipids
- Keywords:
- Chemometrics Method ; Machine Learning ; Breast Cancer ; Hyperspectral Imaging ; Nuclear Magnetic Resonance Spectroscopy ; Pattern Recognition ; Untargeted Metabolomics
-
محتواي کتاب
- view
