Loading...

Analysis of Temporal-Spatial Patterns of Active Neuronal Circuits in Brain during Spoken and Imagined Speech with Semantically Associated Words

Nematbakhsh, Mohammad Jalal | 2025

251 Viewed
  1. Type of Document: M.Sc. Thesis
  2. Language: Farsi
  3. Document No: 58598 (19)
  4. University: Sharif University of Technology
  5. Department: Computer Engineering
  6. Advisor(s): Sameti, Hossein; Karbalaei Aghajan, Hamid
  7. Abstract:
  8. Brain-Computer Interfaces (BCIs) that aim to decode imagined speech from brain signals offer new hope for restoring communication to individuals with severe speech disorders. However, this field faces significant challenges, including the low signal-to-noise ratio in non-invasive Electroencephalography (EEG) signals and a severe lack of publicly available, linguistically rich datasets, especially for less-researched languages such as Persian. This research addresses these gaps by presenting two key contributions. First, the ``Persian Imagined Speech Dataset'' a new data corpus, was designed and collected. This dataset comprises EEG signals recorded from 20 participants imagining and articulating nine distinct classes. The vocabulary was selected to cover complex linguistic relationships such as antonym, synonymy, homophones, and homonym, providing a more challenging testbed than previous datasets. Second, an integrated deep learning framework, titled the ``Spatio-Spectral-Temporal Transformer'' is proposed for the dual tasks of classification and speech synthesis from EEG signals. This architecture utilizes a Convolutional Neural Network (CNN)-based backbone to extract local features from a time-frequency representation (obtained via Discrete Wavelet Transform) and a Transformer-based encoder to model long-range temporal dependencies. To enhance performance, the model is pre-trained using a self-supervised strategy based on masked signal modeling. Subsequently, for speech synthesis, the trained encoder is combined with a conditional diffusion model-based generative decoder to produce the audio signal. The results of comprehensive experiments show that the proposed model achieves high accuracy in subject-dependent classification tasks. Furthermore, quantitative evaluations in the speech synthesis task, using metrics such as Mel-Cepstral Distortion (MCD), demonstrate the significant superiority of the diffusion model-based approach over baseline methods using Generative Adversarial Networks (GANs). By providing a new dataset for the Persian language and a novel computational framework, this research takes a step toward the development of efficient and accessible speech BCI systems
  9. Keywords:
  10. Brain-Computer Interface (BCI) ; Electroencephalography ; Deep Learning ; Diffusion Model ; Speech Synthesis ; Imagined Speech ; Persian Dataset ; Transformer Network

 Digital Object List

 Bookmark

No TOC