Loading...
Search for: sameti--h
0.117 seconds

    Speech enhancement based on hidden markov model with discrete cosine transform coefficients using laplace and gaussian distributions

    , Article 2012 11th International Conference on Information Science, Signal Processing and their Applications, ISSPA 2012, 2 July 2012 through 5 July 2012 ; July , 2012 , Pages 304-309 ; 9781467303828 (ISBN) Aroudi, A ; Veisi, H ; Sameti, H ; Sharif University of Technology
    2012
    Abstract
    This paper presents a novel HMM-based speech enhancement framework based on Laplace and Gaussian distributions in DCT domain. We propose analytical procedures for training clean speech and noise models with the aim of Baum's auxiliary function and present two MMSE estimators based on Gaussian-Gaussian (for clean speech and noise respectively) and Laplace-Gaussian combinations in the HMM framework. The performance evaluation is done using SNR and PESQ measures and the results of the proposed techniques are compared with AR-HMM approach. Higher SNR improvement is achieved for the proposed method in the Gaussian-Gaussian case in comparison with AR-HMM and Laplace-Gaussian techniques for both... 

    Hidden markov model-based speech enhancement using multivariate laplace and gaussian distributions

    , Article IET Signal Processing ; Volume 9, Issue 2 , 2015 , Pages 177-185 ; 17519675 (ISSN) Aroudi, A ; Veisi, H ; Sameti, H ; Sharif University of Technology
    Institution of Engineering and Technology  2015
    Abstract
    In this paper, statistical speech enhancement using hidden Markov model (HMM) is studied and new techniques for applying non-Gaussian distributions are proposed. The superiority of using non-Gaussian distributions in online adaptive noise suppression algorithms has been proven; however, in this study, this approach is formulated in an HMM-based mean-square error estimator (MMSE) estimator in which a priori models are trained in an off-line manner. In addition, an analytical study of using different distributions other than autoregressive (AR) Gaussian distribution, such as Laplace, is presented in order to construct an accurate HMM as a priori model for discrete Fourier transform and... 

    HalluSafe at SemEval-2024 task 6: an NLI-based approach to make LLMs safer by better detecting hallucinations and overgeneration mistakes

    , Article SemEval 2024 - 18th International Workshop on Semantic Evaluation, Proceedings of the Workshop ; 2024 , Pages 139-147 ; 979-889176107-0 (ISBN) Rahimi, Z ; Amirzadeh, H ; Sohrabi, A ; Taghavi, Z ; Sameti, H
    ACL Anthology  2024
    Abstract
    The advancement of large language models (LLMs), their ability to produce eloquent and fluent content, and their vast knowledge have resulted in their usage in various tasks and applications. Despite generating fluent content, this content can contain fabricated or false information. This problem is known as hallucination and has reduced the confidence in the output of LLMs. In this work, we have used Natural Language Inference to train classifiers for hallucination detection to tackle SemEval-2024 Task 6-SHROOM (Mickus et al., 2024) which is defined in three sub-tasks: Paraphrase Generation, Machine Translation, and Definition Modeling. We have also conducted experiments on LLMs to evaluate... 

    PerKey: a persian news corpus for keyphrase extraction and generation

    , Article 9th International Symposium on Telecommunication, IST 2018, 17 December 2018 through 19 December 2018 ; 2019 , Pages 460-465 ; 9781538682746 (ISBN) Doostmohammadi, E ; Bokaei, M. H ; Sameti, H ; Sharif University of Technology
    Institute of Electrical and Electronics Engineers Inc  2019
    Abstract
    Keyphrases provide an extremely dense summary of a text. Such information can be used in many Natural Language Processing tasks, such as information retrieval and text summarization. Since previous studies on Persian keyword or keyphrase extraction have not published their data, the field suffers from the lack of a human extracted keyphrase dataset. In this paper, we introduce PerKey1, a corpus of 553k news articles from six Persian news websites and agencies with relatively high quality author extracted keyphrases, which is then filtered and cleaned to achieve higher quality keyphrases. The resulted data was put into human assessment to ensure the quality of the keyphrases. We also measured... 

    Persian keyphrase generation using sequence-to-sequence models

    , Article 27th Iranian Conference on Electrical Engineering, ICEE 2019, 30 April 2019 through 2 May 2019 ; 2019 , Pages 2010-2015 ; 9781728115085 (ISBN) Doostmohammadi, E ; Bokaei, M. H ; Sameti, H ; Sharif University of Technology
    Institute of Electrical and Electronics Engineers Inc  2019
    Abstract
    Keyphrases are a very short summary of an input text and provide the main subjects discussed in the text. Keyphrase extraction is a useful upstream task and can be used in various natural language processing problems, for example, text summarization and information retrieval, to name a few. However, not all the keyphrases are explicitly mentioned in the body of the text. In real-world examples there are always some topics that are discussed implicitly. Extracting such keyphrases requires a generative approach, which is adopted here. In this paper, we try to tackle the problem of keyphrase generation and extraction from news articles using deep sequence-to-sequence models. These models... 

    The effect of phase information in speech enhancement and speech recognition

    , Article 2012 11th International Conference on Information Science, Signal Processing and their Applications, ISSPA 2012, 2 July 2012 through 5 July 2012 ; 2012 , Pages 1446-1447 ; 9781467303828 (ISBN) Langarani, M. S. E ; Veisi, H ; Sameti, H ; Sharif University of Technology
    2012
    Abstract
    The majority of speech enhancement methods perform noise removal in spectral domain and construct the enhanced speech signal from the estimated magnitude of clean speech and the phase of the noisy speech. In this paper, we show that by incorporating the phase information in the enhancement process, the quality and intelligibility of speech signal are improved. In our investigations, the minimum mean-square error short-time spectral amplitude and MMSE log-spectral amplitude methods are used to estimate the magnitude spectrum of speech signal. By conducting six classes of experiments, it is shown that by taking the phase information into account, overall SNR and PESQ measures are improved. In... 

    Telephony text-prompted speaker verification using i-vector representation

    , Article ICASSP, IEEE International Conference on Acoustics, Speech and Signal Processing - Proceedings, 19 April 2014 through 24 April 2014 ; Volume 2015-August , 2015 , Pages 4839-4843 ; 15206149 (ISSN) ; 9781467369978 (ISBN) Zeinali, H ; Kalantari, E ; Sameti, H ; Hadian, H ; Sharif University of Technology
    Institute of Electrical and Electronics Engineers Inc  2015
    Abstract
    I-vectors have proved to be the most effective features for text-independent speaker verification in recent researches. In this article a new scheme is proposed to utilize i-vectors in text-prompted speaker verification in a simple while effective manner. In order to examine this scheme empirically, a telephony dataset of Persian month names is introduced. Experiments show that the proposed scheme reduces the EER by 31% compared to the state-of-the-art State-GMM-MAP method. Furthermore it is shown that using HMM instead of GMM for universal background modeling leads to 15% reduction in EER  

    Speech signal modeling using multivariate distributions

    , Article Eurasip Journal on Audio, Speech, and Music Processing ; Volume 2015, Issue 1 , 2015 , Pages 1-14 ; 16874714 (ISSN) Aroudi, A ; Veisi, H ; Sameti, H ; Mafakheri, Z ; Sharif University of Technology
    Springer International Publishing  2015
    Abstract
    Using a proper distribution function for speech signal or for its representations is of crucial importance in statistical-based speech processing algorithms. Although the most commonly used probability density function (pdf) for speech signals is Gaussian, recent studies have shown the superiority of super-Gaussian pdfs. A large research effort has focused on the investigation of a univariate case of speech signal distribution; however, in this paper, we study the multivariate distributions of speech signal and its representations using the conventional distribution functions, e.g., multivariate Gaussian and multivariate Laplace, and the copula-based multivariate distributions as candidates.... 

    An evolutionary decoding method for HMM-based continuous speech recognition systems using particle swarm optimization

    , Article Pattern Analysis and Applications ; Vol. 17, issue. 2 , 2014 , pp. 327-339 Najkar, N ; Razzazi, F ; Sameti, H ; Sharif University of Technology
    2014
    Abstract
    The main recognition procedure in modern HMM-based continuous speech recognition systems is Viterbi algorithm. Viterbi algorithm finds out the best acoustic sequence according to input speech in the search space using dynamic programming. In this paper, dynamic programming is replaced by a search method which is based on particle swarm optimization. The major idea is focused on generating initial population of particles as the speech segmentation vectors. The particles try to achieve the best segmentation by an updating method during iterations. In this paper, a new method of particles representation and recognition process is introduced which is consistent with the nature of continuous... 

    Speech synthesis based on gaussian conditional random fields

    , Article Communications in Computer and Information Science ; Vol. 427, issue , 2014 , p. 183-193 Khorram, S ; Bahmaninezhad, F ; Sameti, H ; Sharif University of Technology
    2014
    Abstract
    Hidden Markov Model (HMM)-based synthesis (HTS) has recently been confirmed to be the most effective method in generating natural speech. However, it lacks adequate context generalization when the training data is limited. As a solution, current study provides a new context-dependent speech modeling framework based on the Gaussian Conditional Random Field (GCRF) theory. By applying this model, an innovative speech synthesis system has been developed which can be viewed as an extension of Context-Dependent Hidden Semi Markov Model (CD-HSMM). A novel Viterbi decoder along with a stochastic gradient ascent algorithm was applied to train model parameters. Also, a fast and efficient parameter... 

    Average voice modeling based on unbiased decision trees

    , Article Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics), Mons ; Volume 7911 LNAI , June , 2013 , Pages 89-96 ; 03029743 (ISSN) ; 9783642388460 (ISBN) Bahmaninezhad, F ; Khorram, S ; Sameti, H ; Sharif University of Technology
    2013
    Abstract
    Speaker adaptive speech synthesis based on Hidden Semi-Markov Model (HSMM) has been demonstrated to be dramatically effective in the presence of confined amount of speech data. However, we could intensify this effectiveness by training the average voice model appropriately. Hence, this study presents a new method for training the average voice model. This method guarantees that data from every speaker contributes to all the leaves of decision tree. We considered this fact that small training data and highly diverse contexts of training speakers are considered as disadvantages which degrade the quality of average voice model impressively, and further influence the adapted model and synthetic... 

    A novel approach to HMM-based speech recognition systems using particle swarm optimization

    , Article Mathematical and Computer Modelling ; Volume 52, Issue 11-12 , 2010 , Pages 1910-1920 ; 08957177 (ISSN) Najkar, N ; Razzazi, F ; Sameti, H ; Sharif University of Technology
    2010
    Abstract
    The main core of HMM-based speech recognition systems is Viterbi algorithm. Viterbi algorithm uses dynamic programming to find out the best alignment between the input speech and a given speech model. In this paper, dynamic programming is replaced by a search method which is based on particle swarm optimization algorithm. The major idea is focused on generating an initial population of segmentation vectors in the solution search space and improving the location of segments by an updating algorithm. Several methods are introduced and evaluated for the representation of particles and their corresponding movement structures. In addition, two segmentation strategies are explored. The first... 

    A novel approach to HMM-based speech recognition system using particle swarm optimization

    , Article BIC-TA 2009 - Proceedings, 2009 4th International Conference on Bio-Inspired Computing: Theories and Applications, 16 October 2009 through 19 October 2009 ; 2009 , Pages 296-301 ; 9781424438655 (ISBN) Najkar, N ; Razzazi, F ; Sameti, H ; Sharif University of Technology
    2009
    Abstract
    The main core of HMM-based speech recognition systems is the Viterbi Algorithm. Viterbi is performed using dynamic programming to find out the best alignment between input speech and given speech model. In this paper, dynamic programming is replaced by a search method which is based on particle swarm optimization algorithm. The major idea is focused on generating an initial population of segmentation vectors in the solution search space and improving the location of segments by an updating algorithm. Two methods are introduced for representation of each particle and movement structure. The results show that the effect of these factors is noticeable in finding the global optimum while... 

    Semi-supervised parallel shared encoders for speech emotion recognition

    , Article Digital Signal Processing: A Review Journal ; Volume 118 , 2021 ; 10512004 (ISSN) Pourebrahim, Y ; Razzazi, F ; Sameti, H ; Sharif University of Technology
    Elsevier Inc  2021
    Abstract
    Supervised speech emotion recognition requires a large number of labeled samples that limit its use in practice. Due to easy access to unlabeled samples, a new semi-supervised method based on auto-encoders is proposed in this paper for speech emotion recognition. The proposed method performed the classification operation by extracting the information contained in unlabeled samples and combining it with the information in labeled samples. In addition, it employed maximum mean discrepancy cost function to reduce the distribution difference when the labeled and unlabeled samples were gathered from different datasets. Experimental results obtained on different emotional speech datasets... 

    Using and evaluating new confidence measures in word-based isolated word recognizers

    , Article 2007 9th International Symposium on Signal Processing and its Applications, ISSPA 2007, Sharjah, 12 February 2007 through 15 February 2007 ; 2007 ; 1424407796 (ISBN); 9781424407798 (ISBN) Vaisipour, S ; Babaali, B ; Sameti, H ; Sharif University of Technology
    2007
    Abstract
    In this paper a method for detecting out of vocabulary words in isolated word recognizers is introduced, our method utilized new kinds of confidence measure. After recognition task was completed and consequently confidence measure was extracted, a classifier would accept or reject result of recognition task using this CM. We used two different kinds of confidence measure where for extracting each one a different information source was used. Amount of competition between hypotheses through the recognition task was used for extracting first CM. The second one was extracted using information about manner of distribution of feature vectors in the states of winner HMM model. Both of these CMs... 

    Accent-Invariant Automatic Speech Recognition via Saliency-Driven Spectrogram Masking

    , M.Sc. Thesis Sharif University of Technology Sameti, Mohammad Hossein (Author) ; Sameti, Hossein (Supervisor)
    Abstract
    Automatic speech recognition (ASR) refers to the process of converting human speech signals into their corresponding written text. This task is one of the fundamental problems in speech processing and natural language understanding, and it has witnessed remarkable advancements in recent decades, especially with the emergence of transformer-based architectures. Despite these improvements, modern ASR systems remain highly sensitive to accentual and dialectal variations, an issue that is particularly significant in languages such as Persian and English, where substantial linguistic diversity leads to increased word and character error rates. This sensitivity illustrates that, even with access... 

    Unsupervised induction of persian semantic verb classes based on syntactic information

    , Article Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics), Warsaw ; Volume 7912 LNCS , June , 2013 , Pages 112-124 ; 03029743 (ISSN) ; 9783642386336 (ISBN) Aminian, M ; Rasooli, M. S ; Sameti, H ; Sharif University of Technology
    2013
    Abstract
    Automatic induction of semantic verb classes is one of the most challenging tasks in computational lexical semantics with a wide variety of applications in natural language processing. The large number of Persian speakers and the lack of such semantic classes for Persian verbs have motivated us to use unsupervised algorithms for Persian verb clustering. In this paper, we have done experiments on inducing the semantic classes of Persian verbs based on Levin's theory for verb classes. Syntactic information extracted from dependency trees is used as base features for clustering the verbs. Since there has been no manual classification of Persian verbs prior to this paper, we have prepared a... 

    Mel-scaled Discrete Wavelet Transform and dynamic features for the Persian phoneme recognition

    , Article 2011 International Symposium on Artificial Intelligence and Signal Processing, AISP 2011, 15 June 2011 through 16 June 2011 ; June , 2011 , Pages 138-140 ; 9781424498345 (ISBN) Tavanaei, A ; Manzuri, M. T ; Sameti, H ; Sharif University of Technology
    2011
    Abstract
    In this paper we use a feature vector consisting of the Mel Frequency Discrete Wavelet Coefficients to recognize spoken phonemes in the Persian language. The purpose of using wavelet in feature extraction is to benefit from its multi resolution analysis and localization property in time and frequency domains. The MFDWCs are obtained by applying the Discrete Wavelet Transform (DWT) to the Mel-scaled log filter bank energies of a speech frame. Feature vectors are used for the HMM-based phoneme recognition on a portion of the FarsDat Persian language database consisting of 35 hour recorded data for training and 15 hour for testing. We evaluate the performance of new features for clean speech... 

    Far-field continuous speech recognition system based on speaker localization and sub-band beamforming

    , Article 6th IEEE/ACS International Conference on Computer Systems and Applications, AICCSA 2008, Doha, 31 March 2008 through 4 April 2008 ; 2008 , Pages 495-500 ; 9781424419685 (ISBN) Asaei, A ; Taghizadeh, M. J ; Sameti, H ; Sharif University of Technology
    2008
    Abstract
    This paper proposes a Distant Speech Recognition system based on a novel speaker Localization and Beamforming (SRLB) algorithm. To localize the speaker an algorithm based on Steered Response Power by utilizing harmonic structures of speech signal is proposed. This new scheme has the ability of speaker verification by fundamental frequency variation: therefore it can be utilized in the design of a speech recognition system for verified speakers. Then the performance of the Farsi speech recognition engine is evaluated under notorious conditions of noise and reverberation. Simulation results and tests on real data shows that by utilizing proposed localization scheme, recognition accuracy... 

    Non-speaker information reduction from Cosine Similarity Scoring in i-vector based speaker verification

    , Article Computers and Electrical Engineering ; Volume 48 , November , 2015 , Pages 226–238 ; 00457906 (ISSN) Zeinali, H ; Mirian, A ; Sameti, H ; BabaAli, B ; Sharif University of Technology
    Elsevier Ltd  2015
    Abstract
    Cosine similarity and Probabilistic Linear Discriminant Analysis (PLDA) in i-vector space are two state-of-the-art scoring methods in speaker verification field. While PLDA usually gives better accuracy, Cosine Similarity Scoring (CSS) remains a widely used method due to simplicity and acceptable performance. In this domain, several channel compensation and score normalization methods have been proposed to improve the performance. We investigate non-speaker information in cosine similarity metric and propose a new approach to remove it from the decision making process. I-vectors hold a large amount of non-speaker information such as channel effects, language, and phonetic content. This type...