A Technique for Estimating Intensity of Emotional Expressions and Speaking Styles in Speech Based on Multiple-Regression HSMM

    • NOSE Takashi
    • Interdisciplinary Graduate School of Science and Engineering, Tokyo Institute of Technology
    • KOBAYASHI Takao
    • Interdisciplinary Graduate School of Science and Engineering, Tokyo Institute of Technology

抄録

In this paper, we propose a technique for estimating the degree or intensity of emotional expressions and speaking styles appearing in speech. The key idea is based on a style control technique for speech synthesis using a multiple regression hidden semi-Markov model (MRHSMM), and the proposed technique can be viewed as the inverse of the style control. In the proposed technique, the acoustic features of spectrum, power, fundamental frequency, and duration are simultaneously modeled using the MRHSMM. We derive an algorithm for estimating explanatory variables of the MRHSMM, each of which represents the degree or intensity of emotional expressions and speaking styles appearing in acoustic features of speech, based on a maximum likelihood criterion. We show experimental results to demonstrate the ability of the proposed technique using two types of speech data, simulated emotional speech and spontaneous speech with different speaking styles. It is found that the estimated values have correlation with human perception.

収録刊行物

IEICE transactions on information and systems  

IEICE transactions on information and systems 93(1), 116-124, 2010-01-01 

(社)電子情報通信学会

参考文献:  28件

参考文献を見るにはログインが必要です。ユーザIDをお持ちでない方は新規登録してください。

各種コード

  • NII論文ID(NAID) :
    10026813214
  • NII書誌ID(NCID) :
    AA10826272
  • 本文言語コード :
    ENG
  • 資料種別 :
    ART
  • ISSN :
    09168532
  • 収録DB :
    CJP書誌  J-STAGE