[PDF][PDF] End-to-End Deep Learning Framework for Speech Paralinguistics Detection Based on Perception Aware Spectrum.

D Cai, Z Ni, W Liu, W Cai, G Li, M Li, D Cai, Z Ni… - …, 2017 - isca-archive.org
D Cai, Z Ni, W Liu, W Cai, G Li, M Li, D Cai, Z Ni, W Liu, W Cai
INTERSPEECH, 2017isca-archive.org
In this paper, we propose an end-to-end deep learning framework to detect speech
paralinguistics using perception aware spectrum as input. Existing studies show that speech
under cold has distinct variations of energy distribution on low frequency components
compared with the speech under 'healthy'condition. This motivates us to use perception
aware spectrum as the input to an end-to-end learning framework with small scale dataset.
In this work, we try both Constant Q Transform (CQT) spectrum and Gammatone spectrum in …
Abstract
In this paper, we propose an end-to-end deep learning framework to detect speech paralinguistics using perception aware spectrum as input. Existing studies show that speech under cold has distinct variations of energy distribution on low frequency components compared with the speech under ‘healthy’condition. This motivates us to use perception aware spectrum as the input to an end-to-end learning framework with small scale dataset. In this work, we try both Constant Q Transform (CQT) spectrum and Gammatone spectrum in different end-toend deep learning networks, where both spectrums are able to closely mimic the human speech perception and transform it into 2D images. Experimental results show the effectiveness of the proposed perception aware spectrum with end-to-end deep learning approach on Interspeech 2017 Computational Paralinguistics Cold sub-Challenge. The final fusion result of our proposed method is 8% better than that of the provided baseline in terms of UAR.
isca-archive.org
以上显示的是最相近的搜索结果。 查看全部搜索结果