A comparative study of text preprocessing approaches for topic detection of user utterances
R Sergienko, M Shan, W Minker - Proceedings of the Tenth …, 2016 - aclanthology.org
R Sergienko, M Shan, W Minker
Proceedings of the Tenth International Conference on Language …, 2016•aclanthology.orgThe paper describes a comparative study of existing and novel text preprocessing and
classification techniques for domain detection of user utterances. Two corpora are
considered. The first one contains customer calls to a call centre for further call routing; the
second one contains answers of call centre employees with different kinds of customer
orientation behaviour. Seven different unsupervised and supervised term weighting
methods were applied. The collective use of term weighting methods is proposed for …
classification techniques for domain detection of user utterances. Two corpora are
considered. The first one contains customer calls to a call centre for further call routing; the
second one contains answers of call centre employees with different kinds of customer
orientation behaviour. Seven different unsupervised and supervised term weighting
methods were applied. The collective use of term weighting methods is proposed for …
Abstract
The paper describes a comparative study of existing and novel text preprocessing and classification techniques for domain detection of user utterances. Two corpora are considered. The first one contains customer calls to a call centre for further call routing; the second one contains answers of call centre employees with different kinds of customer orientation behaviour. Seven different unsupervised and supervised term weighting methods were applied. The collective use of term weighting methods is proposed for classification effectiveness improvement. Four different dimensionality reduction methods were applied: stop-words filtering with stemming, feature selection based on term weights, feature transformation based on term clustering, and a novel feature transformation method based on terms belonging to classes. As classification algorithms we used k-NN and a SVM-based algorithm. The numerical experiments have shown that the simultaneous use of the novel proposed approaches (collectives of term weighting methods and the novel feature transformation method) allows reaching the high classification results with very small number of features.
aclanthology.org
以上显示的是最相近的搜索结果。 查看全部搜索结果