User-Adaptive A Posteriori Restoration for Incorrectly Segmented Utterances in Spoken Dialogue Systems

Authors

  • Kazunori Komatani The Institute of Scientific and Industrial Research, Osaka University
  • Naoki Hotta Graduate School of Engineering, Nagoya University
  • Satoshi Sato Graduate School of Engineering, Nagoya University
  • Mikio Nakano Honda Research Institute Japan Co., Ltd.

DOI:

https://doi.org/10.5087/dad.2017.209

Abstract

Ideally, the users of spoken dialogue systems should be able to speak at their own tempo. Thus, the systems needs to interpret utterances from various users correctly, even when the utterances contain pauses. In response to this issue, we propose an approach based on a posteriori restoration for incorrectly segmented utterances. A crucial part of this approach is to determine whether restoration is required. We use a classification-based approach, adapted to each user. We focus on each user’s dialogue tempo, which can be obtained during the dialogue, and determine the correlation between each user’s tempo and the appropriate thresholds for classification. A linear regression function used to convert the tempos into thresholds is also derived. Experimental results show that the proposed user adaptation approach applied to two restoration classification methods, thresholding and decision trees, improves classification accuracies by 3.0% and 7.4%, respectively, in cross validation.

Downloads

Published

2017-12-15

Issue

Section

Articles