查看文章

A study on the impact of fatigue on human raters when scoring speaking responses

作者

Guangming Ling, Pamela Mollaun, Xiaoming Xi

发表日期

2014/10

期刊

Language Testing

卷号

期号

页码范围

479-499

出版商

Sage Publications

简介

The scoring of constructed responses may introduce construct-irrelevant factors to a test score and affect its validity and fairness. Fatigue is one of the factors that could negatively affect human performance in general, yet little is known about its effects on a human rater’s scoring quality on constructed responses. In this study, we compared the scoring quality of 72 raters under four shift conditions differing on the shift length (total scoring time in a day) and session length (time continuously spent on a task). About 14,000 audio responses to four TOEFL iBT speaking tasks were scored, including 5446 validity responses that have pre-assigned “true” scores used to measure scoring accuracy. Our results suggest that the overall scoring accuracy is high for the TOEFL iBT Speaking Test, but varying levels of rating accuracy and consistency exist across shift conditions. The raters working the shorter shifts or shorter sessions …

引用总数

被引用次数：58

2016201720182019202020212022202320244 5 10 8 10 4 3 9 5

学术搜索中的文章

A study on the impact of fatigue on human raters when scoring speaking responses

G Ling, P Mollaun, X Xi - Language Testing, 2014

被引用次数：58 相关文章所有 6 个版本