Evaluating Rater Effects of Large Language Models in Automated Essay Scoring: GPT, Claude, Gemini, and DeepSeek

Λεπτομέρειες βιβλιογραφικής εγγραφής
Τίτλος: Evaluating Rater Effects of Large Language Models in Automated Essay Scoring: GPT, Claude, Gemini, and DeepSeek
Γλώσσα: English
Συγγραφείς: Hong Jiao (ORCID 0000-0001-5014-6698), Dan Song (ORCID 0000-0002-7466-6150), Won-Chan Lee (ORCID 0000-0002-5319-668X)
Πηγή: Educational Measurement: Issues and Practice. 2026 45(2).
Διαθεσιμότητα: Wiley. Available from: John Wiley & Sons, Inc. 111 River Street, Hoboken, NJ 07030. Tel: 800-835-6770; e-mail: cs-journals@wiley.com; Web site: https://www.wiley.com/en-us
Peer Reviewed: Y
Page Count: 17
Ημερομηνία έκδοσης: 2026
Τύπος εγγράφου: Journal Articles
Reports - Research
Descriptors: Artificial Intelligence, Natural Language Processing, Automation, Computer Assisted Testing, Scoring, Evaluators, Interrater Reliability, Accuracy, Item Response Theory
DOI: 10.1111/emip.70018
ISSN: 0731-1745
1745-3992
Περίληψη: Large language models (LLMs) have been widely explored for automated scoring in educational assessment to facilitate learning and instruction. However, empirical evidence regarding which LLMs produce the most reliable scores and induce the least rater effects remains limited. This study compared 10 LLMs (ChatGPT 3.5, ChatGPT 4, ChatGPT 4o, OpenAI o1, Claude 3.5 Sonnet, Gemini 1.5, Gemini 1.5 Pro, Gemini 2.0, DeepSeek V3, and DeepSeek R1) with human expert raters in scoring two types of writing tasks. Their performance was evaluated in terms of score accuracy, intra-rater consistency, and rater effects estimated using the Many-Facet Rasch model. Although the results generally supported the use of ChatGPT 4o, Gemini 1.5 Pro, and Claude 3.5 Sonnet with high scoring accuracy, better intra-rater consistency, and less rater effects, the study is not intended to support substantive comparisons or rankings of LLMs or to identify a single "best" model, given the small sample size.
Abstractor: As Provided
Entry Date: 2026
Αριθμός Καταχώρησης: EJ1507074
Βάση Δεδομένων: ERIC
Περιγραφή
ISSN:0731-1745
1745-3992
DOI:10.1111/emip.70018