효율적 AI 피드백 설계를 위한 LLM 성능 비교 연구: 대학생의 설득적 글쓰기 과제를 중심으로

A comparative evaluation of Large Language Models for efficient AI feedback design: Evidence from undergraduate persuasive writing tasks

초록

This study compares the feedback profiles produced by ChatGPT and Gemini on the same undergraduate persuasive-writing task under an identical prompt. Using an RTPR-based LLM-as-a-Judge framework with counter-balancing, we tested seven rubric criteria and the overall mean via independent two-sample t-tests. Significant differences emerged for most criteria and for the overall mean (p<.001); Lexical Accuracy favored ChatGPT over Gemini (p=.017), whereas Sentence Clarity was not significant (p=.084). These findings indicate systematic, model-specific strengths and weaknesses even under identical inputs, suggesting instructionally targeted model choice (e.g., argument/evidence/organization vs. lexical refinement). The study was limited to two models and a single genre; no multiple-comparison correction was applied.

키워드

LLM; 인공지능 피드백(AI feedback); 설득적 글쓰기(persuasive writing); ChatGPT; Gemini
제목
효율적 AI 피드백 설계를 위한 LLM 성능 비교 연구: 대학생의 설득적 글쓰기 과제를 중심으로
제목 (타언어)
A comparative evaluation of Large Language Models for efficient AI feedback design: Evidence from undergraduate persuasive writing tasks
저자
안미애; 김수연
DOI
10.21296/jls.2025.09.114.113
발행일
2025-09
유형
Y
저널명
언어과학연구
호
114
페이지
113 ~ 141