Evaluating Creativity: Can LLMs Be Good Evaluators in Creative Writing Tasks?

Citations

WEB OF SCIENCE

9
Citations

SCOPUS

19

초록

The evaluation of creative writing has long been a complex and subjective process, made even more intriguing by the rise of advanced Artificial Intelligence (AI) tools like Large Language Models (LLMs). This study evaluates the potential of LLMs as reliable and consistent evaluators of creative texts, directly comparing their performance with traditional human evaluations. The analysis focuses on key creative criteria, including fluency, flexibility, elaboration, originality, usefulness, and specific creativity strategies. Results demonstrate that LLMs provide consistent and objective evaluations, achieving higher Inter-Annotator Agreement (IAA) compared with human evaluators. However, LLMs face limitations in recognizing nuanced, culturally specific, and context-dependent aspects of creativity. Conversely, human evaluators, despite lower consistency and higher subjectivity, exhibit strengths in capturing deeper contextual insights. These findings highlight the need for the further refinement of LLMs to address the complexities of creative writing evaluation.

키워드

large language models (LLMs) evaluation; creative writing evaluation; creativity; AI evaluation; human evaluation
제목
Evaluating Creativity: Can LLMs Be Good Evaluators in Creative Writing Tasks?
저자
Kim, Sungeun; Oh, Dongsuk
DOI
10.3390/app15062971
발행일
2025-03-10
유형
Article
저널명
APPLIED SCIENCES-BASEL
권
15
호
6