Enhancing QA System Evaluation: An In-Depth Analysis of Metrics and Model-Specific Behaviors

Citations

SCOPUS

0

초록

The purpose of this study is to examine how evaluation metrics influence the perception and performance of question answering (QA) systems, particularly focusing on their effectiveness in QA tasks. We compare four different models: BERT, BioBERT, Bio- ClinicalBERT, and RoBERTa, utilizing ten EPIC-QA questions to assess each model’s answer extraction performance. The analysis employs both semantic and lexical metrics. The outcomes reveal clear model-specific behaviors: Bio-ClinicalBERT initially identified irrelevant phrases before focusing on relevant information, whereas BERT and BioBERT continually converge on similar answers, exhibiting a high degree of similarity. RoBERTa, on the other hand, demonstrates effective use of long-range dependencies in text. Semantic metrics outperform lexical metrics, with BERTScore attaining the maximum accuracy (0.97), highlighting the significance of semantic evaluation. Our findings indicate that the choice of evaluation metrics significantly influences the perceived efficacy of models, suggesting that semantic metrics offer more nuanced and insightful assessments of QA system performance. This study contributes to the field of natural language processing and machine learning by providing guidelines for selecting evaluation metrics that align with the strengths and weaknesses of various QA approaches.

키워드

BERT; evaluation metrics; natural language processing; question answering systems; transformer models
제목
Enhancing QA System Evaluation: An In-Depth Analysis of Metrics and Model-Specific Behaviors
저자
Kim, Heesop; Ademola, Aluko
DOI
10.1633/JISTaP.2025.13.1.6
발행일
2025-03
유형
Article
저널명
Journal of Information Science Theory and Practice
권
13
호
1
페이지
85 ~ 98