WEKO3
アイテム
KNUIR at the NTCIR-18 AEOLLM: Automatic Evaluation of LLMs
https://doi.org/10.20736/0002002027
https://doi.org/10.20736/00020020275861b5fb-1206-4ebd-949f-bf92dd7b9a38
| 名前 / ファイル | ライセンス | アクション |
|---|---|---|
|
|
|
| アイテムタイプ | デフォルトアイテムタイプ(フル)(1) | |||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 公開日 | 2025-06-06 | |||||||||||
| タイトル | ||||||||||||
| タイトル | KNUIR at the NTCIR-18 AEOLLM: Automatic Evaluation of LLMs | |||||||||||
| 言語 | en | |||||||||||
| 作成者 |
Yumi Kim
× Yumi Kim
× Meen Chul Kim
× Jongwook Lee
|
|||||||||||
| 内容記述 | ||||||||||||
| 内容記述タイプ | Abstract | |||||||||||
| 内容記述 | In this study, we aim to propose automated evaluation methods of LLMs that approximate human judgment by exploring and comparing two distinct approaches: (1) LLM-based scoring, which utilizes GPT models with prompt engineering, and (2) feature-based machine learning, using transformer-based metrics such as BERTScore, semantic similarity, and keyword coverage. As part of this research, we participated in the NTCIR-18 Automatic Evaluation of LLMs (AEOLLM) task. We submitted the results of the test data set and the reserved data set to NTCIR-18 and analyzed the results obtained. The results show that GPT-4o Mini (with the updated prompt) achieved the highest performance, while the feature-based approach performed competitively, surpassing GPT-3.5 Turbo and showing a small gap with GPT-4o Mini. LLM-based methods offered scalability but lacked explainability, whereas feature-based approaches provided better interpretability but required extensive tuning, highlighting the trade-offs between the two strategies. Throughout the analysis, We expect that the findings of our work will provide insights into the understanding of human judgment and automated evaluation of LLMs. | |||||||||||
| 言語 | en | |||||||||||
| 出版者 | ||||||||||||
| 出版者 | NII Institutional Repository | |||||||||||
| 言語 | en | |||||||||||
| 日付 | ||||||||||||
| 日付 | 2025-06-06 | |||||||||||
| 日付タイプ | Issued | |||||||||||
| 言語 | ||||||||||||
| 言語 | eng | |||||||||||
| 資源タイプ | ||||||||||||
| 資源タイプ識別子 | http://purl.org/coar/resource_type/c_5794 | |||||||||||
| 資源タイプ | conference paper | |||||||||||
| ID登録 | ||||||||||||
| ID登録 | 10.20736/0002002027 | |||||||||||
| ID登録タイプ | JaLC | |||||||||||
| 関連情報 | ||||||||||||
| 関連タイプ | isReferencedBy | |||||||||||
| 識別子タイプ | URI | |||||||||||
| 関連識別子 | https://research.nii.ac.jp/ntcir/ntcir-18/index.html | |||||||||||
| 言語 | en | |||||||||||
| 関連名称 | NTCIR-18 Conference | |||||||||||
| 開始ページ | ||||||||||||
| 開始ページ | none | |||||||||||
| 会議記述 | ||||||||||||
| 会議名 | NTCIR-18 Conference | |||||||||||
| 言語 | en | |||||||||||
| 回次 | 18 | |||||||||||
| 主催機関 | National Institute of Informatics | |||||||||||
| 言語 | en | |||||||||||
| 開始年 | 2025 | |||||||||||
| 開始月 | 6 | |||||||||||
| 開始日 | 10 | |||||||||||
| 終了年 | 2025 | |||||||||||
| 終了月 | 6 | |||||||||||
| 終了日 | 13 | |||||||||||
| 開催期間 | June 10-13, 2025 | |||||||||||
| 言語 | en | |||||||||||
| 開催会場 | National Institute of Informatics | |||||||||||
| 言語 | en | |||||||||||
| 開催国 | JPN | |||||||||||