AI 뉴스로 돌아가기

Judge Arena: Benchmarking LLMs as Evaluators

2024.11.1911원문 보기
Judge Arena: Benchmarking LLMs as Evaluators