claim
Many existing hallucination benchmarks rely on one-dimensional metrics such as Accuracy, Accept/Refusal rates, BLEU, and BERTScore, which limits the interpretability of results and obscures the underlying causes of Large Language Model performance issues.
Authors
Sources
- A Knowledge Graph-Based Hallucination Benchmark for Evaluating ... arxiv.org via serper