claim
Existing benchmarks rarely probe the depth of knowledge beyond surface-level details because they favor closed or multiple-choice questions over open-ended questions, as noted by Hendrycks et al. (2021) and Rahman et al. (2024).
Authors
Sources
- A Knowledge Graph-Based Hallucination Benchmark for Evaluating ... arxiv.org via serper
Referenced by nodes (1)
- hallucination concept