measurement
Applying difficulty-based weighting to the benchmark decreases the mean accuracy by 0.05%, from 45.30% to 45.25%.
Authors
Sources
- A Knowledge Graph-Based Hallucination Benchmark for Evaluating ... arxiv.org via serper
Referenced by nodes (1)
- KGHaluBench concept