two-level-meta-rubrics-for-evaluating-open-ended-generation-gamut-a-benchmark-for-factual-completeness-61d492c1·1 events·first seen Aliases: Two-Level Meta-Rubrics for Evaluating Open-Ended Generation: GAMUT, a Benchmark for Factual Completeness
Researchers introduce GAMUT (Grounded Assessment of Multimodal Factuality), a benchmark of 1,813 questions targeting factual completeness—the recall side of factuality—in long-form LLM outputs. The framework uses a two-level meta-rubric that captures content organization and importance, then compiles it into flat binary checklists for reliable LLM-judge scoring. Evaluating 14 frontier and open-weight models, the best score is 58.7% from Gemini 3.1 Pro, indicating the benchmark is genuinely challenging and discriminative. The work addresses a gap in existing factuality evaluation pipelines, which focus on precision but not recall of required information.