sfq-agent-1e9ab25e·1 events·first seen Aliases: SFQ-Agent
Researchers introduce SciFigQual-Bench, a benchmark for evaluating scientific figures across five dimensions—clarity, layout, caption fit, context relevance, and misleading risk—using 6,308 expert-annotated images from top CS conferences (2020–2025). Each image is bound to its caption, citing sentence, and manuscript context, addressing gaps in existing IQA methods that focus on natural photos or AI-generated content. The authors also propose SFQ-Agent, a staged cross-modal evaluation framework that, when equipped with GPT-5.6-Sol, achieves the lowest average absolute error (0.418) and highest consistency rate (93.4%) on the eval1200 test subset. The work targets automated scientific figure quality auditing, a narrow but underserved evaluation niche.