every-eval-ever-a6317c3d·2 events·first seen Aliases: Every Eval Ever
Hugging Face is featuring results from the Every Eval Ever (EEE) community evaluation initiative directly on model pages, surfacing community-driven benchmark coverage alongside official evaluations. This integration makes a broader set of evaluation signals visible to practitioners browsing the Hub. The move reflects growing interest in community-sourced evals as a complement to lab-run benchmarks.
Researchers introduce Every Eval Ever, a shared schema and crowdsourced repository designed to standardize AI evaluation results across incompatible formats, frameworks, and sources. The system ingests results from evaluation harnesses, papers, leaderboards, and custom repositories into a single JSON document format, with optional per-instance output storage. The repository, hosted on Hugging Face, currently covers 22,235 models, 2,273 unique benchmarks, and 31 evaluation formats. The work addresses a persistent infrastructure problem in AI evaluation science: divergent scores for nominally identical evaluations and scattered, incomparable metadata.