opik-4dc4053d·2 events·first seen Aliases: Opik
A preprint from arXiv reports an empirical study comparing RAG evaluation metrics from four libraries—Ragas, DeepEval, RAGChecker, and Opik—against human annotator scores and standard metrics like recall on a business question-answering dataset. The study conducts correlation analysis between automated metrics and human evaluators, finding and documenting limitations of current RAG evaluation methodology. The paper is an English translation of work originally presented at the French-language EvalLLM workshop.
Opik is an open-source toolkit from Comet ML for debugging, evaluating, and monitoring LLM applications, RAG systems, and agentic workflows. It provides tracing, automated evaluations, and production dashboards. The project has accumulated nearly 20K GitHub stars, indicating meaningful adoption in the practitioner community.