researchqa-732afaba·1 events·first seen Aliases: ResearchQA
SERPO (Self-Evolving Rubric Policy Optimization) is a new method for test-time reinforcement learning that extends self-evolution to open-ended generation tasks without requiring labeled feedback, external reward models, or stronger judge models. The approach co-evolves response archives, query-specific rubrics, and policy parameters in a closed loop using a Good-Normal-Bad response organization scheme and probabilistic criterion scoring. Evaluated across six benchmarks, SERPO improves HealthBench and ResearchQA by up to 20.63 and 20.31 points over base models and raises the macro-average by up to 8.06 points. The method addresses a key limitation of existing TTRL approaches that rely on answer voting and cannot generalize to tasks without canonical answers.