craft-a7e44554·1 events·first seen Aliases: CRAFT
CRAFT is a new method that converts rubric-based evaluation datasets into model-specific capability diagnoses by extracting capability descriptions from prompt-rubric pairs, clustering them into a hierarchical tree, and scoring models at each node to identify weak capabilities. The identified weaknesses then direct targeted supervised fine-tuning data generation. Evaluated across four open-source models, two professional domains (finance and legal), and 13 held-out benchmarks, CRAFT outperforms prompt-level clustering and random data generation baselines in most settings. The approach addresses a gap in evaluation pipelines by explaining not just where models fail but why, and linking that diagnosis directly to post-training improvement.