freeform-preference-learning-for-robotic-manipulation-199824bb·1 events·first seen Aliases: Freeform Preference Learning for Robotic Manipulation
Freeform Preference Learning (FPL) replaces binary preference labels in robot learning with natural-language preference axes (e.g., speed, safety, carefulness), allowing annotators to provide pairwise preferences along each axis. A language-conditioned reward model is trained from these annotations and used to optimize a reward-conditioned policy across multiple human-specified dimensions. Evaluated on four real-world and two simulated long-horizon manipulation tasks, FPL outperforms sparse-reward and binary-preference baselines by 38 percentage points, and enables test-time behavioral steering without retraining.