improving-llm-generated-process-model-quality-through-reinforcement-learning-the-role-of-reward-function-design-c64ee1c2·1 events·first seen Aliases: Improving LLM-Generated Process Model Quality Through Reinforcement Learning: The Role of Reward Function Design
Researchers present a systematic study of reward function design for reinforcement learning applied to LLM-based BPMN process model generation, training Llama 3.1 8B and Qwen 2.5 14B across 48 configurations using Group Sequence Policy Optimization. Key findings: RL substantially improves syntactic and pragmatic quality while preserving semantic fidelity, equal reward weighting outperforms targeted weighting, and reward design effects interact with model architecture in non-trivial ways. The paper argues reward composition is as consequential as the decision to apply RL at all, with implications for any multi-dimensional structured generation task.