arxivcs.CLcs.AIcs.LG2026-07-07
Improving LLM-Generated Process Model Quality Through Reinforcement Learning: The Role of Reward Function Design
Alexander Rombach, Chantale Lauer, Nijat Mehdiyev
Large language models (LLMs) can generate BPMN process models from natural-language descriptions, yet supervised fine-tuning (SFT) limits their output quality to the patterns present in the training data. Reinforcement learning (RL) can optimize beyond this ceiling using external…