arxivcs.LGcs.AI2026-07-23
Multi-turn RL with Structural and Performance Aware Rewards for CUDA Kernel Generation
Quazi Ishtiaque Mahmud, Nesreen K. Ahmed, Ali Jannesari
Reinforcement Learning with Verifiable Rewards (RLVR) has emerged as a powerful technique to enhance the reasoning capacity of LLMs for optimized code generation. However, existing RLVR approaches primarily rely on outcome-based signals such as correctness and speedup, overlookin…