CORTEXA
← Browse
arxivcs.RO2026-07-22

GRACE: Gradient-Free Robot Action Generation via Combined Diffusion-MPPI Posterior Mean Estimation

Leesai Park, Jiho HOng, Sanghyun Kim

Diffusion policies generate multimodal robot action sequences from demonstrations, but steering them toward deployment-time constraints typically relies on differentiable guidance costs. This excludes many practical safety constraints, such as binary collision checks, joint limits, and black-box rollout costs that are nondifferentiable. We propose Gradient-free Robot Action generation via Combined diffusion-MPPI posterior mean Estimation (GRACE), which guides a pretrained diffusion policy with Model Predictive Path Integral (MPPI) control using only forward cost evaluations. Building on the common score-ascent structure of diffusion and MPPI, GRACE constructs a cost-conditioned guidance posterior at each reverse step and estimates its mean with a single MPPI update centered at the diffusion reverse mean. For differentiable costs, GRACE recovers conventional gradient guidance under a first-order, matched-covariance approximation. GRACE attains higher success rates than diffusion-based and sampling-based baselines in simulation. On a real 7-DoF manipulator, GRACE avoids a deployment-time obstacle that the unguided prior collides with in every trial. Code and experiment videos are available at https://anonymous.4open.science/w/grace-70BB/.

View free PDFSource page

Related papers

arxivcs.ROcs.IR2026-07-23

StARS: Socially Appropriate Robot Actions via a Recommender System-Driven Approach

Erencem Ozbey, Fethiye Irmak Dogan, Jin Huang, Hatice Gunes

Social appropriateness in human-robot interaction (HRI) is not universal: different people can judge the same robot action differently in the same situation. To capture this inter-subject variability, we reformulate socially appropriate action generation as a preference modelling…

View free PDFSource page
arxivcs.RO2026-07-31

Temporal Policy: History-Initialized Action Generation for Robotic Learning from Demonstration

Dylan Miller, Martin Jagersand

By relying on independent couplings from uninformative Gaussian priors, standard diffusion and flow matching models are forced to learn complex, high-cost vector fields to reach the physical action space. Generative models excel at capturing multimodal behaviors for robotic Learn…

View free PDFSource page
arxivcs.ROcs.LG2026-07-10

GenVid2Robot: From Video Generation to Robot Manipulation via Rigid-Geometric Consistency

Haohui Huang, Xi Yuan, Panpan Liao, Tao Teng, Chenguang Yang, Jing Guo, et al.

Generated videos provide useful visual motion priors for robot manipulation, but their visual plausibility does not imply physical executability. A generated video usually lacks metric geometry, grasp grounding, robot kinematic feasibility, and execution-time feedback, which make…

View free PDFSource page
arxivcs.RO2026-06-28

CORE: Common Outcome Regularities from Action-Free Visual Demonstrations for Robot Manipulation

Juyi Sheng, Jincheng Li, Mingxin Tan, Mengyuan Liu

Robot imitation learning often relies on costly robot demonstrations, while abundant action-free visual demonstrations, such as human videos, are difficult to use because they lack robot-executable actions and suffer from embodiment gaps. We propose CORE, a policy learning framew…

View free PDFSource page
arxivcs.RO2026-07-06

PRISM: Personalized Robotic Dataset Generation via Image-based Scene and Motion Synthesis

Dogyu Ko, Haneul Kim, Chanyoung Yeo, Dowoon Lee, Taeho Park, Hyoseok Hwang

Recent advances in large-scale pretrained vision-language-action models have improved robot policy learning, but directly deploying such policies in user-specific environments remains challenging due to limited generalization, which inevitably requires collecting a dataset tailor…

View free PDFSource page