arxivcs.CVcs.GR2026-07-17
Per-Stroke Temporal Control for Text-to-Motion via Action Units and Action-Detection Guidance
Text-to-motion models are competent at the action a prompt names but unreliable at when each stroke lands: four punches alternating left and right rarely return four separable strokes. We introduce typed temporal events called Action Units (AUs) that make the individual stroke --…