CORTEXA
← Browse
arxivcs.ROcs.AI2026-06-30

LLM-Powered Interactive Robotic Action Synthesis from Multimodal Speech, Gestures, and Music

Snehasis Banerjee, Ranjan Dasgupta

The quest for intuitive and natural human-robot interaction (HRI) remains a significant challenge in robotics. Traditional methods often rely on rigid, pre-programmed commands that limit the robot's expressiveness and adaptability. This paper introduces a novel framework that leverages the reasoning capabilities of Large Language Models (LLMs) to synthesize complex robotic actions from a rich tapestry of multimodal human inputs: natural speech, hand gestures, and music/sound beats. Our system architecture integrates a speech transcription model, a gesture recognition module, and a signal processing pipeline for beat detection. These processed inputs are contextualized using prompt templates and fed into a LLM. The LLM, informed by a predefined robot action space, reasons over the combined inputs to generate a coherent sequence of actions. This sequence is dispatched to an action queue for execution on a quadruped robot over ROS. The framework has ability to interpret and fuse semantic commands from speech, deictic information from gestures, and rhythmic cues from music. This work represents a step towards creating robots that can interact with humans in a more fluid, creative, and context-aware manner.

View free PDFSource page

Related papers

arxivcs.ROcs.AI2026-07-01

From Technical Metrics to User Perception: A User Study of a Multimodal Human-Robot Interaction System for Object Detection and Grasping

Jian Song, Tian Zi, Shen Guanting

Improvements in the technical performance of human--robot interaction (HRI) systems do not automatically translate into differences that human users can detect during live interaction. This paper investigates whether a 15 percentage point gain in end-to-end task success (from 75%…

View free PDFSource page
arxivcs.ROcs.AIcs.CV2026-07-15

Semantic Anchoring for Robotic Action Representations

Yuan Xu, Youheng Shi, Chengyang Li, Wentao Zhu, Yizhou Wang

Vision-Language-Action (VLA) models inherit rich semantic representations from pretrained Vision-Language Models, yet fine-tuning on limited robot demonstrations degrades this structure and undermines generalization. A fundamental question therefore arises: what constitutes a goo…

View free PDFSource page
arxivcs.ROcs.AI2026-07-16

Interventional Causal Circuits for Safe Robot Action Testing and Failure Recovery

Naren Vasantakumaar, Tom Schierenbeck, Michael Beetz

Safe physical AI for robot actions are required not only likely to succeed but tested to be safe before execution. In practice, however, formal testing of motion parameters is computationally expensive, and the cost scales poorly with the dimensionality of the action space. When…

View free PDFSource page
arxivcs.ROcs.AIcs.HC2026-07-07

Responsible Personalisation: The Double-Edged Sword of Personalisation in Human-Robot Interaction

Antonio Andriella, Jauwairia Nasir, Andrea Rezzani, Alyssa Kubota, Dimitri Lacroix, Tamlin Love, et al.

While personalisation is becoming a defining capability in human-robot interaction (HRI), the existing literature on responsible personalisation remains fragmented, offering isolated accounts of ethical risks without a structured understanding of how they emerge across interactio…

View free PDFSource page
arxivcs.ROcs.AI2026-06-30

A Modular Vision-Language-Action Robotics Framework for Indoor Environments

Anindya Jana, Snehasis Banerjee, Arup Sadhu, Ranjan Dasgupta

This paper presents an integrated system for the CMU Vision-Language-Action (VLA) Challenge, designed to enable an autonomous agent to perform complex tasks based on natural language instructions. Our framework employs a modular architecture that orchestrates environment mapping,…

View free PDFSource page