Editorial: Embodied perspectives on sound and music AI
Çağrı Erdem, Anna Xambó Sedó, Stefania Serafin, Carsten Griwodz
Contemporary AI research often treats body and mind as separate domains, overlooking how our bodies shape mental and emotional states and how perception, environment, and action are deeply intertwined.The Embodied Perspectives on Musical Artificial Intelligence workshop, held at the University of Oslo, brought these concerns into focus by examining embodied cognition and its potential to reshape how we design, experience, and evaluate emerging technologies in the context of sound and music. The discussions at this workshop motivated the present Research Topic, which aims to strengthen the incorporation of embodied principles into AI across research, education, innovation, well-being, and the arts.Rather than treating AI solely in technical terms, the sixteen contributions collected here explore its potential to enhance human expressive capabilities, spanning original research, conceptual analyses, theoretical frameworks, perspectives, a mini-review, and a pedagogical study.We can roughly organize these contributions into four thematic areas. Neural Audio Instruments and the Body brings together work on how deep learning audio models serve as material for embodied musical practice (Donnarumma, 2025;Reus, 2026;Stefánsdóttir and Magnusson, 2025;Zappi and Tatar, 2025).Gesture, Notation, and Compositional Process focuses on how bodily movement becomes both an input to and an organizing principle for AI-mediated musical creation (Bacon et al., 2025;Kirby, 2025;Vassallo et al., 2025;Zheng et al., 2025). Embodied Cognition and Affect addresses how physiological signals and affective states mediate human-AI collaboration (Hopkins et al., 2025;Ng et al., 2025;Noufi et al., 2025;Wallace et al., 2025). Finally, Multimodality and Learning extends embodied sound AI into crossmodal, everyday, and educational contexts (Galvão and Sagesser, 2025;Rhodes et al., 2025;Riaz et al., 2026;Spanio et al., 2025).The first group of articles focuses on AI systems that reshape the relationship between sound, embodiment, and instrumental agency. Together, they show how neural audio and voice technologies can move beyond technical synthesis toward more corporeal, crossmodal, and relational forms of musical practice. Donnarumma (2025) draws on the author's artistic practice to propose embodied approaches to AI that open new forms of corporeal knowledge and new ways of designing musical AI systems. The essay examines how such systems can enable the transgression of musical and bodily boundaries and support learning beyond normative corporeal experience through a transdisciplinary analytical framework. Reus (2026) examine the live performance i: goU weI as a case study for understanding AI voices as materials, mediators, and gifts. Drawing on embodied cognition, voice studies, and material anthropology, the article analyzes real-time AI-mediated voice as a form of embodied cognition. It further discusses how AI voice transfer systems reshape vocal embodiment, agency, and the ethical responsibilities associated with data-driven voices. Stefánsdóttir and Magnusson (2025) explore how neural audio synthesis reshapes the relationship between performer and instrument through a practice-led investigation of an AI-augmented violin. Rather than acting as a passive tool, the instrument becomes a semi-autonomous collaborator, redistributing agency between human and algorithm. By emphasizing small, curated datasets and embodied interaction, the work highlights a shift toward more personal and co-creative uses of AI in artistic practice.The study offers a timely perspective on how intelligent systems can transform not only musical expression, but also broader notions of control, authorship, and human-machine interaction. Zappi and Tatar (2025) examine the integration of artificial intelligence into crossmodal sensory experiences, focusing on how data-driven models can bridge traditionally separate perceptual domains. By combining computational techniques with insights from human-media interaction, the work highlights the potential of AI to mediate and translate between different sensory modalities, enabling richer and more adaptive forms of interaction.At the same time, it underscores key challenges related to interpretability, evaluation, and user experience, pointing toward the need for more human-centered approaches in the design of multimodal intelligent systems.The second group shifts from instrumentality and embodiment toward the compositional process itself.These articles examine how gestures, notation, interfaces, and constraints can become active parts of musical thinking, especially when creative work unfolds through interaction with AI systems. Bacon et al. (2025) investigate embodied (prescriptive) and sonic (descriptive) notation as approaches to notation for a neural network-based musical instrument. Through a user study involving 11 musicians, the article explores the use of these two forms of graphic notation. It also introduces the conceptual and empirical framework "Space of Notational Strategies," which supports music composition and creative engagement with novel digital musical instruments. Kirby (2025) introduces a movement-led methodology for instrumental composition using pose estimation technology. In the Body Fragmented project, a collaboration between composer and violinist, pose estimation serves as both a notational and a collaborative tool, centering the body in the compositional process while supporting non-linear, iterative creative workflows. Vassallo et al. (2025) present NeuralConstraints, a computer-assisted composition library that integrates a feedforward neural network as a rule within a constraint-based composition framework.By combining the predictive generative capabilities of neural networks with a backtracking constraint algorithm, the tool offers composers higher-level creative control than conventional neural generation, bridging rule-based and inferential approaches within a single compositional environment. Zheng et al. (2025) explore the gestural affordances enabled by navigating the audio latent spaces of generative AI models. The paper presents a user study involving 18 musicians who used a stylus-and-tablet interface designed for latent-space navigation in a neural audio synthesis model. The participants' interactions and musical score creations reveal new gesture-based approaches to latent-space navigation informed by embodied music cognition.The third group foregrounds affect, bodily response, and social perception. Across brain sensing, physiological synchrony, vocal persona, and robotic movement, these articles investigate how AI systems may respond to, shape, or model embodied and emotional dimensions of musical interaction. Hopkins et al. (2025) introduce BrAIn Jam, a system based on functional near-infrared spectroscopy (fNIRS), used in an experimental study to monitor human drummers' brain states while collaborating with an AI-driven virtual musician. The article's key contributions include the presentation of rhythmic predictability as a quantifiable metric, the development of BrAIn Jam as an adaptive system for capturing real-time neural activity, and a discussion of the challenges and opportunities of brain-computer interfaces for communication between human and AI-driven musicians. Ng et al. (2025) The final group broadens the discussion toward multimodal interaction, learning, and everyday environments. These articles show how AI can support musical and sonic engagement not only through performance systems, but also through pedagogy, public spaces, multisensory translation, and accessible music education. Galvão and Sagesser (2025) present a pedagogical approach in which ChatGPT serves as an interlocutor during collaborative soundwalk script creation. Conducted with four undergraduate students, the experiment shows how AI can facilitate collective auditory imagination and embodied listening practice, while students critically assessed AI's strengths in streamlining collaboration and its limitations in emotional depth. Rhodes et al. (2025) present an AI guitar assistant designed to improve accessibility in UK music education. The system is evaluated through a survey of 21 guitarists, which identifies themes that highlight both the challenges and opportunities of using AI in music education. The challenges include providing nuanced feedback, addressing the social and emotional dimensions of learning, and adapting to technological limitations. The opportunities include supporting diverse learning approaches beyond traditional systems and positioning AI as a complementary teaching tool. Riaz et al. (2026) examine indirect and inverse mappings as alternative design strategies for embodied AI systems in everyday environments.Through four musicking technologies (a reactive birdbox, a reactive painting, self-playing guitars, and the Muzziball), the article explores how sound-and music-based interactions can emerge from incidental, unconscious, or involuntary human actions. By combining rule-based and learning-based approaches with multimodal sensing and actuation, the work argues that minimalistic systems can support passive musicking, perceived agency, and subtle forms of engagement without requiring explicit musical intention or complex AI behavior. Spanio et al. (2025) explore the emerging intersection of gustatory and auditory experience through generative AI, proposing a novel framework for integrating taste and sound into a unified multimodal expression. By leveraging data-driven models to translate sensory features across modalities, the work envisions a "multimodal symphony" in which flavor profiles and sonic structures coevolve in real time. Beyond its artistic implications, the study raises important questions about crossmodal perception, authorship, and the role of AI in shaping embodied experience. It offers a forward-looking perspective on how generative systems can expand the design space of multisensory interaction, with potential applications spanning creative practice, immersive media, and sensory augmentation.When the call for the Embodied Perspectives on Musical AI (EmAI) workshop was first circulated in the summer of 2022, the public landscape of AI was quite different. This was only a few months before the launch of ChatGPT and before the rapid rise of large language models and generative AI across creative domains. Since then, the state of the art in generating text, images, sound, and other media has advanced dramatically. Although multimodality has become more visible since then, the embodied realm of cognition is still often treated as secondary to representation, prediction, and content generation. This also means that the recurring question "Can AI be creative?" may be less useful than it first appears. In musical and sonic practices, creativity is rarely confined to the final output. It emerges through situated relations among attention, action, timing, gesture, posture, gaze, touch, voice, listening, and shared cultural cues. From this perspective, AI is most compelling not as a rival creator, but as a multimodal partner that helps shape what is attended to, selected, measured, and fed back into the creative loop. This orientation also brings ethical responsibilities. If collaboration is reduced to optimization and narrow benchmarks, AI systems may amplify what is easy to measure while making slower and less visible forms of influence harder to contest, including the circulation of moods, norms, and behaviors through social and technological networks. The contributions in this special issue point toward a different path: embodied perspectives on AI as a means of making relations among bodies, sounds, systems, and environments more perceptible, negotiable, and accountable.