CORTEXA
← Browse
arxivcs.CVcs.MMcs.SDeess.ASeess.IV2026-07-15

Bring Music The Horizon: Music-Driven 360$^\circ$ Video Generation

Kai Hsu Tsai, Yong Wei Fu, Hung I Yang, Yu-Chih Chen

Music visualization offers a powerful way to enhance listeners' understanding and experience of music by translating auditory signals into visual forms. However, most existing approaches either rely heavily on lyrics or generate flat, non-immersive videos similar to conventional music videos, which limits their ability to convey the emotional dynamics of music and provide an immersive listening experience. We propose Bring Music The Horizon, an emotion-aware pipeline for music-driven 360$^\circ$ video generation. Given an input song, our work first estimates its emotional trajectory by predicting valence-arousal values at the level of every four bars. These values are then converted into emotion-aware visual guidance using EmotiCrafter, and these guidance vectors can be manipulated by the SEGA framework, which provides fine-grained semantic control for keyframe generation. Finally, image-to-video models are applied to the generated keyframes to synthesize temporally continuous 360$^\circ$ videos for immersive music visualization. Our pipeline generates 360$^\circ$ music visualization videos that reflect the emotional progression and temporal structure of the input song. We demonstrate its capability using songs from different genres and provide qualitative comparisons with From-Sound-To-Sight, a representative audio-to-visual generation baseline, on our project page at https://etoile-et-toi-mp3.github.io/BMTH_Project_Page/.

View free PDFSource page

Related papers

arxivcs.CVcs.MMcs.SDeess.AS2026-06-29

AVTok: 1D Unified Tokenization for Holistic Audio-Video Generation

Kien T. Pham, I Chieh Chen, Qifeng Chen, Long Chen

Audio-video generation has recently gained unprecedented research attention, aiming to synthesize high-quality sounding video content with fine-grained synchronization and semantic alignment between the auditory and visual components. The preceding methods predominantly adopt a d…

View free PDFSource page
arxiveess.IVcs.AIcs.MMcs.SD2026-07-02

Data-driven Video Codec with Implicit Neural Representations

Nishan Khanal, Saugat Neupane, Abhinav Chalise, Nimesh Gopal Pradhan, Dinesh Baniya Kshatri

A conventional codec stores a video as compressed pixel data. We instead store the video, together with its audio track, as the weights of a single sinusoidal representation network (SIREN) that maps space-time coordinates to RGB values and audio amplitudes. The network uses sepa…

View free PDFSource page
arxiveess.IVcs.CVcs.MM2026-07-21

Group-of-Latents: Perceptual Video Compression at Extreme Bitrates via Masked Latent Generative Modeling

Shaokang Wang, Jinchang Xu, Peidong Jia, Zhijian Hao, Siyuan Qian, Fei Zhao, et al.

Most existing video compression algorithms follow a paradigm of transformation and quantization, optimizing the trade-off between distortion and bitrate. However, extremely low-bitrate compression remains an underexplored frontier where perceptual quality optimization under sever…

View free PDFSource page
arxivcs.CVcs.SDeess.AS2026-06-29

SyncCache: Exploiting Asymmetric Dynamics for Fast Audio-Driven Portrait Animation

Juncheng Ma, Yuxuan Du, Yanan Sun, Zhening Xing, Changlin Li, Zhenyu Tang, et al.

Diffusion Transformers (DiTs) have significantly advanced audio-driven portrait animation, but their high computational cost leads to substantial inference latency. Although training-free diffusion caching accelerates inference significant, existing methods are primarily develope…

View free PDFSource page
arxivcs.CVcs.MMcs.ROeess.IV2026-07-11

Label-Free Target-Domain Adaptation for Unconstrained Event-Image Feature Matching via Dual-Stage Distillation

Zhonghua Yi, Hao Shi, Qi Jiang, Yufan Zhang, Kailun Yang, Kaiwei Wang

Building pixel-level correspondence between event and image data is a fundamental task for multi-sensor systems. However, existing cross-modal matching methods are largely restricted by their reliance on either matching labels or strictly aligned hardware, which limits them to un…

View free PDFSource page
arxiveess.IVcs.CVcs.MM2026-07-06

CompressedVQA-AEV: Full-Reference and No-Reference Quality Assessment Models for Asymmetric Encoded Videos

Wei Sun, Xingwei Liu, Dandan Zhu, Xiangyang Zhu, Weixia Zhang, Guangtao Zhai

This report presents our solutions to the QoMEX 2026 Grand Challenge on Video Quality Assessment for Asymmetric Encoded Videos, comprising a full-reference (FR) model, CompressedVQA-AEV-FR, and a no-reference (NR) model, CompressedVQA-AEV-NR. The FR approach leverages a Swin-B ba…

View free PDFSource page