arxiveess.AScs.AI2026-07-02
An Efficient vLLM-Based Inference Pipeline for Unified Audio Understanding and Generation
Haoran Wang, Jinchuan Tian, Siddhant Arora, Shinji Watanabe
While Large Multimodal Models excel in comprehension, high-throughput inference engines lack native support for multimodal generation. This is severe in Speech Language Models, where generating multi-layered audio tokens via decoupled AR+NAR or synchronous Multi-Token Prediction…