Multimodal Large Language Models (MLLMs) excel in diverse vision tasks, but full-parameter retraining is computationally expensive as real-world knowledge evolves. Existing continual learning methods often suffer from semantic entanglement in parameter spaces across tasks, impedi…
Background. The widespread deployment of ambient digital scribes is driving large-scale capture of clinician-patient dialogues. Human coding of clinical communication data remains costly, inconsistent, and difficult to scale, motivating AI-driven communication coding systems. How…
Visual recognition and localization of underwater optical beacons are critical for AUV docking, but traditional beacons are limited by fixed directionality and light attenuation in water. To extend the range of optical docking, this study designs a novel omnidirectional rotating…