High-resolution (HR) image perception remains a key challenge in multimodal large language models (MLLMs), as fine-grained details are often lost when the image is processed as a whole. Existing methods either require training to teach models where to look or heuristically divide…
Autoregressive (AR) streaming models have emerged as a powerful paradigm for long video generation. However, the linearly growing Key-Value (KV) cache poses a significant bottleneck, leading to memory overload and degraded inference throughput. A common compression method is to d…
This study combines machine vision technology and deep learning models to rapidly assess the activity of anaerobic ammonium oxidation (Anammox) granular sludge. As a highly efficient nitrogen removal technology for wastewater treatment, the Anammox process has been widely applied…