As large multimodal models move from understanding content to operating on digital environments, mobile GUI has emerged as a challenging and consequential testbed for digital embodied intelligence. Mobile agents operate under three coupled constraints: precise perception of compl…
We present HunyuanOCR-1.5, a lightweight end-to-end OCR-specialized vision-language model. HunyuanOCR unifies document parsing, text spotting, information extraction, text-image translation, and multi-image document understanding within a single end-to-end VLM. Building upon the…
In the context of interdisciplinary integration between advanced manufacturing and artificial intelligence, the optimization of complex part machining parameters has become a key challenge in intelligent production systems. Combining mechanical engineering, data science, and cont…