Vision foundation models are increasingly reused as frozen backbones for downstream visual recognition, making parameter-efficient adaptation a central problem. Prompt-based adaptation, including Visual Prompt Tuning (VPT), provides a lightweight way to specialize these models, b…
Background Grounded in self-determination theory, basic psychological needs are intrinsic factors that regulate employees' occupational mental health. Pediatric nurses face heavy work pressure and high levels of emotional labor, leading to prevalent occupational burnout. This stu…
Vision-Language-Action (VLA) models augmented with world modeling represent a promising paradigm for end-to-end autonomous driving. While pixel-level future prediction enables fine-grained spatiotemporal reasoning, it compromises robustness in noisy driving scenarios. Conversely,…
UV seam placement is a critical yet labor-intensive step in 3D content creation, requiring artists to balance chart shape, seam concealment, and alignment with semantic and geometric features. Existing automatic methods are primarily based on per-object optimization, relying on h…
We present KAT-Coder-V2.5, a coding-focused agentic model trained to act autonomously inside real, executable repositories rather than as a single-turn code generator. Its capability is bottlenecked less by model scale than by the scarcity of reproducible environments, verifiable…
We introduce Gemma 4, a new generation of open-weight, natively multimodal language models in the Gemma model family. Designed to advance compute efficiency and reasoning, the Gemma 4 model suite features dense and Mixture-of-Experts architectures, ranging from 2.3B to 31B parame…
Dual-pixel (DP) imaging enables metric depth estimation from a single camera using sub-aperture disparity. However, the extremely small effective baseline limits disparity observability, leading to structural degradation and depth failure in textureless, low-contrast, or downsamp…
Discovering governing equations directly from observational data is a key step towards interpretable scientific machine learning. Current data-driven approaches typically operate on a single dataset, inherently limiting their performance when faced with restricted observations. I…
This article surveys spatial-domain-enhanced Physical-layer Authentication (PLA), with Dual-polarized Antennas (DPA), Massive Multiple-Input Multiple-Output (MIMO), and Reconfigurable Intelligent Surfaces (RIS) as the primary focus. With the rapid growth of wireless deployments,…
This paper addresses the lack of explicit memory mechanisms in current object detection models and proposes Hippocampus-DETR, a novel detection framework based on biological hippocampal memory modeling. This framework integrates a hippocampal memory network module, HipNet, into t…
The digital economy plays a pivotal role in advancing green productivity; however, the specific configurations driving this relationship remain underexplored. Employing the TOE theoretical framework alongside k-means clustering and fuzzy-set qualitative comparative analysis (fsQC…