CORTEXA
← Browse

Harry Yang

6 papers indexed

arxivcs.CV2026-07-17

Distributional Matching for Vector Quantization: A Unified Theoretical and Empirical Framework

Xianghong Fang, Litao Guo, Hengchao Chen, Yuxuan Zhang, XiaofanXia, Dingjie Song, et al.

The effectiveness of modern visual representation learning and autoregressive models critically depends on vector quantization (VQ), which discretizes continuous feature representations using a learnable codebook. Despite its widespread use, existing VQ methods often suffer from…

View free PDFSource page
arxivcs.AI2026-07-09

Game Theory Driven Multi-Agent Framework Mitigates Language Model Hallucination

Runzhe Liu, Biquan Bie, Zihao Wang, Yuchao Ma, Yexin Liu, Xinghai Li, et al.

The application of lightweight Large Language Models in rule-based scientific domains remains severely limited by their tendency to mimic linguistic patterns rather than reproduce axiomatic reasoning, causing frequent hallucinations. Here, we show that G-Frame, an adaptive multi-…

View free PDFSource page
arxivcs.AI2026-07-05

Do GUI Agents Believe Their Eyes? Diagnosing State-Belief Reliance on Pixels versus Structure

Guijia Zhang, Yuxun Chen, Yuheng Qi, Harry Yang

Multimodal GUI agents read an interface through two redundant channels: the rendered pixels of a screenshot and a serialized structure such as a document object model or accessibility tree. Before acting, an agent forms a belief about the current interface state, but existing ben…

View free PDFSource page
arxivcs.CV2026-07-01

GKDT: General Keypoint Detection Transformer

Changsheng Lu, Yuxin Chen, Haokun Gui, Rong Wang, Jie Yang, Harry Yang, et al.

With the emergence of various pre-trained vision and language models, computer vision is shifting from narrow-domain to open-domain recognition. The construction of a more powerful yet general keypoint detection (GKD) model to support diverse tasks has become increasingly importa…

View free PDFSource page
arxivcs.CV2026-07-01

RotateAttention: RoPE-Aware Rotation and Range Rectification for INT4 Quantized Attention in Video Generation

Yaofu Liu, Wanli Lan, Jinxi Li, Binhang Yuan, Harry Yang

In $\textbf{DiT-based video generation models equipped with 3D Rotary Position Embeddings (3D RoPE)}$, the attention mechanism remains a primary computational bottleneck due to its quadratic complexity with respect to sequence length. While quantized $\textbf{FlashAttention}$ off…

View free PDFSource page