A machine learning framework for ERA5-PWV bias calibration
Kai Li, Yun Xie, Jie Tang, De-rui Luo, Junjie Xu, Hanquan Cheng, et al.
4 papers indexed
Kai Li, Yun Xie, Jie Tang, De-rui Luo, Junjie Xu, Hanquan Cheng, et al.
Zhenyu Hou, Yujiang Li, Jie Tang, Yuxiao Dong
Reinforcement learning (RL) is becoming increasingly important for post-training large language models (LLMs). Previous RL pipelines for LLMs were mostly synchronous and batch-interleaved, which is inefficient for long-horizon agentic tasks. Recently, asynchronous RL has emerged…
Yujiang Li, Zhenyu Hou, Yi Jing, Jie Tang, Yuxiao Dong
Long-horizon agentic LLMs are increasingly limited by finite context windows, as extended interaction trajectories can exceed the maximum context length before a task is completed. Context compaction offers a natural solution by summarizing previous interaction states and continu…
Dijia Zhan, Xuemiao Xu, Jinyi Li, Jie Tang
Vision-Language-Action (VLA) policies that execute fixed-length action chunks can exhibit multimodal bifurcation: a cross-chunk inconsistency in which adjacent chunks generated from independent Gaussian latents can converge to incompatible trajectory modes, producing abrupt disco…