LLM preference alignment aims to optimize models toward human preferences across diverse user instructions. Reinforcement learning has become a major post-training approach for this goal, but existing proxy rewards are often outcome-level, mainly evaluating the final response whi…
This paper investigates the ranging performance of a bistatic integrated sensing and communications (ISAC) system employing orthogonal frequency-division multiplexing (OFDM), in which an ISAC transmitter emits a communication waveform carrying random data symbols, and a separate…
This paper investigates the statistical ambiguity functions (AFs) of orthogonal frequency division multiplexing (OFDM) waveforms that incorporate deterministic unit-modulus pilot symbols and random data payloads for integrated sensing and communication (ISAC). We derive analytica…
In this paper, we propose a novel tamed stochastic gradient Hamiltonian Monte Carlo (tSGHMC) algorithm for sampling and stochastic optimization problems with superlinearly growing stochastic gradients. Under a certain continuity in average condition and a strong convexity conditi…
Black-box modeling of inverter-based resources (IBRs) has attracted growing interest for real-time grid operation and control in the presence of proprietary electronic control architectures. Existing machine learning (ML)-based online dynamic trajectory prediction approaches usin…
Reinforcement Learning like Group Relative Policy Optimization (GRPO) has significantly advanced text-to-image post-training. However, current methods often favor superficial aesthetics, such as over-saturated colors, leaving critical flaws like AI artifacts and biological implau…