Optical neural networks (ONNs) promise ultra-fast and energy-efficient computing but are hampered by the critical challenge of on-chip training. Here, we propose an on-chip training distillation-guided optical neural network (DGONN) and introduce a forward distilled algorithm to…
Large Language Models (LLMs) often struggle to navigate value conflicts when trained with the compressed scalar rewards of Reinforcement Learning from Human Feedback (RLHF). To address this challenge, we investigate how chain-of-thought (CoT) reasoning can help improve performanc…