STARE: Surprisal-Guided Token-Level Advantage Reweighting for Policy Entropy Stability
Reinforcement Learning with Verifiable Rewards algorithms like GRPO have emerged as the dominant post-training paradi...
每天自动聚合 AI 领域最新动态
Reinforcement Learning with Verifiable Rewards algorithms like GRPO have emerged as the dominant post-training paradi...
The Network Data Analytics Function (NWDAF) is central to enabling zero-touch network management in fifth-generation ...
Predictive code completion greatly accelerates how quickly developers work. In spreadsheets, despite being much more ...
Multi-turn tool-use RL is bottlenecked by the rapid depletion of informative samples in static datasets. We observe t...
Current benchmarks for computer-use agents evaluate models in impersonal environments. This leaves a gap between eval...
A useful phone agent needs to be personally intelligent. It should reason over a user's identity, history, and prefer...
We show the standard basis of transformer hidden states already provides a training-free, architecture-general featur...
Vision Transformers (ViTs) have become a dominant architecture for visual representation learning, providing exceptio...
Motion forecasting is central to visual intelligence: agents must anticipate how objects will move in order to plan a...
On-policy self-distillation (OPSD) trains a model on its own rollouts and uses a frozen copy to provide dense token-l...
As an increasing majority of global video content is consumed on social platforms for interactive social purposes, vi...
Score- and flow-matching models often rely on preference-based reinforcement learning for two purposes: aligning with...