Understanding Evolution Strategies for LLM Reasoning: Broader Reasoning Coverage than GRPO
Evolution Strategies (ES) have recently emerged as a memory-efficient post-training paradigm for LLM reasoning. Howev...
每天自动聚合 AI 领域最新动态
Evolution Strategies (ES) have recently emerged as a memory-efficient post-training paradigm for LLM reasoning. Howev...
On-policy distillation (OPD), which leverages a pre-trained, specialized teacher model to provide dense supervisory s...
Agent skills package specialized knowledge and workflows into reusable resources that extend AI agent capabilities. R...
Zero-shot cross-task generalization, where a policy must execute manipulation tasks never seen during training, remai...
Multimodal large language models (MLLMs) can interpret a street view, but urban agency depends on whether such local ...
Recent video generation models are increasingly framed as world models. Many physical processes can unfold in more th...
Large-scale vision-language models (VLMs) have demonstrated remarkable versatility across a wide range of multimodal ...
AI-powered digital avatar streamers must answer product questions, engage viewers, and execute marketing strategies i...
Reusable skill libraries allow large language model (LLM) agents to reuse procedural knowledge across tasks, but they...
Explicit visual intermediates can help multimodal large language models (MLLMs) externalize spatial evidence and upda...
Recent prominent post-training methods, such as Reinforcement Learning (RL) and On-Policy Self-Distillation (OPSD), h...
Modern video games combine first-person perception, rapid visual changes, persistent world state, and heterogeneous n...