How to Train a Critic Stably and Efficiently
Group-based reinforcement learning methods such as GRPO for large language models avoid training a critic by sampling...
每天自动聚合 AI 领域最新动态
Group-based reinforcement learning methods such as GRPO for large language models avoid training a critic by sampling...
On-policy distillation (OPD) has emerged as an effective framework for post-training language models by pairing stude...
Language models are sequential processors, but long-horizon agency requires external information and computation beyo...
An interactive world model must follow the user's actions, remember the places it has shown, and stream in real time....
Existing image editing frameworks predominantly follow the training paradigm of text-to-image diffusion models. Howev...
Recent large language models (LLMs) can operate as coding agents that build complete games from natural language requ...
Open-ended real-world interaction admits multiple valid behaviors: an agent may answer directly, ask for clarificatio...
We present a novel approach to efficient LLM harness optimization through adaptive validation task selection. Harness...
Speech-based applications pass spoken queries through automatic speech recognition (ASR) before any retrieval module,...
As on-device LLM agents evolve into personal copilots, the mobile operating system has become a key testbed for this ...
While text-to-3D generation has advanced rapidly, achieving high geometric fidelity at low inference cost remains cha...
Search-augmented LLMs increasingly mediate everyday consumer recommendations by retrieving live web content. This cre...