DOPD: Dual On-policy Distillation
On-policy distillation (OPD) offers superior capacity transfer by supervising student-sampled trajectories with dense...
每天自动聚合 AI 领域最新动态
On-policy distillation (OPD) offers superior capacity transfer by supervising student-sampled trajectories with dense...
Conservative offline training is widely advocated as a safe foundation for subsequent online adaptation: if a policy ...
Can the robot use a plate to cut a cake if no knife is available? Tool use greatly expands robot capabilities, but to...
World models offer a principled way to equip long-horizon LLM agents with foresight: predictions of action consequenc...
Full-length song generation must preserve coherence and musicality, render detailed vocal and accompaniment acoustics...
Perception-based humanoid loco-manipulation requires connecting egocentric observations and task instructions to whol...
Self-collision remains a persistent challenge in SMPL-based human pose estimation and motion generation. Under extrem...
MLLM-based GUI grounding methods commonly formulate target localization as autoregressive coordinate generation, enab...
In real-world applications, guardrails are often expected to identify unsafe user-model interactions according to app...
Interactive video generation systems for camera-controlled world exploration roll out growing sequences of latent vid...
Existing computer-use benchmarks fail to capture the realism, complexity, and long-horizon demands of real-world comp...
As large language models and harness frameworks continue to advance, agents operating in terminals are increasingly c...