arXiv

DOPD: Dual On-policy Distillation

On-policy distillation (OPD) offers superior capacity transfer by supervising student-sampled trajectories with dense...

AI 聚合
2026-06-30
arXiv

Pessimism's Paradox: Conservative Offline Training Amplifies Reward Hacking During Online Adaptation in Reasoning Models

Conservative offline training is widely advocated as a safe foundation for subsequent online adaptation: if a policy ...

AI 聚合
2026-06-30
arXiv

GROW$^2$: Grounding Which and Where for Robot Tool Use

Can the robot use a plate to cut a cake if no knife is available? Tool use greatly expands robot capabilities, but to...

AI 聚合
2026-06-30
arXiv

Self-Evolving World Models for LLM Agent Planning

World models offer a principled way to equip long-horizon LLM agents with foresight: predictions of action consequenc...

AI 聚合
2026-06-30
arXiv

LeVo 2: Stable and Melodious Song Generation via Hierarchical Representation Modeling and Progressive Post-Training

Full-length song generation must preserve coherence and musicality, render detailed vocal and accompaniment acoustics...

AI 聚合
2026-06-30
arXiv

VLK: Learning Humanoid Loco-Manipulation from Synthetic Interactions in Reconstructed Scenes

Perception-based humanoid loco-manipulation requires connecting egocentric observations and task instructions to whol...

AI 聚合
2026-06-30
HuggingFace

PoseShield: Neural Collision Fields for Human Self-Collision Resolution

Self-collision remains a persistent challenge in SMPL-based human pose estimation and motion generation. Under extrem...

AI 聚合
2026-06-30
HuggingFace

One Forward Beats Two: InnerZoom for Accurate and Efficient GUI Grounding

MLLM-based GUI grounding methods commonly formulate target localization as autoregressive coordinate generation, enab...

AI 聚合
2026-06-30
HuggingFace

SafePyramid: A Hierarchical Benchmark for In-context Policy Guardrailing

In real-world applications, guardrails are often expected to identify unsafe user-model interactions according to app...

AI 聚合
2026-06-30
HuggingFace

Walking in the Implicit: Interactive World Exploration via Neural Scene Representation

Interactive video generation systems for camera-controlled world exploration roll out growing sequences of latent vid...

AI 聚合
2026-06-30
HuggingFace

OSWorld2.0: Benchmarking Computer Use Agents on Long-Horizon Real-World Tasks

Existing computer-use benchmarks fail to capture the realism, complexity, and long-horizon demands of real-world comp...

AI 聚合
2026-06-30
HuggingFace

TUA-Bench: A Benchmark for General-Purpose Terminal-Use Agents

As large language models and harness frameworks continue to advance, agents operating in terminals are increasingly c...

AI 聚合
2026-06-30
首页 上一页 第 167 / 215 页 下一页 末页