Desktop-Delta Bench: Do Computer-Use Models Understand Desktop GUI Transitions?
Computer-use agents (CUAs) increasingly act through desktop GUIs to complete long-horizon tasks. Current benchmarks p...
每天自动聚合 AI 领域最新动态
Computer-use agents (CUAs) increasingly act through desktop GUIs to complete long-horizon tasks. Current benchmarks p...
Generalist manipulation policies increasingly take the form of action-chunking flow policies built on large pretraine...
On-policy distillation (OPD) grounds token-level supervision in the student's own trajectory, yet suffers from prefix...
The alignment of Small Language Models (SLMs) in the 70--500M parameter range using reinforcement learning is often c...
We introduce Shieldstral, a 3B-parameter policy-adaptive multimodal safety classifier that matches or outperforms mod...
We present Wonder, a general-purpose video world model for real-time, camera-controllable world exploration. Given an...
Emerging Omni-modal Large Language Models (OmniLLMs) enable unified understanding of text, audio, and video, but thei...
In the age of foundation models, a model is only as good as its prompt. For this reason, prompt engineering has becom...
Long-term memory systems store what a user says in an external store and retrieve it when a related query arrives. Th...
We introduce PerceptionBench, a benchmark specifically designed to evaluate the atomic visual perception capabilities...
Learning deployable manipulation policies is bottlenecked by the scarcity of data that is both high-fidelity and scal...
Relevance is a query-dependent estimate of whether a document or excerpt contains useful evidence. Existing retrieval...