arXiv

Desktop-Delta Bench: Do Computer-Use Models Understand Desktop GUI Transitions?

Computer-use agents (CUAs) increasingly act through desktop GUIs to complete long-horizon tasks. Current benchmarks p...

AI 聚合
2026-07-29
arXiv

$π\mathbf{R}^2$: Reactive Real-time Flow Policies

Generalist manipulation policies increasingly take the form of action-chunking flow policies built on large pretraine...

AI 聚合
2026-07-29
arXiv

Pass the Baton: Trajectory-Relayed On-Policy Distillation

On-policy distillation (OPD) grounds token-level supervision in the student's own trajectory, yet suffers from prefix...

AI 聚合
2026-07-29
HuggingFace

Towards Robust Reinforcement Learning for Small-Scale Language Model Agents

The alignment of Small Language Models (SLMs) in the 70--500M parameter range using reinforcement learning is often c...

AI 聚合
2026-07-29
HuggingFace

Shieldstral

We introduce Shieldstral, a 3B-parameter policy-adaptive multimodal safety classifier that matches or outperforms mod...

AI 聚合
2026-07-29
HuggingFace

Wonder: Video World Model Done Better

We present Wonder, a general-purpose video world model for real-time, camera-controllable world exploration. Given an...

AI 聚合
2026-07-29
HuggingFace

OmniDelta: Skill-Driven Budget Allocation for Token Compression in OmniLLMs

Emerging Omni-modal Large Language Models (OmniLLMs) enable unified understanding of text, audio, and video, but thei...

AI 聚合
2026-07-29
HuggingFace

Visual prompt engineering for video models

In the age of foundation models, a model is only as good as its prompt. For this reason, prompt engineering has becom...

AI 聚合
2026-07-29
HuggingFace

Keep It InMind: Benchmarking the Implicit-Association Blind Spot in Agent Memory

Long-term memory systems store what a user says in an external store and retrieve it when a related query arrives. Th...

AI 聚合
2026-07-29
HuggingFace

PerceptionBench: Evaluating Atomic Visual Perception in Multimodal Large Language Models

We introduce PerceptionBench, a benchmark specifically designed to evaluate the atomic visual perception capabilities...

AI 聚合
2026-07-29
HuggingFace

HiFi-UMI: Learning Deployable Manipulation Policies from High-Fidelity UMI Data Alone

Learning deployable manipulation policies is bottlenecked by the scarcity of data that is both high-fidelity and scal...

AI 聚合
2026-07-29
HuggingFace

A New Role for Relevance: Guiding Corpus Interaction in Agentic Search

Relevance is a query-dependent estimate of whether a document or excerpt contains useful evidence. Existing retrieval...

AI 聚合
2026-07-29
首页 上一页 第 101 / 216 页 下一页 末页