GUI vs. CLI: Execution Bottlenecks in Screen-Only and Skill-Mediated Computer-Use Agents
Computer-use agents can execute software tasks through either graphical interfaces or programmatic command interfaces...
每天自动聚合 AI 领域最新动态
Computer-use agents can execute software tasks through either graphical interfaces or programmatic command interfaces...
A classical intuition holds that verifying a solution is easier than producing one. For today's coding agents, this i...
Modern generative world models render increasingly realistic action-controllable futures, yet they frequently halluci...
Scientific reasoning models for biology combine language models with foundation models trained on multimodal biologic...
Speculative decoding (SD) accelerates autoregressive Large Language Models (LLMs) by drafting multiple tokens and ver...
Despite their widespread use, the role of reward models in shaping reinforcement learning is poorly understood. Rewar...
As LLM agents become capable of increasingly long-horizon tasks, evaluating their performance in economic systems is ...
On-policy distillation (OPD) improves LLM reasoning by training a student model on its own generated outputs, but sta...
Continual Test-Time Adaptation (CTTA) aims to maintain model performance under evolving target domains by adapting on...
Jailbreak attacks reveal a persistent weakness in aligned Large Language Models: carefully crafted prompts can elicit...
AI agents acting on behalf of users are constantly making decisions, and for users to trust their agents, those decis...
Recent advances in stereo matching have achieved remarkable accuracy, but often rely on large models, heavy computati...