E-Commerce Bench: Evaluating LLM Agents on Long-Horizon Autonomous Business Operation
Long-horizon agentic tasks go beyond chaining short tasks over more interaction turns. Their evolving dynamic environ...
每天自动聚合 AI 领域最新动态
Long-horizon agentic tasks go beyond chaining short tasks over more interaction turns. Their evolving dynamic environ...
Robotic manipulation faces a fundamental scaling challenge: robust generalization demands broad physical experience, ...
We present Qwen-Drive-1.0, an initial step towards a vision-language foundation model for autonomous driving. Qwen-Dr...
Multimodal GUI agents have emerged as a promising paradigm for digital task automation, yet transitioning from benchm...
This paper studies autonomous software development, in which LLM-based coding agents transform high-level requirement...
Looped Transformers increase effective depth by iterating a shared block of layers, but most evaluations compare at f...
Long-horizon complex tasks require foundation models to accumulate information, maintain internal states, and adapt o...
Self-play is an effective paradigm for language-model self-evolution, but without guidance, solver performance can pl...
Cognitive language agents have achieved substantial progress by equipping language models with memory, tools, and dec...
AI is increasingly used in the R\&D process that produces future AI systems. We study the conditions under which this...
End-to-end weather forecasting systems produce skillful global gridded and station forecasts directly from raw Earth ...
Large language models can generate interactive web interfaces, but reliable generative UI requires maintaining an exe...