StateM: Reaching 95.3% Raw Accuracy, or a \$15 Frontier Run, on Terminal-Bench 2.1 via Harness Scaling
Long-horizon agents can fail even when their underlying models can solve the constituent steps. They may lose track o...
每天自动聚合 AI 领域最新动态
Long-horizon agents can fail even when their underlying models can solve the constituent steps. They may lose track o...
Multimodal large language models increasingly use visual chain-of-thought (Visual CoT) to reason about spatial, tempo...
Multiple-choice benchmarks are widely used to evaluate large language models, but MCQ scores conflate knowledge with ...
Frontier large language models (LLMs) safety evaluation has largely treated harmful generation as an attack outcome r...
We study how teams of AI coding agents coordinate while solving programming tasks. Current evaluations usually report...
Sign language serves as a vital means of communication for individuals with hearing impairments, yet recognition reso...
Large vision-language models can recognize the objects and attributes in a crowded scene yet assign an attribute to t...
Large Language Models (LLMs) have demonstrated capabilities in in-context learning, task decomposition, step-by-step ...
Agents now write knowledge graphs, but knowledge-graph stores still carry defaults set when humans curated them: acce...
Video world models approximate the stochastic distribution of physical outcomes through generative sampling, but exis...
Generative pretraining established reusable task representations; later work on language-based task conditioning and ...
We demonstrate that AI models are broadly susceptible to a phenomenon we call model hypnosis, in which individually w...