Rethinking the Evaluation of Harness Evolution for Agents
We revisit the evaluation of automatic harness evolution for LLM agents. Existing harness evolution methods use unit ...
每天自动聚合 AI 领域最新动态
We revisit the evaluation of automatic harness evolution for LLM agents. Existing harness evolution methods use unit ...
We present a novel viewpoint for uncertainty quantification. Uncertainty measures are not primitives, in need of axio...
Reinforcement learning has proven effective for improving reasoning in large language models, but extending it to Mas...
Annotation quality is a major bottleneck in building reliable and explainable artificial intelligence (XAI) systems f...
Real repository issues routinely include visual evidence such as screenshots, error dialogs, rendered UI states, and ...
Multimodal large language models (MLLMs) often introduce errors when generating image captions, resulting in misalign...
Large language models (LLMs) increasingly serve as high-level planners for embodied agents, where linguistically beni...
A tokenizer fixed at the start of pre-training allocates vocabulary in proportion to the pre-training corpus, reflect...
Evidence synthesis is crucial for turning primary research into reliable knowledge for science, medicine, education, ...
Traffic agencies now have access to large volumes of video-derived data for studying safety and congestion. Most of t...
Recent advances in Tool-Integrated Large Language Models have made web search a core capability of information-seekin...
Security-agent evaluations commonly measure peak offensive capability under generous inference budgets, emphasizing v...