AgentBeats: Agentifying Agent Assessment for Openness, Standardization, and Reproducibility
Agent systems are advancing quickly across domains, but their evaluation remains fragmented. Most benchmarks rely on ...
每天自动聚合 AI 领域最新动态
Agent systems are advancing quickly across domains, but their evaluation remains fragmented. Most benchmarks rely on ...
Search-augmented LLMs increasingly mediate everyday consumer recommendations by retrieving live web content. This cre...
Shielded reinforcement learning is typically presented as a runtime safety mechanism that compiles temporal-logic spe...
There is a proliferation of work arguing for the use of synthetic data in scientific research. For example, social sc...
We introduce SkMTEB, the first comprehensive MTEB-style text embedding benchmark for Slovak, a low-resource West Slav...
This paper examines three recent frameworks for understanding the cognitive and epistemic consequences of artificial ...
LLM-based agents have shown increasing potential in automating scientific discovery. Given an optimizable metric and ...
Current LLM-based research agents have advanced through agent orchestration, yet largely overlook scientific knowledg...
Reproducibility in the social and behavioral sciences is typically evaluated by independent researchers who reanalyze...
Spatial reasoning, the ability to determine where objects are, how they relate, and how they move in 3D, remains a fu...
Articulated tool manipulation remains a major challenge in dexterous robotics due to the need to coordinate internal ...
Retrieval-augmented generation (RAG) has become a standard mechanism for grounding language models in external knowle...