AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities
As Large Language Models (LLMs) evolve into autonomous agents, the need for unified evaluation infrastructure becomes...
每天自动聚合 AI 领域最新动态
As Large Language Models (LLMs) evolve into autonomous agents, the need for unified evaluation infrastructure becomes...
Large Language Model (LLM) agents have moved beyond generating responses to executing multi-step tasks by calling too...
Vision Transformers (ViTs) are known to exhibit high-norm patch-token outliers that degrade feature map quality, a pr...
Modern AI models achieve strong performance on many established benchmarks, yet they still fail on tasks that humans ...
Starting from the utilization of deep neural networks to approximate the state-action value function that led to winn...
Large language model (LLM) agents are beginning to automate machine learning engineering (MLE) by coupling planning, ...
Existing methods for automatic music transcription are often limited to single-instrument recordings or fail on compl...
This work introduces a unified formulation for vision models, where diverse forms of visual information beyond natura...
Mainstream visual encoders are pretrained on natural images and cannot be effectively applied to document images with...
Visual generators excel at rendering, but they confidently fabricate what they do not know. User requests are unbound...
Vision language models (VLMs) have achieved strong performance on visual document understanding benchmarks such as Do...
Large language models fine-tuned for forecasting can be accurate yet poorly calibrated, and their chain-of-thought (C...