arXiv

Benchmarking the Benchmarks: Evaluating Benchmarks for Conversational Agents

Task-oriented conversational agents are evaluated using curated or automatically generated benchmarks, yet benchmark ...

AI 聚合
2026-08-07
arXiv

Tytan: Interactive Neurosymbolic Construction of Analytic Semantic Schemas from Relational Data

From natural-language query interfaces to automated report generation, data analysis tools need a description of the ...

AI 聚合
2026-08-07
arXiv

TRAJDEBUG: Tracing Error Lifecycle to Identify Critical Failures in Long-Horizon Agent Trajectories

LLM-based agentic systems have shown remarkable capabilities in complex domains, while suffering from cascading error...

AI 聚合
2026-08-07
arXiv

Challenges in Evaluating Explanation Methods for Static and Evolving Data

This paper addresses the limitations of Explainable Artificial Intelligence (XAI) with respect to insufficient evalua...

AI 聚合
2026-08-07
arXiv

Resourced Authority A Mechanism-Design Model for Participatory Governance of Deployed AI Agents

We give a formal mechanism design model for the continuous participatory governance of a deployed AI agent. The mecha...

AI 聚合
2026-08-07
arXiv

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping

Real-world video benchmarks provide broad coverage, but their fixed clips entangle event count, rate, duration, and v...

AI 聚合
2026-08-07
arXiv

AV-AIVAT: 74x Cheaper Agent Evaluation with Certified Anytime-Valid Stopping in Imperfect-Information Games

Deciding which of two agents is stronger means playing games until skill outweighs luck, and every game costs money, ...

AI 聚合
2026-08-07
arXiv

An Optimal Agnostic PAC Algorithm

Let $H\subseteq\{-1,+1\}^X$ be a class of finite VC dimension $d\ge1$. Writing $L$ for the binary risk and $L^*=\min_...

AI 聚合
2026-08-07
arXiv

Investigating Artificial Intelligence Digital Sovereignty in Mobile Shopping Apps: A Case Study of Nigeria

The use of e-commerce mobile applications is expanding in Nigeria, creating both opportunities and risks, including f...

AI 聚合
2026-08-07
arXiv

Tracing the Heart: An Evidence-Linked Pipeline for Heart-Failure Feature Engineering

Electronic health record (EHR) feature engineering is a major bottleneck in clinical research and AI, accounting for ...

AI 聚合
2026-08-07
arXiv

Learning When to Trust via Selective Context Preference Optimization

Language models increasingly condition their answers on external signals, and a single misleading one can turn a corr...

AI 聚合
2026-08-07
HuggingFace

ChronoVision: Temporal Reasoning via Latent State Reconstruction

Multimodal large language models excel at passive perception but struggle with complex visual cognitive tasks requiri...

AI 聚合
2026-08-07
首页 上一页 第 72 / 212 页 下一页 末页