Ignorance or Incompetence? Constructing Knowledge-Gated, Verifiable Tasks for LLM Agents
Professional agent tasks often depend on conventions that are absent from public corpora, yet benchmarks rarely contr...
每天自动聚合 AI 领域最新动态
Professional agent tasks often depend on conventions that are absent from public corpora, yet benchmarks rarely contr...
Data-residency constraints force enterprises to self-host LLMs, but continuous adoption of newer models without decom...
Multimodal geometry reasoning requires VLMs to extract precise visual relations and preserve them through multi-step ...
Commercial short-drama production follows a multi-stage chain: script, storyboard, keyframe imagery, shot-level video...
Test-time adaptation (TTA) typically assumes that model parameters can be updated at inference time. This assumption ...
Valuable data remains embedded in unstructured sources: web pages, reports, contracts, filings, earnings calls, and P...
LLM judges are widely used to evaluate agentic tool-calling systems, yet their reliability on structured, dependency-...
Logit-based knowledge distillation (KD) is used to train smaller language models (LMs) via supervision from stronger ...
Dexterous manipulation policies learned by imitation are typically evaluated for robustness to variation in scenes, o...
Long-running LLM agents rely on persistent memory to carry state across interactions, including permissions, restrict...
Coding agents are now commonly evaluated on the SWE-bench family of benchmarks, whose tasks are built from curated Gi...
This paper explores whether large language models (LLMs) can design video coding tools, a highly challenging task due...