SciFigQual-Bench: A Benchmark for Scientific Figure Quality Assessment with Full-Manuscript Context
Scientific images are the core elements of presenting experimental conclusions, elaborating system architecture, and ...
每天自动聚合 AI 领域最新动态
Scientific images are the core elements of presenting experimental conclusions, elaborating system architecture, and ...
With the development of audio large language models (AudioLLMs), audio captioning needs to move from brief descriptio...
Ontology matching (OM) has traditionally been formulated as either equivalence discovery or subsumption matching. The...
Writing and communication are increasingly mediated by large language models (LLMs) that are being used to draft, rev...
Vision-language-action (VLA) models remain constrained by scarce action-labeled robot data, whereas action-free video...
High-stakes decision systems in credit scoring, fraud detection, healthcare, and industrial safety require reliable u...
CT vision-language foundation models have demonstrated promising performance across downstream tasks, but are typical...
Large language model (LLM) agents are increasingly expected to assist users in completing tasks. However, existing be...
Traditional search systems are optimized to retrieve items that strictly match a query, often prioritizing precision ...
Effective collaboration with novel and diverse partners is a crucial skill for autonomous agents. Most current ad-hoc...
Conversational AI is increasingly positioned as a teammate rather than a tool, yet we know little about how its prese...
We introduce APEX-Accounting, a benchmark built by Mercor in partnership with Ramp, to assess whether frontier models...