S3Gym: Can LLMs Turn Self-Testing and Self-Judging into Self-Improvement?
Large language models (LLMs) increasingly interact with external environments and accumulate substantial behavioral e...
每天自动聚合 AI 领域最新动态
Large language models (LLMs) increasingly interact with external environments and accumulate substantial behavioral e...
Many important forms of human learning begin with a vague goal, such as "become a better physicist" or "improve at re...
As agents move from research prototypes to deployed tools, their capability increasingly depends on model-external ex...
Historical newspapers are an abundant record of public life, but their dense, irregular and sometimes noisy layouts m...
The attention prefilling phase of long-context LLM inference scales quadratically, making self-attention a severe com...
An image may be worth a thousand words, but most captioning models describe it in only a few. Modern vision-language ...
Traditional speaker-attributed ASR systems treated ASR and speaker diarization as two separate tasks. Recently, end-t...
Mobile AI acts as a visual oracle, empowering users to snap a picture of something and ask for information. Snap-and-...
Compact token sequences are essential for efficient 3D generation. However, existing 3D tokenizers typically organize...
Large language models (LLMs) trained only on text and code can sometimes generate programs that draw recognizable ima...
Knowledge-Based Visual Question Answering (KB-VQA) relies on retrieving external information to answer queries involv...
LLMs are increasingly deployed as orchestrators that coordinate specialized subagents to solve complex tasks through ...