WorldCupArena: Fine-Grained Evaluation of Language Models and Deep-Research Agents on Football Forecasting
Predicting a football match before kickoff requires more than knowing past results: a model must use changing informa...
每天自动聚合 AI 领域最新动态
Predicting a football match before kickoff requires more than knowing past results: a model must use changing informa...
Training tool-use agents to improve from their own experience remains challenging, as supervised fine-tuning relies o...
Despite recent scaling successes, multilingual ASR performance remains highly uneven, with long-tail languages suffer...
Temporal grounding in long recordings remains challenging for audio-conditioned LLMs. We present a time-aware audio L...
We present a continuous geometric framework that models the discrete algebraic operations of the Transformer architec...
Model merging is promoted as a substitute for joint multi-task training, yet in the reinforcement-learning setting th...
Retinal layer segmentation in Optical Coherence Tomography (OCT) is a fundamental step for extracting quantitative bi...
Agentic Artificial Intelligence (AI), enabled by Large Language Models, marks a shift from rule-based automation towa...
The post-training of Vision-Language-Action (VLA) models is essential due to the diversity of simulators, robot embod...
Multimodal sarcasm and cyberbullying detection remain challenging because the intended meaning often emerges from inc...
Transferring policies across domains poses a vital challenge in reinforcement learning, due to the dynamics mismatch ...
Reinforcement learning (RL) has become central to improving large language models (LLMs) on complex reasoning tasks, ...