Co-RL: Unsupervised Reasoning Emerges from Diverse Cohort in Multi-agent RL
Reinforcement learning (RL) has emerged as a powerful approach for improving reasoning in language and vision-languag...
每天自动聚合 AI 领域最新动态
Reinforcement learning (RL) has emerged as a powerful approach for improving reasoning in language and vision-languag...
Token-level hallucination detectors score each token independently from a single signal, and fail exactly when the ge...
Popular facts are memorised more deeply during pretraining and resist removal longer than rare ones, yet existing LLM...
Agent frameworks increasingly package procedural knowledge as skills: instruction files an agent reads on demand, whi...
Electrocardiography (ECG), photoplethysmography (PPG), and phonocardiography (PCG) provide complementary views of the...
We introduce PTXBench, a benchmark for evaluating and adapting large language models (LLMs) to use architecture-speci...
Reinforcement learning for coding agents increasingly relies on long-running agent harnesses to manage tool integrati...
AI systems are increasingly capable of contributing to mathematical research. In research practice, frontier-model re...
Training terminal agents requires scalable executable supervision, yet synthesizing high-quality terminal tasks remai...
In the Code World Model paradigm an LLM synthesizes an executable world model that a classical planner searches, and ...
State-of-the-art model-based reinforcement learning methods learn neural world models that allow policy improvement b...
Online log anomaly detection is critical for maintaining the reliability of large-scale computing systems. Although r...