FM-Bench: A Benchmark for Long-Horizon Management with Competing Agents
Language model agents now execute bounded tasks reliably. Whether they can sustain effective decision-making over lon...
每天自动聚合 AI 领域最新动态
Language model agents now execute bounded tasks reliably. Whether they can sustain effective decision-making over lon...
Looped language models have shown promising results on reasoning benchmarks, yet their potential for agentic tool use...
Continuous self-improvement requires an ever-expanding pool of self-generated, diverse, adaptive goals. For language ...
Programmable logic controllers (PLCs) run industrial plants, and large language models can already generate independe...
Single-step retrosynthesis is a central component of computer-aided synthesis planning, yet its intrinsically one-to-...
Open-weight language models are fine-tuned, quantized, pruned, and merged, yet their provenance is often undocumented...
We introduce Semantic Task Completion Video Generation, an outcome-oriented video generation task. Under this formula...
Physical interaction quality is central to deformable-object manipulation, yet most benchmarks evaluate task success ...
JEPA-style latent world models can use Euclidean distance to a goal latent as the cost for model-predictive control (...
Embodied agents are increasingly used to close the gap left by end-to-end policy models. Yet the agentic path has not...
Recent advances in foundation models have enabled AI scientists to automate increasingly complete research workflows,...
High-quality creative writing data for large language models (LLMs) remains dominated by story-centric data, limiting...