MemSyco-Bench: Benchmarking Sycophancy in Agent Memory
Memory has emerged as a cornerstone of modern LLM-based agents, supporting their evolution from single-turn assistant...
每天自动聚合 AI 领域最新动态
Memory has emerged as a cornerstone of modern LLM-based agents, supporting their evolution from single-turn assistant...
Traditional robot programming is challenging: it requires orchestrating multimodal perception, managing physical cont...
We present Seed2.0, a model series that takes a meaningful step toward solving complex, real-world tasks. Our approac...
In prefill-decode (PD) disaggregated LLM serving, each request is assigned to a decode worker after prefill. Existing...
Multimodal Large Language Models (MLLMs) are often constrained by a language-space bottleneck, forcing complex visual...
Classic 3D scene graph generation approaches fail to work in real-time due to the heavy computational cost of environ...
We introduce PerceptionRubrics, a rubric-based evaluation framework that addresses the gap between saturated benchmar...
Slide design requires personalizing both deck themes and page layouts. Yet, current AI agent-based methods struggle w...
Lightweight machine learning models are increasingly proposed for intrusion detection in Industrial Internet of Thing...
AI translation of literary works is increasingly common. While the content may be rendered adequately, we do not know...
Accelerating materials discovery requires AI systems that can generate scientifically valid hypotheses through multi-...
Procedural memory is increasingly used to improve LLM agents on recurring workplace tasks, yet its ability to produce...