MPIE-Bench: Benchmarking Anatomically Plausible Multi-Person Interaction Editing
Text-to-image and personalized editing models now synthesize high-fidelity single-subject images with ease. Yet placi...
每天自动聚合 AI 领域最新动态
Text-to-image and personalized editing models now synthesize high-fidelity single-subject images with ease. Yet placi...
Role-playing agents (RPAs) have become one of the most important consumer applications of large language models. User...
This paper proposes AI Tour Meeting, a group travel planning framework powered by multiple Large Language Model (LLM)...
Speculative Decoding (SD) accelerates large language model inference by allowing a lightweight draft model to propose...
The deep learning revolution, kicked off by AlexNet, taught us that end-to-end training beats decomposing a problem i...
Multimodal agents for visual question answering increasingly operate as multi-step trajectories that interleave perce...
Forward latent world models predict how actions change a scene, but recover actions for a desired change only through...
Transformers propagate information across depth through a single additive residual stream: every sublayer reads only ...
On-device speech emotion recognition (SER) is critical for real-time applications, yet large self-supervised models t...
We present ShadowDancer, a novel approach to any-action, frame-level control of interactive video world models. The o...
Deep Research agents extend LLM-based assistants into long-horizon workflows involving planning, retrieval, evidence ...
In our prior work, Pedestrian Archetypes, we defined pedestrian archetypes as collections of behaviors that uniquely ...