Learning from Failures: Retrieval-Centric CoT via Hard Negatives for Unified Multimodal Retrieval
Unified multimodal retrieval aims to identify candidates that satisfy complex user intent expressed through heterogen...
每天自动聚合 AI 领域最新动态
Unified multimodal retrieval aims to identify candidates that satisfy complex user intent expressed through heterogen...
On-Policy Distillation (OPD) is emerging as a promising alternative to reinforcement learning for LLM post-training, ...
Spatial intelligence is fundamental to embodied agents, yet existing benchmarks focus on local spatial perception fro...
Generating large-scale, freely explorable 3D worlds from open-ended text remains challenging because a system must jo...
Reinforcement learning (RL) with verifiable rewards constructs trajectory-level advantage estimates, yet it often fai...
Training terminal agents requires executable and verifiable tasks that are not merely solvable, but appropriately cha...
Understanding 3D scenes is fundamental to embodied intelligence, requiring joint reasoning over heterogeneous informa...
Training large language model agents for long-horizon tool use typically relies on interactions with real or synthesi...
Vision-language-action (VLA) models often treat main-view and wrist-view observations as parallel visual inputs, over...
End-to-end document parsers provide a unified interface, but serialize page layouts and regional contents into one au...
Computer-using agents (CUAs) are advancing rapidly across the digital world. A CUA trajectory records the agent's act...
Deep vision models exploit shortcuts, relying on cues that correlate with supervision signals. Prior work has focused...