VGI-BENCH: Probing Visual Intelligence in Video Generation Models
Recent studies suggest that video generation models can exhibit certain forms of zero-shot visual reasoning through g...
每天自动聚合 AI 领域最新动态
Recent studies suggest that video generation models can exhibit certain forms of zero-shot visual reasoning through g...
Agent capability is not determined by the model alone. The agent harness, encompassing memory management, planning st...
Vision-language models can produce fluent answers that are insufficiently grounded in the visual evidence: a single u...
Multi-teacher on-policy distillation (M-OPD) has emerged as a promising paradigm for consolidating domain-specialized...
Vision-Language-Action (VLA) models have demonstrated effectiveness in robot manipulation, yet state-of-the-art model...
Multi-arm collaboration is becoming a core capability in embodied manipulation. Recent vision-language-action (VLA) m...
Massively parallel simulation changes the data regime in which off-policy reinforcement learning (RL) is trained, cha...
Scaling transformer language models creates an inherent tension between expressivity and memory efficiency. While uni...
Conversational systems, such as duplex speech language models (SLMs), still lack a streaming, accurate, and empatheti...
Streaming autoregressive diffusion models enable real-time, long-horizon video generation, but their training objecti...
Large vision-language models have shown strong progress in UI-to-code generation, yet their test-time self-evolution ...
Reliable reward models are essential for text-to-video evaluation and alignment. However, the trade-off between evalu...