GridVQA-X: A Framework for Evaluating Multimodal Explainability Methods
With the increasing development of Vision-Language Models, it becomes imperative that their predictions are readily e...
每天自动聚合 AI 领域最新动态
With the increasing development of Vision-Language Models, it becomes imperative that their predictions are readily e...
We present PhysiFormer, a diffusion transformer for physically-plausible 3D object motion. Unlike video world models ...
Modern image generation demands a single model that unifies diverse capabilities, including text-to-image (T2I), loca...
While text-to-image (T2I) models have achieved remarkable progress, they struggle with real-world requests that are o...
While generative AI has achieved remarkable success in solving problems with verifiable solutions, generating physica...
Joint-Embedding Predictive Architectures (JEPAs), including recent LeWorldModel (LeWM), have become a promising found...
Outcome-based reinforcement learning provides a stable optimization backbone for language agents, but its sparse traj...
A unified representation for text and vision is a natural pursuit, as it enables simpler multimodal modeling and more...
Video reasoning language models implicitly assume that every input frame is equally reliable. This leads to what we t...
Modern Vision-Language-Action (VLA) models often fail to generalize to novel setups, such as altered camera viewpoint...
A working citation looks like proof -- but the fact that a link resolves does not mean the cited paper supports the c...
Tool use enables large language models (LLMs) to perform complex tasks, and recent agentic reinforcement learning (RL...