Marionette: Predicting World States, Rendering Geometry, Painting Appearance
Interactive game world models typically autoregress visual observations directly in pixel or latent space, forcing st...
每天自动聚合 AI 领域最新动态
Interactive game world models typically autoregress visual observations directly in pixel or latent space, forcing st...
Determining the biological sex of the individuals who created Upper Paleolithic hand stencils remains a challenging p...
Spatial perception and reasoning from visual observations require recovering geometric structure, establishing corres...
We propose claim-level falsification as a principle for test-time scaling and instantiate it through Claim-Level Reli...
Recent video generators can fabricate realistic depictions of wars, disasters, public emergencies, and other real-wor...
Humanoid motion tracking is central to teleoperation and whole-body imitation, yet evaluation often disagrees with wh...
As large language models scale, their training-token budgets must also increase to maintain an appropriate tokens-per...
We show that on-policy reinforcement learning with verifiable rewards (RLVR) can improve the current objective while ...
While Multimodal Large Language Models (MLLMs) have achieved remarkable progress, visual understanding and generation...
With the rapid advancement of image editing models and their widespread application across various domains, there is ...
Multimodal Large Language Models (MLLMs) exhibit strong visual understanding, yet the internal features that cause th...
The next generation of AI agents is increasingly moving beyond systems that answer isolated questions toward persiste...