arXiv

Marionette: Predicting World States, Rendering Geometry, Painting Appearance

Interactive game world models typically autoregress visual observations directly in pixel or latent space, forcing st...

AI 聚合
2026-08-17
arXiv

Decoding the Past: An Uncertainty-Aware Deep Learning Framework for Sex Attribution in Prehistoric Hand Stencils

Determining the biological sex of the individuals who created Upper Paleolithic hand stencils remains a challenging p...

AI 聚合
2026-08-17
HuggingFace

SPARGen: Unifying Spatial Perception and Reasoning through Native Multimodal Generation

Spatial perception and reasoning from visual observations require recovering geometric structure, establishing corres...

AI 聚合
2026-08-17
HuggingFace

Claim-Level Reliability Assessment for Efficient Test-Time Reasoning

We propose claim-level falsification as a principle for test-time scaling and instantiate it through Claim-Level Reli...

AI 聚合
2026-08-17
HuggingFace

Can We Defend Against AI-Generated Video Attacks on Real-World Crisis Events? A Systematic Evaluation of Detectors, Generators and Social Dissemination

Recent video generators can fabricate realistic depictions of wars, disasters, public emergencies, and other real-wor...

AI 聚合
2026-08-17
HuggingFace

HumanTracker: Towards Comprehensive and Human-Aligned Motion Tracking Benchmark

Humanoid motion tracking is central to teleoperation and whole-body imitation, yet evaluation often disagrees with wh...

AI 聚合
2026-08-17
HuggingFace

Scaling Domain Data Repetition in LLM Pretraining

As large language models scale, their training-token budgets must also increase to maintain an appropriate tokens-per...

AI 聚合
2026-08-17
HuggingFace

Verifier-Induced Support Reshaping in On-Policy Optimization

We show that on-policy reinforcement learning with verifiable rewards (RLVR) can improve the current objective while ...

AI 聚合
2026-08-17
HuggingFace

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction

While Multimodal Large Language Models (MLLMs) have achieved remarkable progress, visual understanding and generation...

AI 聚合
2026-08-17
HuggingFace

CPI-Bench: A Comprehensive,Practical and Intelligent Benchmark for Real-World Image Editing

With the rapid advancement of image editing models and their widespread application across various domains, there is ...

AI 聚合
2026-08-17
HuggingFace

Multimodal Model Diffing for Feature Discovery and Control

Multimodal Large Language Models (MLLMs) exhibit strong visual understanding, yet the internal features that cause th...

AI 聚合
2026-08-17
HuggingFace

MobileMem: Learning from a Year of Mobile Experiences

The next generation of AI agents is increasingly moving beyond systems that answer isolated questions toward persiste...

AI 聚合
2026-08-17
首页 上一页 第 51 / 212 页 下一页 末页