PointDiT: Pixel-Space Diffusion for Monocular Geometry Estimation
State-of-the-art single-image 3D reconstruction methods often rely on complex hybrid architectures and loss functions...
每天自动聚合 AI 领域最新动态
State-of-the-art single-image 3D reconstruction methods often rely on complex hybrid architectures and loss functions...
LLM agents increasingly rely on retrieval buffers to store and reuse past experience, yet the cache management polici...
We introduce Gemma 4, a new generation of open-weight, natively multimodal language models in the Gemma model family....
Mathematical reasoning has become a central task for evaluating and tuning reasoning Large Language Models (LLMs), ye...
Embodied navigation aims to build agents that interpret multimodal goals, reason in 3D space, and reach target destin...
Despite recent progress of VLA foundation models, the disparity between laboratory conditions and real-world applicat...
Academic output is produced across a fragmented toolchain: literature discovery in one application, reference managem...
We present the AI Wizards submission to EXIST 2026 for multimodal sexism identification in memes. The task is compose...
Concept erasure aims to remove a target concept from a representation while preserving the other information encoded ...
In longitudinal clinical practice, every chest X-ray is read in the context of the patients prior exam, and much of w...
Crossmodal correspondences between sound and taste are well established in psychology and neuroscience, but largely a...
Audio intelligence involves understanding, reasoning about, and generating both audio and speech. In this work, we in...