3DZip: Spatial-Aware Feature Diversity-Guided Token Compression for 3D Question Answering
Recent 3D vision-language models (3D VLMs) construct geometry aware tokens by projecting 2D visual features into worl...
每天自动聚合 AI 领域最新动态
Recent 3D vision-language models (3D VLMs) construct geometry aware tokens by projecting 2D visual features into worl...
Video motion transfer aims to animate a target object using dynamics from a reference video. Existing formulations la...
Multimodal large language models (MLLMs) have advanced visual understanding and reasoning, yet their static parametri...
Vision-language MoE batches contain different numbers of image and text tokens. Image resolution, image count, tiling...
Adaptive rounding methods such as GPTQ, or equivalently Babai's nearest plane algorithm, round a real matrix to integ...
Optimization-based latent reasoning improves large language model outputs by optimizing instance-specific continuous ...
Accurate prediction of object trajectories during manipulation is essential for closing the perception-action loop. P...
Optimizing modern recommender models still depends heavily on engineers manually iterating over architectural, object...
Despite the impressive visuomotor capabilities enabled by Vision-Language-Action (VLA) models, their performance ofte...
Reinforcement learning with verifiable rewards (RLVR) broadcasts a single response-level reward to every token, while...
Emotional dialogue research includes two influential strategy traditions. Empathetic dialogue prioritizes understandi...
The web is increasingly accessed by AI agents rather than humans. Every agent needs knowledge, especially in the life...