Quantifying and Expanding the Theoretical Capacity of Late-Interaction Retrieval Models
Late-interaction retrieval models that use the MaxSim similarity function have shown strong empirical performance, of...
每天自动聚合 AI 领域最新动态
Late-interaction retrieval models that use the MaxSim similarity function have shown strong empirical performance, of...
Dense video captioning aims to generate temporally grounded descriptions of video events, benefiting both event-level...
Recent progress in large-scale generative models has substantially advanced video generation, yet existing methods re...
Audiovisual arts encompass diverse creative disciplines, including cinema, visual arts, stage performance, and game d...
Speculative decoding accelerates Large Language Model (LLM) inference by decoupling draft generation from target veri...
Scaling modern large language models (LLMs) to long contexts is limited by the quadratic computation cost, and poor l...
Significant disparities exist in the diagnosis and clinical presentation of depression across different linguistic po...
Challenges remain in ego-centric 3D scene generation due to limited view overlap and the dominant influence of indivi...
We formulate computer vision as unified multimodal generation, where heterogeneous visual tasks are expressed in the ...
We introduce MentalThink, a visual-symbolic reasoning paradigm that equips Multimodal LLMs (MLLMs) with an executable...
Hierarchical Vision-Language-Action (VLA) models decouple high-level planning from low-level control to improve gener...
On-policy distillation (OPD) trains a student policy by matching a stronger teacher on the student's own trajectories...