Video-Oasis: Rethinking Evaluation of Video Understanding
The inherent complexity of video understanding makes it difficult to determine whether Video-LLM benchmark performanc...
每天自动聚合 AI 领域最新动态
The inherent complexity of video understanding makes it difficult to determine whether Video-LLM benchmark performanc...
The rapid development of large language models and multimodal large language models has accelerated the emergence of ...
Zero-Shot Compositional Action Recognition (ZS-CAR) requires recognizing novel verb-object combinations composed of p...
Semantic audio applications increasingly require controllable generation on commodity and embedded hardware rather th...
In this work, we present Canvas360, a two-stage framework for in-context panoramic generation that combines geometry-...
Recovering high-quality video from sparse event streams is a challenging task. Regression methods often blur textures...
Current computational approaches for drug design typically focus on generating molecules conditioned on specific targ...
Self-attention lets each token retrieve information from the full context, but its quadratic cost in sequence length ...
Embodied agents are typically built as hand-designed compositions of perception, memory, planning, and action modules...
Long-horizon failure in world models is conventionally attributed to compounding error, a generic framing that does n...
Linear attention models allow a fixed state size and a fixed amount of compute per token. However, due to their limit...
Reinforcement learning (RL) is becoming increasingly important for post-training large language models (LLMs). Previo...