Listening Forward: Next Patch Embedding Prediction Enables Scalable Audio Learners
Self-supervised learning (SSL) has driven substantial progress in audio representation learning, though existing meth...
每天自动聚合 AI 领域最新动态
Self-supervised learning (SSL) has driven substantial progress in audio representation learning, though existing meth...
Long-form video understanding encompasses tasks that go beyond retrieving isolated events, including tracking an evol...
LLM-based agents can act on behalf of a user to access cloud services, call tools, or invoke agents. At session start...
Language-specific competency (LSC) is the phenomenon of a language model performing better or worse depending on the ...
Object detectors often produce over-confident predictions for objects outside their training categories, leading to s...
LiDAR scene completion is a key component of 3D perception in autonomous driving, where the scene must be completed i...
Music editing plays a vital role in modern music production, with applications in film, broadcasting, and game develo...
Using reinforcement learning to post-train joint video-audio generation models requires a reward signal. Existing met...
Anticipating how scenes evolve under ego actions is fundamental to safe autonomous driving, yet the full potential of...
Object detection models deployed in safety-critical applications remain vulnerable to backdoor attacks that cause tar...
Multi-teacher on-policy distillation (M-OPD) has emerged as a promising paradigm for consolidating domain-specialized...
Time series imputation is a crucial area for reliable time series analysis, yet it remains challenging due to the com...