The Missing Temporal Link: Temporal Context Routing for Script-Driven Audio-Video Generation
Joint audio-video generation models have made substantial progress in visual quality and audio-visual synchronization...
每天自动聚合 AI 领域最新动态
Joint audio-video generation models have made substantial progress in visual quality and audio-visual synchronization...
Streaming video understanding requires multimodal large language models (MLLMs) to process continuous visual inputs a...
We propose Puffin-World, a unified multimodal architecture that integrates physical understanding, spatial simulation...
Large language models offer broad capabilities, but adapting them to evolving domains, tools, and requirements often ...
Personalized assistants should not only comply with user requests but also assess whether those requests are appropri...
We study the dynamics of Stochastic Gradient Descent (SGD), which is known to steer deep neural networks toward invar...
Many recurring text functions are easy to describe but difficult to implement with rules, while calling a large remot...
Post-hoc calibration corrects reported confidence, yet a multiclass calibrator can also change the associated top-1 p...
Online 3D reconstruction models perform poorly on long videos. This happens because regressing poses relative to a fi...
Understanding agent behavior requires methods that scale to thousands of trajectories and surface new patterns in lon...
We present FoldingAgent, an agentic framework for inferring explicit parametric folding programs directly from origam...
Multimodal models often build on architectures designed for generative vision-language modeling, typically combining ...