Thinking on Shots: Consistent Multi-Shot Video Editing with Agentic Reasoning
While generative AI has significantly advanced video editing, existing methods primarily focus on single-shot or shor...
每天自动聚合 AI 领域最新动态
While generative AI has significantly advanced video editing, existing methods primarily focus on single-shot or shor...
Native 3D generators now recover impressive mesh geometry from a single image. However, a dense mesh stays soft where...
A common strategy for scaling world models is to train on more crawled video with more compute. We argue that this st...
Long-horizon agent runs generate experience that can improve both the current run and future work. Most self-improvem...
We introduce LibriBrain100, a large-scale MEG dataset for speech decoding designed from the ground up for reproducibl...
The development of 0.1^{circ} global weather forecasting models based on machine learning (ML) is constrained by the ...
Test-time scaling uses extra test-time compute to improve performance, such as letting language models reason longer ...
Large language models access knowledge inconsistently across languages, but to what extent do they differ in their sk...
Computer-use agents ground natural-language instructions in screenshots to locate interface elements, yet existing be...
On-policy self-distillation (OPSD) uses a privileged copy of the student model to provide dense supervision without a...
Answer accuracy is an insufficient reliability signal for LLM data agents. In structured-data tasks, a benchmark-corr...
We consider optimization applications with unknown parameters where the decision maker believes that the optimal valu...