Revisiting Local Context for Long-Horizon Streaming 3D Reconstruction
Streaming 3D reconstruction from extremely long videos requires estimating camera motion and scene geometry online un...
每天自动聚合 AI 领域最新动态
Streaming 3D reconstruction from extremely long videos requires estimating camera motion and scene geometry online un...
Self-evolving language models have recently emerged as a promising path toward superintelligence, with the advantage ...
We present StarHarness, a framework for evolving environment-specific agent harnesses while keeping model weights fix...
Evaluation is shifting from static QA toward agentic settings where models act through external tools. We identify a ...
Language model pretraining has become almost synonymous with prohibitive cost, placing it out of reach for much of th...
Autoregressive video diffusion enables scalable long-video generation by producing chunks from a bounded recent conte...
Scaling video generation to long durations reveals a critical bottleneck: current models lack robust long-term memory...
Recurrent fast-weight memories and selective state-space models compress an expanding context into a fixed-size recur...
Spatio-temporal video grounding (STVG) requires models to identify when a referred event occurs and localize the targ...
Equipping Large Language Models (LLMs) with multi-turn tool-calling capabilities is essential for building autonomous...
Generative models can turn natural-language prompts into images, text, code, and other content, lowering the cost of ...
The alignment of Large Language Models heavily relies on English-centric high-quality preference data, which often le...