AI4AI at Test-Time: Strong-to-Weak Capability Transfer via Harnesses
Recent work on distillation transfers the capabilities of large models to smaller ones often by updating the latter's...
每天自动聚合 AI 领域最新动态
Recent work on distillation transfers the capabilities of large models to smaller ones often by updating the latter's...
Aerial vision-language navigation (VLN) requires an embodied agent to integrate visual evidence over time, plan futur...
The dominant paradigm treats AI safety as a property to be instilled during model training via RLHF, DPO, or Constitu...
Previsualization is an intermediate layer between ideas and production in film, games, architecture, and urban design...
Creative agents still lack an effective way to learn from high-quality human films, limiting their ability to produce...
World modeling is an unsettled field: architectures, training objectives, and state representations interact in compl...
Agentic systems powered by large language models (LLMs) have opened new opportunities for business ideation. Yet exis...
Videos captured through glass often contain reflections that degrade visual quality and interfere with downstream vis...
Recent work on distillation transfers the capabilities of large models to smaller ones often by updating the latter's...
Multimodal expansion of large language models (LLMs) enables new perceptual capabilities but often compromises the la...
The "thinking-with-images" paradigm equips multimodal LLMs with active visual operations such as crop-and-zoom. Howev...
The rapid advancement of Large Language Models (LLMs) is revolutionizing AI for Games by enabling open-ended and flui...