SANA-Video 2.0: Hybrid Linear Attention with Attention Residuals for Efficient Video Generation
We introduce SANA-Video 2.0, a hybrid video diffusion transformer instantiated at 5B and 14B scales under a unified a...
每天自动聚合 AI 领域最新动态
We introduce SANA-Video 2.0, a hybrid video diffusion transformer instantiated at 5B and 14B scales under a unified a...
We study on-policy distillation (OPD) for agentic tasks, where an LLM agent interacts with an environment over multip...
We revisit dataset distillation from an outcome-centric perspective. Rather than aligning process surrogates (per-ste...
Modern AI agents rely on elaborate inference harnesses such as Claude Code, Codex, and OpenClaw to drive multi-turn r...
Even a current high-capability LLM can appear safer when shown a dangerous objective directly than when other agents ...
Creating dynamic and physically realistic 4D worlds from natural language descriptions is both fascinating and challe...
Test-Time Tuning (TTT) on pretrained diffusion models has emerged as a powerful paradigm for video editing. However, ...
Concurrent stateful library APIs expose behavior through evolving resource ownership, lifecycle states, and competing...
The rapid progress of AI has intensified the long-standing pursuit of automation: replacing human participation with ...
Unlike large language models (LLMs) that exhibit strong reasoning capabilities, vision-language models (VLMs) struggl...
On-policy self-distillation (OPSD) is promising as it removes the external teacher required by on-policy distillation...
Modern AI agents rely on elaborate inference harnesses such as Claude Code, Codex, and OpenClaw to drive multi-turn r...