Demystifying On-Policy Distillation: Roles, Pathologies, and Regulations
On-policy distillation (OPD) has become a key paradigm in LLM post-training, yet its training dynamics remain poorly ...
每天自动聚合 AI 领域最新动态
On-policy distillation (OPD) has become a key paradigm in LLM post-training, yet its training dynamics remain poorly ...
Video generation increasingly relies on keyframe-based workflows, where creators specify a sequence of reference imag...
Multi-reference-to-audio-video (MR2AV) generation aims to generate coherent audio-video content conditioned on multip...
Recent advances in video understanding have spanned motion, long video, and streaming interaction, driving this field...
Looped Transformers scale sequential computation by applying a compact stack of physical blocks for multiple rounds, ...
Learning broad world knowledge directly from raw visual data is a fundamental capability of intelligence. We introduc...
Building interactive worlds that respond coherently to player actions has long been a shared goal of computer graphic...
In-context learning is commonly interpreted as a form of conditional inference, in which the prompt specifies a conte...
Reinforcement learning has become a standard post-training recipe for large language models, but dense full-parameter...
Agentic retrieval-augmented generation (RAG) extends static RAG by allowing language models to iteratively reason, ge...
Visually impaired individuals (VIIs) encounter significant daily challenges due to limited access to visual informati...
A growing gap separates inference context lengths from RL post-training: inference systems are approaching million-to...