On-Policy Self-Distillation in Diffusion Models
Reinforcement learning can align diffusion models with human preferences and task-specific objectives, but endpoint r...
每天自动聚合 AI 领域最新动态
Reinforcement learning can align diffusion models with human preferences and task-specific objectives, but endpoint r...
LLM agents remain unreliable on long-horizon tasks, where small local failures can compound over extended interaction...
Reinforcement Learning with Verifiable Rewards (RLVR) and on-policy distillation (OPD) have become two widely adopted...
Group-based reinforcement learning methods such as GRPO for large language models avoid training a critic by sampling...
Machine translation tests masked diffusion language models (dLLMs) because every source token must be rendered faithf...
Morphological transforms are long-standing tools for shape and mask processing, but the de facto reference implementa...
Universal multimodal embeddings are becoming a core component of modern AI systems, enabling heterogeneous content to...
We present LAION-BVD, a large-scale open video dataset for multimodal learning, which contains 1.3B platform-specific...
Smart glasses are evolving from capture and display accessories into first-person intelligence platforms that connect...
Multimodal large language models (MLLMs) have become a prevailing paradigm for unified video perception. However, pos...
Video games provide a scalable source of training data for video world models, offering diverse environments, complex...
As large language models (LLMs) continue to advance in coding capabilities, their potential in cybersecurity has draw...