On-Policy Delta Distillation
On-policy distillation is an alternative post-training method in reinforcement learning that alleviates the constrain...
每天自动聚合 AI 领域最新动态
On-policy distillation is an alternative post-training method in reinforcement learning that alleviates the constrain...
We present Xiaomi-Robotics-1, a foundational vision-language-action (VLA) model capable of (1) following diverse lang...
Video generative models commonly rely on latent spaces learned by 3D Variational Autoencoders (3D-VAEs). However, con...
In this report, we introduce Qwen-Music, a powerful music generation model capable of producing highly musical and hi...
Graph retrieval-augmented generation (GraphRAG) enhances large language models with structured knowledge, yet existin...
We present S1-Omni, a unified multimodal reasoning model for scientific understanding, prediction, and generation. AI...
Muon is competitive with AdamW in large-scale pre-training, but its value for reinforcement-learning (RL) post-traini...
Reinforcement learning with verifiable rewards (RLVR) commonly uses entropy for advantage shaping. However, entropy c...
Skills are a useful abstraction for software agents, turning human and agent experience into reusable procedural know...
Reinforcement learning from verifiable rewards (e.g. GRPO) is the engine behind today's reasoning models, yet it grad...
Vision-language-action (VLA) models predict robot actions from visual observations and language instructions. These a...
Autonomous negotiation agents are increasingly deployed in high-stakes settings such as insurance and procurement. Wh...