Hierarchical Advantage Weighting for Online RL Fine-Tuning of VLAs from Sparse Episode Outcomes
When pretrained VLA policies are fine-tuned through online RL, each rollout episode produces only a single binary out...
每天自动聚合 AI 领域最新动态
When pretrained VLA policies are fine-tuned through online RL, each rollout episode produces only a single binary out...
Advanced reasoning typically requires Chain-of-Thought prompting, which is accurate but incurs prohibitive latency an...
Polymarket has emerged as a prominent prediction market platform and one of the fastest-growing applications in DeFi....
Efficient and scalable agentic intelligence requires models that can deliver both low-latency responses and strong re...
We introduce the Massive Video Embedding Benchmark (MVEB), a 23-task benchmark for video embeddings spanning classifi...
Vision-Language-Action models (VLAs) leverage large-scale vision-language pretraining for semantic robot control, but...
Sparse autoencoders (SAEs) are widely used to interpret neural network representations, but their utility depends on ...
Humans naturally understand object physics through everyday interactions, but faithfully predicting complex deformabl...
Humans can grasp objects effortlessly, whereas multi-fingered robots are far from this level of generality. We argue ...
Sparse reward reinforcement learning (RL) has become a standard tool for improving LLM reasoning, but its success dep...
Re-rendering an existing video from a novel camera viewpoint requires the output to follow the prescribed camera traj...
Progress in AI has largely been driven by methods that assume less. As compute and data increase, approaches with wea...