Hy-Embodied-0.5-VLA: From Vision-Language-Action Models to a Real-World Robot Learning Stack
In this report, we present Hy-Embodied-0.5-VLA, abbreviated as HyVLA-0.5, an end-to-end system that spans the full ro...
每天自动聚合 AI 领域最新动态
In this report, we present Hy-Embodied-0.5-VLA, abbreviated as HyVLA-0.5, an end-to-end system that spans the full ro...
Generating avatar videos that are not merely visually similar to a target individual but behaviorally recognizable, f...
The recent success of agent swarms has shifted the paradigm of large language model (LLM)-based agents from single-ag...
Coding agents powered by large language models have demonstrated strong performance on software engineering tasks. Ye...
Large language models (LLMs) are widely used in text-to-image (T2I) systems, but they are typically limited to text e...
Cloning camera motion from reference videos is an important task in video generation, as videos provide intuitive and...
Recent advancements in video-based world models have demonstrated an unprecedented ability to synthesize high-fidelit...
Users rely on execution traces to observe agent behavior, diagnose failures, and ensure accountability. These traces ...
Video generation models based on Diffusion Transformers (DiTs) have achieved remarkable performance in video synthesi...
Recent advances in agentic Reinforcement Learning (RL) have substantially improved the multi-turn tool-use capabiliti...
AI agents are increasingly being developed to accelerate scientific discovery, yet their practical capabilities in re...
We identify a new dimension for enhancing rollout diversity in Group Relative Policy Optimization (GRPO) for LLMs. Wh...