Hunyuan3D-Buffalo 1.0: A Unified Multimodal Model for Scalable 3D Generation, Understanding, and Editing
Recent advances in image generation have demonstrated the potential of unified multimodal models that integrate under...
每天自动聚合 AI 领域最新动态
Recent advances in image generation have demonstrated the potential of unified multimodal models that integrate under...
Real-time video editing requires low-latency causal generation with bounded computational resources while preserving ...
Tool-Integrated Reasoning (TIR) enables LLMs to solve complex tasks through iterative tool interactions. However, exi...
Captions serve as a primary supervision signal for both multimodal understanding and text-to-image generation. Howeve...
Language remains an outlier in generative modeling: while images, video, and audio are increasingly modeled in contin...
Omni-modal large language models (Omni-LLMs) have achieved remarkable performance on audio-visual understanding tasks...
Large language model agents have shown strong potential in complex interactive tasks, yet their reinforcement learnin...
Recursive self-improvement requires agents to turn accumulated experience into better future behavior. Personal AI ag...
Industrial recommenders increasingly adopt the pretrain-then-transfer paradigm, yet behavioral distribution drift rai...
Large Language Model (LLM) agents have seen rapid adoption in software engineering. As agents take a greater role in ...
Continually improving agents require dynamic interaction feedback beyond static supervision, yet direct real-environm...
World Action Models (WAMs) have emerged as a promising paradigm by jointly modeling robot actions and future visual d...