Dual Latent Memory in Vision-Language-Action Models for Robotic Manipulation
Mainstream Vision-Language-Action (VLA) models predict actions primarily from the current observation under a Markovi...
每天自动聚合 AI 领域最新动态
Mainstream Vision-Language-Action (VLA) models predict actions primarily from the current observation under a Markovi...
Generalist robot manipulation policies have advanced rapidly, yet existing benchmarks remain limited in systematicall...
Large Language Models (LLMs) unlocked new possibilities in automated code writing, becoming the backbone of most code...
Structure-property relationships are foundational to biology, chemistry and materials science, where function, reacti...
Complex image creation and editing often require more than a single generation or editing model. A user request may i...
Every chemical language model reading SMILES begins with a tokenizer, yet the field has inherited byte-pair encoding ...
We present HunyuanOCR-1.5, a lightweight end-to-end OCR-specialized vision-language model. HunyuanOCR unifies documen...
Recently, Joint Embedding Predictive Architectures (JEPAs) have attracted significant attention in the computer visio...
We introduce Rank-Then-Act (RTA), a framework for learning control policies from expert video demonstrations without ...
Coding agents increasingly generate pull requests (PRs) for real-world software issues, yet one-shot PR generation re...
Multimodal large language models (MLLMs) generate responses autoregressively, integrating visual and linguistic infor...
We present RuleChef, a framework that uses large language models (LLMs) to generate executable rules for NLP tasks su...