Push-Wiper: Toward General-Purpose Robotic Cleaning across Varied Stains and Surfaces with Segmented Pushing Trajectories
Viscous stains, characterized by high viscosity and complex rheological properties, remain a major challenge for robo...
每天自动聚合 AI 领域最新动态
Viscous stains, characterized by high viscosity and complex rheological properties, remain a major challenge for robo...
We introduce Video-DeepResearch (Video-DR), extending multimodal agents from static images to continuous video stream...
Persona skills distill personal interaction histories into portable and executable artifacts for downstream agents. W...
Diffusion language models (dLLMs) offer an alternative to autoregressive (AR) language modeling, yet the scaling beha...
Scientific poster construction compresses a long multimodal paper into a readable, editable canvas. Existing systems ...
Modern agent frameworks equip large language models with external skill libraries to solve complex tasks. However, it...
Video world models predict future observations conditioned on historical observations and control signals, enabling l...
We introduce a new problem domain for human action recognition: the fine-grained analysis of children's gait behavior...
Existing evaluations often reduce agent failures to system-level outcomes, obscuring where the fault originated and w...
Pixel-space diffusion models aim to learn an end-to-end generator directly over raw pixels. This is challenging becau...
Geospatial foundation models aim to learn representations that transfer across regions and sensors, yet evaluating th...
Scientific figure comprehension and reasoning using multimodal AI requires integrating visual perception with domain-...