HuggingFace

Beyond the Current Observation: Evaluating Multimodal Large Language Models in Controllable Non-Markov Games

Deploying multimodal foundation models as closed-loop policies increasingly requires conditioning actions on observat...

AI 聚合
2026-06-18
HuggingFace

Guava: An Effective and Universal Harness for Embodied Manipulation

Language models trained on large-scale vision-language data have demonstrated strong potential for embodied agents. H...

AI 聚合
2026-06-18
HuggingFace

Externalizing Research Synthesis and Validation in AI Scientists through a Research Harness

AI systems can increasingly automate scientific workflows, but the reasoning that links prior evidence, generated ide...

AI 聚合
2026-06-18
HuggingFace

CEO-Bench: Can Agents Play the Long Game?

Language model agents are becoming proficient executors at isolated, short-horizon tasks such as software engineering...

AI 聚合
2026-06-18
HuggingFace

Physics-IQ Verified

Video generative models ( VGMs) have become a new frontier that can be used not just for video generation but for a m...

AI 聚合
2026-06-18
HuggingFace

PAIWorld: A 3D-Consistent World Foundation Model for Robotic Manipulation

World foundation models (WFMs) are powerful simulators, yet they predominantly operate in a single-view setting and l...

AI 聚合
2026-06-18
HuggingFace

Learning User Simulators with Turing Rewards

Learning to simulate human users in interactive settings could advance the training of agent assistants, evaluation o...

AI 聚合
2026-06-18
HuggingFace

Kairos: A Native World Model Stack for Physical AI

World models are transitioning from passive visual generators to foundational, operational infrastructure for Physica...

AI 聚合
2026-06-18
HuggingFace

IndustryBench-MIPU: Benchmarking Multi-Image Attribute Value Extraction for Industrial Products

Industrial products such as valves and circuit breakers are defined by dense technical specifications that govern pro...

AI 聚合
2026-06-18
HuggingFace

Sumi: Open Uniform Diffusion Language Model from Scratch

Diffusion models have become a promising alternative to autoregressive models. Among these, uniform diffusion languag...

AI 聚合
2026-06-18
HuggingFace

Trust the Right Teacher: Quality-Aware Self-Distillation for GUI Grounding

Graphical user interface (GUI) grounding requires vision-language models (VLMs) to identify small target elements in ...

AI 聚合
2026-06-18
HuggingFace

SAE Interventions are Unreliable: Post-Intervention Recovery of Suppressed Behavior

Sparse Autoencoders (SAEs) decompose residual-stream activations into interpretable features. Recent latent-space def...

AI 聚合
2026-06-18
首页 上一页 第 193 / 212 页 下一页 末页