HuggingFace

The Next Screenshot Knows: Gated Hindsight Distillation for Mobile GUI Agents

GUI agents are commonly trained offline from successful interaction trajectories. Standard training decomposes each t...

AI 聚合
2026-08-12
arXiv

Agentic Auto-Research is Fuzz Testing

Autonomous research agents can generate experiments faster than researchers can validate them. Researchers have respo...

AI 聚合
2026-08-11
arXiv

Agentic Harnesses: LLM-Driven Verification Layers for Robot Autonomy

Advances in advanced artificial intelligence tools have sparked research in robot autonomy, but the development of su...

AI 聚合
2026-08-11
arXiv

Towards Expert-level Medical AI for Real-time Video Consultations

Audio-visual interaction is the standard for patient-physician consultations, enabling natural communication and effe...

AI 聚合
2026-08-11
arXiv

Stealing Reasoning Traces from Proprietary LLM APIs

Leading large language model providers now conceal their models' step-by-step reasoning, or chain-of-thought, to prot...

AI 聚合
2026-08-11
arXiv

Sci-VBench: Evaluating Knowledge- and Reasoning-Intensive Video Generation in Science Domains

We introduce Sci-VBench, a comprehensive benchmark for evaluating knowledge- and reasoning-intensive video generation...

AI 聚合
2026-08-11
arXiv

ArchAgent v2: A Case Study with the Data Prefetching Championship

Agentic artificial intelligence has shown great promise in automating algorithm design, but scaling similar technique...

AI 聚合
2026-08-11
arXiv

Energy-Structured Latent World Models with Neural Time Fields for Physically Constistent Open-World Motion Planning

Physically consistent motion planning remains a fundamental challenge in embodied AI, as generated trajectories must ...

AI 聚合
2026-08-11
arXiv

SHE: Trajectory-driven Safety Harness Evolution for LLM Agents

The safety of large language model (LLM) agents depends not only on model weights but also on the agent harness that ...

AI 聚合
2026-08-11
arXiv

BDH-CQ: In-Context Learning with Recurrent Latent Reasoning

We introduce BDH-CQ, a reasoning model that combines in-context learning with recurrent latent reasoning. Inputs pres...

AI 聚合
2026-08-11
arXiv

Fusion Training for Mathematical Generalization in Large Language Models

Thinking Mode Fusion (TMF) enables large language models to support both concise responses and long-form reasoning by...

AI 聚合
2026-08-11
arXiv

DSLE: A Learning Environment for Dark Souls Boss Encounters

We introduce the Dark Souls Learning Environment (DSLE), a containerized platform that presents all 22 boss encounter...

AI 聚合
2026-08-11
首页 上一页 第 64 / 212 页 下一页 末页