arXiv

Argus: A General-Purpose Agentic Runtime for Long-Horizon Reasoning

Long-horizon reasoning requires an agentic runtime that can persist when evidence supports its current approach and p...

AI 聚合
2026-08-06
HuggingFace

UniWorld-View: Large-Baseline View Synthesis via Video Diffusion Models

The abundance of casually captured monocular videos and images on social media provides a valuable source for immersi...

AI 聚合
2026-08-06
HuggingFace

SKILL-KD: Contrastive Skill Distillation for LLM Agents

Skill-based prompting has become a practical mechanism for improving large language model (LLM) agents, yet existing ...

AI 聚合
2026-08-06
HuggingFace

NOLLI: A Difficulty-Calibrated Puzzle Benchmark for Diagnosing the English-Korean Performance Gap

We introduce NOLLI, a procedurally generated English-Korean puzzle benchmark designed to diagnose where Korean perfor...

AI 聚合
2026-08-06
HuggingFace

When Memory Lies: An Empirical Study of Spatial Memory Staleness in VLM Agents

Memory-augmented VLM agents act on persistent spatial knowledge, yet that knowledge silently goes stale as the enviro...

AI 聚合
2026-08-06
HuggingFace

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning

Long-horizon reasoning in recent LLMs demands that the model switch between distinct skills inside a reasoning chain,...

AI 聚合
2026-08-06
HuggingFace

The Personalization Mirage: How LLMs Fabricate User Profiles, and Why Self-Monitoring Misleads

Personalized LLMs with persistent memory are increasingly deployed, yet the faithfulness of their user models remains...

AI 聚合
2026-08-06
HuggingFace

When Teachers Mislead: Spurious-Signal-Aware On-Policy Distillation

On-Policy distillation (OPD) transfers teacher capabilities by supervising student-sampled trajectories with dense to...

AI 聚合
2026-08-06
HuggingFace

OneDayAgent: Towards a Long-Horizon Harness for Autonomous Agents

LLM agents are increasingly applied to open-ended everyday requests that span work, study, and life. These tasks are ...

AI 聚合
2026-08-06
HuggingFace

WorldCycle: Self-Verifiable Reinforcement Learning for Long-Horizon Video World Models

Interactive video world models are essential for long-horizon planning and exploration, yet they suffer from compound...

AI 聚合
2026-08-06
HuggingFace

BridgeVLA++: A Data-Efficient, Generalizable, and Memory-Augmented Vision-Language-Action Framework for 3D Manipulation

Leveraging pre-trained vision-language models (VLMs) to construct vision-language-action (VLA) models has emerged as ...

AI 聚合
2026-08-06
HuggingFace

Distill Where You Fail: Recovering Learning Signals of Negative RL-Groups from Adaptive Teacher Guidance

Reinforcement learning with verifiable rewards (RLVR) has become a standard paradigm for post-training large language...

AI 聚合
2026-08-06
首页 上一页 第 79 / 215 页 下一页 末页