arXiv

How to Train a Critic Stably and Efficiently

Group-based reinforcement learning methods such as GRPO for large language models avoid training a critic by sampling...

AI 聚合
2026-08-25
HuggingFace

Beyond Imitation: Filtering On-Policy Distillation by Reasoning Progress

On-policy distillation (OPD) has emerged as an effective framework for post-training language models by pairing stude...

AI 聚合
2026-08-25
HuggingFace

Prime Agent: A Self-Improving RLM Harness

Language models are sequential processors, but long-horizon agency requires external information and computation beyo...

AI 聚合
2026-08-25
HuggingFace

ReWorld: An Interactive World Model with Long-Horizon Memory

An interactive world model must follow the user's actions, remember the places it has shown, and stream in real time....

AI 聚合
2026-08-25
HuggingFace

Unlocking the Potential of Image Editing via Concept Scaling and Dense Supervision

Existing image editing frameworks predominantly follow the training paradigm of text-to-image diffusion models. Howev...

AI 聚合
2026-08-25
HuggingFace

GameXpert-Bench: How Far Are Coding Agents from Expert Game Development?

Recent large language models (LLMs) can operate as coding agents that build complete games from natural language requ...

AI 聚合
2026-08-25
HuggingFace

ARC: Fair Relative Advantage Comparison in Open-Ended Real-World Interaction

Open-ended real-world interaction admits multiple valid behaviors: an agent may answer directly, ask for clarificatio...

AI 聚合
2026-08-25
HuggingFace

Task-CoEvolve: Efficient Harness Optimization via Adaptive Validation Task Selection

We present a novel approach to efficient LLM harness optimization through adaptive validation task selection. Harness...

AI 聚合
2026-08-25
HuggingFace

Better Retrieval, Worse Robustness:How Multi-hop RAG Amplifies Upstream ASR Errors

Speech-based applications pass spoken queries through automatic speech recognition (ASR) before any retrieval module,...

AI 聚合
2026-08-25
HuggingFace

MobilePA-Bench: Benchmarking Mobile Planner Agents on Complex Real-World Tasks

As on-device LLM agents evolve into personal copilots, the mobile operating system has become a key testbed for this ...

AI 聚合
2026-08-25
HuggingFace

Block3D: Efficient Text-to-3D Generation via Block-Wise Diffusion

While text-to-3D generation has advanced rapidly, achieving high geometric fidelity at low inference cost remains cha...

AI 聚合
2026-08-25
HuggingFace

One Polluted Page Is Enough: Evaluating Web Content Pollution in LLM Recommenders

Search-augmented LLMs increasingly mediate everyday consumer recommendations by retrieving live web content. This cre...

AI 聚合
2026-08-25
首页 上一页 第 31 / 212 页 下一页 末页