HuggingFace

Rethinking the Evaluation of Harness Evolution for Agents

We revisit the evaluation of automatic harness evolution for LLM agents. Existing harness evolution methods use unit ...

AI 聚合
2026-07-18
arXiv

Subjective Risk Decomposition: A New View for Uncertainty Quantification

We present a novel viewpoint for uncertainty quantification. Uncertainty measures are not primitives, in need of axio...

AI 聚合
2026-07-17
arXiv

Mask-Aware Policy Gradients for Diffusion Language Models

Reinforcement learning has proven effective for improving reasoning in large language models, but extending it to Mas...

AI 聚合
2026-07-17
arXiv

Self-Evolving Human-Centered Framework for Explainable Depression Symptom Annotation

Annotation quality is a major bottleneck in building reliable and explainable artificial intelligence (XAI) systems f...

AI 聚合
2026-07-17
arXiv

MM-IssueLoc: A Controlled Benchmark for Evaluating Visual Evidence in Multimodal Repository-Level Issue Localization

Real repository issues routinely include visual evidence such as screenshots, error dialogs, rendered UI states, and ...

AI 聚合
2026-07-17
arXiv

Symbal: Detecting Systematic Misalignments in Model-Generated Captions

Multimodal large language models (MLLMs) often introduce errors when generating image captions, resulting in misalign...

AI 聚合
2026-07-17
arXiv

When Words Are Safe But Actions Kill: Probing Physical Danger Beyond Text Safety in Hidden-State Risk Space

Large language models (LLMs) increasingly serve as high-level planners for embodied agents, where linguistically beni...

AI 聚合
2026-07-17
arXiv

In-Place Tokenizer Expansion for Pre-trained LLMs

A tokenizer fixed at the start of pre-training allocates vocabulary in proportion to the pre-training corpus, reflect...

AI 聚合
2026-07-17
arXiv

AutoSynthesis: An agentic system for automated meta-analysis

Evidence synthesis is crucial for turning primary research into reliable knowledge for science, medicine, education, ...

AI 聚合
2026-07-17
arXiv

teLLMe Why (Ain't Nothing but a Jam): Exploratory Causal Analysis of Urban Driving Data

Traffic agencies now have access to large volumes of video-derived data for studying safety and congestion. Most of t...

AI 聚合
2026-07-17
arXiv

SearchOS-V1: Towards Robust Open-Domain Information-Seeking Agent Collaboration

Recent advances in Tool-Integrated Large Language Models have made web search a core capability of information-seekin...

AI 聚合
2026-07-17
arXiv

Beyond Success Rate: Cost-Aware Evaluation of Offensive and Defensive Security Agents

Security-agent evaluations commonly measure peak offensive capability under generous inference budgets, emphasizing v...

AI 聚合
2026-07-17
首页 上一页 第 126 / 216 页 下一页 末页