arXiv

Measurement-Driven Sub-Network Selection for On-Premise Retrieval-Augmented Factory Agents

On-premise assistants can give factory workers conversational access to machine documentation, but models capable of ...

AI 聚合
2026-09-03
arXiv

From Reweighting to Rewriting: Unlocking the Intervention Effects of Influential Samples in Training Data Attribution

Training data attribution (TDA) aims to identify training examples that shape model behavior, but its intervention va...

AI 聚合
2026-09-03
arXiv

SafeEvolve: Harness-Policy Co-Evolution from Agent Experience for Safety Alignment

The performance of LLM-based agents is jointly shaped by the base model and the harness used when interacting with th...

AI 聚合
2026-09-03
arXiv

Dutch Books for Language Models

People increasingly use language models to support life decisions. Many such decisions involve a probabilistic foreca...

AI 聚合
2026-09-03
arXiv

frb100-40 After Two Decades: An Optimality Certificate and a Preregistered Search Study

For more than 20 years, the Model-RB benchmark frb100-40 remained an open challenge; since 2014, its public record ha...

AI 聚合
2026-09-03
arXiv

Large Language Models (LLMs) for Telecom Root Cause Analysis (RCA): A Structured Reasoning Framework for Evidence-Grounded Diagnosis

Root cause analysis (RCA) is a critical task in telecom network operations, but diagnosing performance degradations i...

AI 聚合
2026-09-03
arXiv

AI Contextual Measurement for Recovering Individual and Group-Level Effects: Validation Against Survey Measures and an Occupational Application

Researchers increasingly use artificial intelligence to construct measures of social, organizational, and occupationa...

AI 聚合
2026-09-03
arXiv

Post-Training Language Models for Gold-Medal Performance in Coding Competitions

Competitive programming has become a key test of large language model reasoning, with international competitions such...

AI 聚合
2026-09-03
arXiv

Towards Trustworthy Autonomous Robots: An Explainable AI-Based Decision Framework

Autonomous robots powered by deep learning face a fundamental auditability challenge: when incidents occur, investiga...

AI 聚合
2026-09-03
arXiv

Discriminative World Models for Web Agents

Recent web agents use world models for test-time action selection by sampling candidate actions, predicting the resul...

AI 聚合
2026-09-03
HuggingFace

Cliff: Learning Process Rewards from the First Mistake

Reinforcement learning with verifiable rewards (RLVR) has emerged as a powerful paradigm for large language model (LL...

AI 聚合
2026-09-03
HuggingFace

Kirin: Animal Motion Generation from In-the-Wild Video

Understanding animal motion is fundamental to modeling animal behavior and biomechanics, yet progress in this area la...

AI 聚合
2026-09-03
首页 上一页 第 5 / 211 页 下一页 末页