HuggingFace

AgentRoom: Concurrent Multi-Agent Coding in a CRDT-Backed Shared Workspace

Concurrent multi-agent coding promises division of labor across modules, robustness through redundancy, and parallel ...

AI 聚合
2026-08-27
HuggingFace

SecOPD: Mitigating Adaptive Prompt Injections by On-Policy Distillation

Prompt injection is listed as the \#1 threat to AI agents. When an agent accesses external data from websites, files,...

AI 聚合
2026-08-27
HuggingFace

Autonomous Mathematical Discovery in an Open-World Multi-Agent Environment

We study autonomous mathematical discovery in the Station, an open-world multi-agent environment in which AI agents f...

AI 聚合
2026-08-27
HuggingFace

When "Must" Becomes "Maybe": Constraint Weakening in LLM Agent Workflows

Large language model (LLM) agents coordinate complex tasks through multi-role and multi-stage workflows. Upstream sta...

AI 聚合
2026-08-27
HuggingFace

MARS: Multi-Specialist LLM Relay System for Competitive Programming

Large Language Models excel at code generation, yet competitive programming exposes a persistent failure mode: existi...

AI 聚合
2026-08-27
arXiv

Score-Based Ideal Observer Approximation via Denoising Score Matching for Signal-Known-Exactly Detection Tasks

The Bayesian Ideal Observer (IO) establishes the theoretical upper bound on task performance for binary detection tas...

AI 聚合
2026-08-26
arXiv

Ensemble of Convolutional Neural Networks for StrokePrediction: Towards Improved Diagnostic Accuracy

Brain stroke, known for its high mortality and incidence rates, poses significant health risks and requires rapid int...

AI 聚合
2026-08-26
arXiv

StepGuard: Learning Step-Level Guardrails with Scalable Supervision and Safety-Utility Balancing

LLM-based agents can interact with external environments through tool invocation, but this capability also introduces...

AI 聚合
2026-08-26
arXiv

Right Diagnoses, Decorative Reasoning:A Perturbation Audit of Medical Chain-of-Thought

Clinicians read chain-of-thought (CoT) rationales as evidence of medical reasoning, but whether the visible chain pla...

AI 聚合
2026-08-26
arXiv

CAFE: Self-Improving Search Agents Need Co-Evolving Feedback

Outcome-supervised search agents learn when and how to retrieve evidence, but terminal rewards neither localize inter...

AI 聚合
2026-08-26
arXiv

StarHarness: Evolving Harnesses with Stratified Search for Enterprise Environments

We present StarHarness, a framework for evolving environment-specific agent harnesses while keeping model weights fix...

AI 聚合
2026-08-26
arXiv

Automatic Model Card Generation Using an LLM

Model cards are structured documents that summarize key information about machine learning models to improve transpar...

AI 聚合
2026-08-26
首页 上一页 第 26 / 212 页 下一页 末页